EDBT 2026 Demo / reviewers in the wild / expert
Kyle A. Carey
dblp:330/8634
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-5038-6074ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Explaining alerts from a pediatric risk prediction model using clinical textabstractOBJECTIVE: Risk prediction models are used in hospitals to identify pediatric patients at risk of clinical deterioration, enabling timely interventions and rescue. The objective of this study was to develop a new explainer algorithm that uses a patient's clinical notes to generate text-based explanations for risk prediction alerts. MATERIALS AND METHODS: We conducted a retrospective study of 39 406 patient admissions to the American Family Children's Hospital at the University of Wisconsin-Madison (2009-2020). The pediatric Calculated Assessment of Risk and Triage (pCART) validated risk prediction model was used to identify children at risk for deterioration. A transformer model was trained to use clinical notes from the 12-hour period preceding each pCART score to predict whether a patient was flagged as at risk. Then, label-aware attention highlighted text phrases most important to an at-risk alert. The study cohort was randomly split into derivation (60%) and validation (20%) data, and a separate test (20%) was used to evaluate the explainer's performance. RESULTS: Our pCART Explainer algorithm performed well in discriminating at-risk pCART alert vs no alert (c-statistic 0.805). Sample explanations from pCART Explainer revealed clinically important phrases such as "rapid breathing," "fall risk," "distension," and "grunting," thereby demonstrating excellent face validity. DISCUSSION: The pCART Explainer could quickly orient clinicians to the patient's condition by drawing attention to key phrases in notes, potentially enhancing situational awareness and guiding decision-making. CONCLUSION: We developed pCART Explainer, a novel algorithm that highlights text within clinical notes to provide medically relevant context about deterioration alerts, thereby improving the explainability of the pCART model. Samuel Nycklemoe, Sriharsha Devarapu, Yanjun Gao, Kyle A. Carey, Nicholas Kuehnel, Neil Munjal, Priti Jani, Matthew M. Churpek, Dmitriy Dligach, Majid Afshar, Anoop M. Mayampurath |
J. Am. Medical Informatics Assoc. | 4 |
| 2024 | Development and external validation of deep learning clinical prediction models using variable-length time series dataabstractOBJECTIVES: To compare and externally validate popular deep learning model architectures and data transformation methods for variable-length time series data in 3 clinical tasks (clinical deterioration, severe acute kidney injury [AKI], and suspected infection). MATERIALS AND METHODS: This multicenter retrospective study included admissions at 2 medical centers that spanned 2007-2022. Distinct datasets were created for each clinical task, with 1 site used for training and the other for testing. Three feature engineering methods (normalization, standardization, and piece-wise linear encoding with decision trees [PLE-DTs]) and 3 architectures (long short-term memory/gated recurrent unit [LSTM/GRU], temporal convolutional network, and time-distributed wrapper with convolutional neural network [TDW-CNN]) were compared in each clinical task. Model discrimination was evaluated using the area under the precision-recall curve (AUPRC) and the area under the receiver operating characteristic curve (AUROC). RESULTS: The study comprised 373 825 admissions for training and 256 128 admissions for testing. LSTM/GRU models tied with TDW-CNN models with both obtaining the highest mean AUPRC in 2 tasks, and LSTM/GRU had the highest mean AUROC across all tasks (deterioration: 0.81, AKI: 0.92, infection: 0.87). PLE-DT with LSTM/GRU achieved the highest AUPRC in all tasks. DISCUSSION: When externally validated in 3 clinical tasks, the LSTM/GRU model architecture with PLE-DT transformed data demonstrated the highest AUPRC in all tasks. Multiple models achieved similar performance when evaluated using AUROC. CONCLUSION: The LSTM architecture performs as well or better than some newer architectures, and PLE-DT may enhance the AUPRC in variable-length time series data for predicting clinical outcomes during external validation. Fereshteh S. Bashiri, Kyle A. Carey, Jennie Martin, Jay L. Koyner, Dana P. Edelson, Emily R. Gilbert, Anoop M. Mayampurath, Majid Afshar, Matthew M. Churpek |
J. Am. Medical Informatics Assoc. | 2 |
| 2024 | Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning modelsabstractOBJECTIVE: The timely stratification of trauma injury severity can enhance the quality of trauma care but it requires intense manual annotation from certified trauma coders. The objective of this study is to develop machine learning models for the stratification of trauma injury severity across various body regions using clinical text and structured electronic health records (EHRs) data. MATERIALS AND METHODS: Our study utilized clinical documents and structured EHR variables linked with the trauma registry data to create 2 machine learning models with different approaches to representing text. The first one fuses concept unique identifiers (CUIs) extracted from free text with structured EHR variables, while the second one integrates free text with structured EHR variables. Temporal validation was undertaken to ensure the models' temporal generalizability. Additionally, analyses to assess the variable importance were conducted. RESULTS: Both models demonstrated impressive performance in categorizing leg injuries, achieving high accuracy with macro-F1 scores of over 0.8. Additionally, they showed considerable accuracy, with macro-F1 scores exceeding or near 0.7, in assessing injuries in the areas of the chest and head. We showed in our variable importance analysis that the most important features in the model have strong face validity in determining clinically relevant trauma injuries. DISCUSSION: The CUI-based model achieves comparable performance, if not higher, compared to the free-text-based model, with reduced complexity. Furthermore, integrating structured EHR data improves performance, particularly when the text modalities are insufficiently indicative. CONCLUSIONS: Our multi-modal, multiclass models can provide accurate stratification of trauma injury severity and clinically relevant interpretations. Jifan Gao, Guanhua Chen 0002, Ann P. O'Rourke, John R. Caskey, Kyle A. Carey, Madeline Oguss, Anne Stey, Dmitriy Dligach, Timothy A. Miller, Anoop M. Mayampurath, Matthew M. Churpek, Majid Afshar |
J. Am. Medical Informatics Assoc. | 5 |
| 2022 | Explaining Alerts from a Pediatric Deterioration Prediction Model Using Clinical Text
Anoop M. Mayampurath, Kyle A. Carey, Priti Jani, Majid Afshar, Matthew M. Churpek, Dmitriy Dligach |
AMIA | 2 |
| 2022 | Identifying infected patients using semi-supervised and transfer learningabstractOBJECTIVES: Early identification of infection improves outcomes, but developing models for early identification requires determining infection status with manual chart review, limiting sample size. Therefore, we aimed to compare semi-supervised and transfer learning algorithms with algorithms based solely on manual chart review for identifying infection in hospitalized patients. MATERIALS AND METHODS: This multicenter retrospective study of admissions to 6 hospitals included "gold-standard" labels of infection from manual chart review and "silver-standard" labels from nonchart-reviewed patients using the Sepsis-3 infection criteria based on antibiotic and culture orders. "Gold-standard" labeled admissions were randomly allocated to training (70%) and testing (30%) datasets. Using patient characteristics, vital signs, and laboratory data from the first 24 hours of admission, we derived deep learning and non-deep learning models using transfer learning and semi-supervised methods. Performance was compared in the gold-standard test set using discrimination and calibration metrics. RESULTS: The study comprised 432 965 admissions, of which 2724 underwent chart review. In the test set, deep learning and non-deep learning approaches had similar discrimination (area under the receiver operating characteristic curve of 0.82). Semi-supervised and transfer learning approaches did not improve discrimination over models fit using only silver- or gold-standard data. Transfer learning had the best calibration (unreliability index P value: .997, Brier score: 0.173), followed by self-learning gradient boosted machine (P value: .67, Brier score: 0.170). DISCUSSION: Deep learning and non-deep learning models performed similarly for identifying infection, as did models developed using Sepsis-3 and manual chart review labels. CONCLUSION: In a multicenter study of almost 3000 chart-reviewed patients, semi-supervised and transfer learning models showed similar performance for model discrimination as baseline XGBoost, while transfer learning improved calibration. Fereshteh S. Bashiri, John R. Caskey, Anoop M. Mayampurath, Nicole Dussault, Jay Dumanian, Sivasubramanium Bhavani, Kyle A. Carey, Emily R. Gilbert, Christopher J. Winslow, Nirav Shah 0004, Dana P. Edelson, Majid Afshar, Matthew M. Churpek |
J. Am. Medical Informatics Assoc. | 7 |
| 2021 | Sepsis Prediction Using Semi-Supervised and Transfer Learning
John R. Caskey, Fereshteh S. Bashiri, Anoop M. Mayampurath, Nicole Dussault, Jay Dumanian, Sivasubramanium Bhavani, Kyle A. Carey, Emily R. Gilbert, Christopher J. Winslow, Nirav Shah 0004, Dana P. Edelson, Majid Afshar, Matthew M. Churpek |
AMIA | 7 |