EDBT 2026 Demo / reviewers in the wild / expert
Anoop M. Mayampurath
dblp:16/366
· DBLP profile ↗
15ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0002-3010-6960ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explainable multimodal deep learning models for variable-length sequences in critically ill patients
Jennifer Martin, Majid Afshar, Askar Safipour Afshar, John R. Caskey, Dmitriy Dligach, Yanjun Gao, Jifan Gao, Guanhua Chen 0002, Anoop M. Mayampurath, Matthew M. Churpek |
J. Biomed. Informatics | 9 |
| 2025 | Simple Yet Effective: An Information-Theoretic Approach to Multi-LLM Uncertainty QuantificationabstractLarge language models (LLMs) often behave inconsistently across inputs, indicating uncertainty and motivating the need for its quantification in high-stakes settings. Prior work on calibration and uncertainty quantification often focuses on individual models, overlooking the potential of model diversity. We hypothesize that LLMs make complementary predictions due to differences in training and the Zipfian nature of language, and that aggregating their outputs leads to more reliable uncertainty estimates. To leverage this, we propose MUSE (Multi-LLM Uncertainty via Subset Ensembles), a simple information-theoretic method that uses Jensen-Shannon Divergence to identify and aggregate well-calibrated subsets of LLMs. Experiments on binary prediction tasks demonstrate improved calibration and predictive performance compared to single-model and naïve ensemble baselines. In addition, we explore using MUSE as guided signals with chain-of-thought distillation to fine-tune LLMs for calibration. MUSE is available at:https://github.com/LARK-NLP-Lab/MUSE. Maya Kruse, Majid Afshar, Saksham Khatwani, Anoop M. Mayampurath, Yanjun Gao |
EMNLP | 4 |
| 2025 | Lessons learned on information retrieval in electronic health records: a comparison of embedding models and pooling strategiesabstractOBJECTIVES: Applying large language models (LLMs) to the clinical domain is challenging due to the context-heavy nature of processing medical records. Retrieval-augmented generation (RAG) offers a solution by facilitating reasoning over large text sources. However, there are many parameters to optimize in just the retrieval system alone. This paper presents an ablation study exploring how different embedding models and pooling methods affect information retrieval for the clinical domain. MATERIALS AND METHODS: Evaluating on 3 retrieval tasks on 2 electronic health record (EHR) data sources, we compared 7 models, including medical- and general-domain models, specialized encoder embedding models, and off-the-shelf decoder LLMs. We also examine the choice of embedding pooling strategy for each model, independently on the query and the text to retrieve. RESULTS: We found that the choice of embedding model significantly impacts retrieval performance, with BGE, a comparatively small general-domain model, consistently outperforming all others, including medical-specific models. However, our findings also revealed substantial variability across datasets and query text phrasings. We also determined the best pooling methods for each of these models to guide future design of retrieval systems. DISCUSSION: The choice of embedding model, pooling strategy, and query formulation can significantly impact retrieval performance and the performance of these models on other public benchmarks does not necessarily transfer to new domains. The high variability in performance across different query phrasings suggests that the choice of query may need to be tuned and validated for each task, or even for each institution's EHR. CONCLUSION: This study provides empirical evidence to guide the selection of models and pooling strategies for RAG frameworks in healthcare applications. Further studies such as this one are vital for guiding empirically-grounded development of retrieval frameworks, such as in the context of RAG, for the clinical domain. Skatje Myers, Timothy A. Miller, Yanjun Gao, Matthew M. Churpek, Anoop M. Mayampurath, Dmitriy Dligach, Majid Afshar |
J. Am. Medical Informatics Assoc. | 5 |
| 2025 | Explaining alerts from a pediatric risk prediction model using clinical textabstractOBJECTIVE: Risk prediction models are used in hospitals to identify pediatric patients at risk of clinical deterioration, enabling timely interventions and rescue. The objective of this study was to develop a new explainer algorithm that uses a patient's clinical notes to generate text-based explanations for risk prediction alerts. MATERIALS AND METHODS: We conducted a retrospective study of 39 406 patient admissions to the American Family Children's Hospital at the University of Wisconsin-Madison (2009-2020). The pediatric Calculated Assessment of Risk and Triage (pCART) validated risk prediction model was used to identify children at risk for deterioration. A transformer model was trained to use clinical notes from the 12-hour period preceding each pCART score to predict whether a patient was flagged as at risk. Then, label-aware attention highlighted text phrases most important to an at-risk alert. The study cohort was randomly split into derivation (60%) and validation (20%) data, and a separate test (20%) was used to evaluate the explainer's performance. RESULTS: Our pCART Explainer algorithm performed well in discriminating at-risk pCART alert vs no alert (c-statistic 0.805). Sample explanations from pCART Explainer revealed clinically important phrases such as "rapid breathing," "fall risk," "distension," and "grunting," thereby demonstrating excellent face validity. DISCUSSION: The pCART Explainer could quickly orient clinicians to the patient's condition by drawing attention to key phrases in notes, potentially enhancing situational awareness and guiding decision-making. CONCLUSION: We developed pCART Explainer, a novel algorithm that highlights text within clinical notes to provide medically relevant context about deterioration alerts, thereby improving the explainability of the pCART model. Samuel Nycklemoe, Sriharsha Devarapu, Yanjun Gao, Kyle A. Carey, Nicholas Kuehnel, Neil Munjal, Priti Jani, Matthew M. Churpek, Dmitriy Dligach, Majid Afshar, Anoop M. Mayampurath |
J. Am. Medical Informatics Assoc. | 11 |
| 2024 | Development and external validation of deep learning clinical prediction models using variable-length time series dataabstractOBJECTIVES: To compare and externally validate popular deep learning model architectures and data transformation methods for variable-length time series data in 3 clinical tasks (clinical deterioration, severe acute kidney injury [AKI], and suspected infection). MATERIALS AND METHODS: This multicenter retrospective study included admissions at 2 medical centers that spanned 2007-2022. Distinct datasets were created for each clinical task, with 1 site used for training and the other for testing. Three feature engineering methods (normalization, standardization, and piece-wise linear encoding with decision trees [PLE-DTs]) and 3 architectures (long short-term memory/gated recurrent unit [LSTM/GRU], temporal convolutional network, and time-distributed wrapper with convolutional neural network [TDW-CNN]) were compared in each clinical task. Model discrimination was evaluated using the area under the precision-recall curve (AUPRC) and the area under the receiver operating characteristic curve (AUROC). RESULTS: The study comprised 373 825 admissions for training and 256 128 admissions for testing. LSTM/GRU models tied with TDW-CNN models with both obtaining the highest mean AUPRC in 2 tasks, and LSTM/GRU had the highest mean AUROC across all tasks (deterioration: 0.81, AKI: 0.92, infection: 0.87). PLE-DT with LSTM/GRU achieved the highest AUPRC in all tasks. DISCUSSION: When externally validated in 3 clinical tasks, the LSTM/GRU model architecture with PLE-DT transformed data demonstrated the highest AUPRC in all tasks. Multiple models achieved similar performance when evaluated using AUROC. CONCLUSION: The LSTM architecture performs as well or better than some newer architectures, and PLE-DT may enhance the AUPRC in variable-length time series data for predicting clinical outcomes during external validation. Fereshteh S. Bashiri, Kyle A. Carey, Jennie Martin, Jay L. Koyner, Dana P. Edelson, Emily R. Gilbert, Anoop M. Mayampurath, Majid Afshar, Matthew M. Churpek |
J. Am. Medical Informatics Assoc. | 7 |
| 2024 | Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning modelsabstractOBJECTIVE: The timely stratification of trauma injury severity can enhance the quality of trauma care but it requires intense manual annotation from certified trauma coders. The objective of this study is to develop machine learning models for the stratification of trauma injury severity across various body regions using clinical text and structured electronic health records (EHRs) data. MATERIALS AND METHODS: Our study utilized clinical documents and structured EHR variables linked with the trauma registry data to create 2 machine learning models with different approaches to representing text. The first one fuses concept unique identifiers (CUIs) extracted from free text with structured EHR variables, while the second one integrates free text with structured EHR variables. Temporal validation was undertaken to ensure the models' temporal generalizability. Additionally, analyses to assess the variable importance were conducted. RESULTS: Both models demonstrated impressive performance in categorizing leg injuries, achieving high accuracy with macro-F1 scores of over 0.8. Additionally, they showed considerable accuracy, with macro-F1 scores exceeding or near 0.7, in assessing injuries in the areas of the chest and head. We showed in our variable importance analysis that the most important features in the model have strong face validity in determining clinically relevant trauma injuries. DISCUSSION: The CUI-based model achieves comparable performance, if not higher, compared to the free-text-based model, with reduced complexity. Furthermore, integrating structured EHR data improves performance, particularly when the text modalities are insufficiently indicative. CONCLUSIONS: Our multi-modal, multiclass models can provide accurate stratification of trauma injury severity and clinically relevant interpretations. Jifan Gao, Guanhua Chen 0002, Ann P. O'Rourke, John R. Caskey, Kyle A. Carey, Madeline Oguss, Anne Stey, Dmitriy Dligach, Timothy A. Miller, Anoop M. Mayampurath, Matthew M. Churpek, Majid Afshar |
J. Am. Medical Informatics Assoc. | 10 |
| 2022 | Explaining Alerts from a Pediatric Deterioration Prediction Model Using Clinical Text
Anoop M. Mayampurath, Kyle A. Carey, Priti Jani, Majid Afshar, Matthew M. Churpek, Dmitriy Dligach |
AMIA | 1 |
| 2022 | Identifying infected patients using semi-supervised and transfer learningabstractOBJECTIVES: Early identification of infection improves outcomes, but developing models for early identification requires determining infection status with manual chart review, limiting sample size. Therefore, we aimed to compare semi-supervised and transfer learning algorithms with algorithms based solely on manual chart review for identifying infection in hospitalized patients. MATERIALS AND METHODS: This multicenter retrospective study of admissions to 6 hospitals included "gold-standard" labels of infection from manual chart review and "silver-standard" labels from nonchart-reviewed patients using the Sepsis-3 infection criteria based on antibiotic and culture orders. "Gold-standard" labeled admissions were randomly allocated to training (70%) and testing (30%) datasets. Using patient characteristics, vital signs, and laboratory data from the first 24 hours of admission, we derived deep learning and non-deep learning models using transfer learning and semi-supervised methods. Performance was compared in the gold-standard test set using discrimination and calibration metrics. RESULTS: The study comprised 432 965 admissions, of which 2724 underwent chart review. In the test set, deep learning and non-deep learning approaches had similar discrimination (area under the receiver operating characteristic curve of 0.82). Semi-supervised and transfer learning approaches did not improve discrimination over models fit using only silver- or gold-standard data. Transfer learning had the best calibration (unreliability index P value: .997, Brier score: 0.173), followed by self-learning gradient boosted machine (P value: .67, Brier score: 0.170). DISCUSSION: Deep learning and non-deep learning models performed similarly for identifying infection, as did models developed using Sepsis-3 and manual chart review labels. CONCLUSION: In a multicenter study of almost 3000 chart-reviewed patients, semi-supervised and transfer learning models showed similar performance for model discrimination as baseline XGBoost, while transfer learning improved calibration. Fereshteh S. Bashiri, John R. Caskey, Anoop M. Mayampurath, Nicole Dussault, Jay Dumanian, Sivasubramanium Bhavani, Kyle A. Carey, Emily R. Gilbert, Christopher J. Winslow, Nirav Shah 0004, Dana P. Edelson, Majid Afshar, Matthew M. Churpek |
J. Am. Medical Informatics Assoc. | 3 |
| 2021 | Sepsis Prediction Using Semi-Supervised and Transfer Learning
John R. Caskey, Fereshteh S. Bashiri, Anoop M. Mayampurath, Nicole Dussault, Jay Dumanian, Sivasubramanium Bhavani, Kyle A. Carey, Emily R. Gilbert, Christopher J. Winslow, Nirav Shah 0004, Dana P. Edelson, Majid Afshar, Matthew M. Churpek |
AMIA | 3 |
| 2019 | A structured approach to working with wearable health data: lessons learned from the University of Chicago IBD biosensor study
Philip H. Sossenheimer, Olivia Yvellez, Anoop M. Mayampurath, Samuel L. Volchenboum, David T. Rubin |
AMIA | 3 |
| 2017 | A Data-driven Framework for Sub-Typing Stem Cell Transplant Recipients at Risk for Infection
Anoop M. Mayampurath, Jonathan Matthews, Samuel L. Volchenboum, L. Nelson Sanchez-Pinto |
AMIA | 1 |
| 2013 | Automated annotation and quantification of glycans using liquid chromatography-mass spectrometryabstractUNLABELLED: As a common post-translational modification, protein glycosylation plays an important role in many biological processes, and it is known to be associated with human diseases. Mass spectrometry (MS)-based glycomic profiling techniques have been developed to measure the abundances of glycans in complex biological samples and applied to the discovery of putative glycan biomarkers. To automate the annotation of glycomic profiles in the liquid chromatography-MS (LC-MS) data, we present here a user-friendly software tool, MultiGlycan, implemented in C# on Windows systems. We tested MultiGlycan by using several glycomic profiling datasets acquired using LC-MS under different preparations and show that MultiGlycan executes fast and generates robust and reliable results. AVAILABILITY: MultiGlycan can be freely downloaded at http://darwin.informatics.indiana.edu/MultiGlycan/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chuan-Yih Yu, Anoop M. Mayampurath, Yunli Hu, Shiyue Zhou, Yehia Mechref, Haixu Tang |
Bioinform. | 2 |
| 2010 | Machine learning based prediction for peptide drift times in ion mobility spectrometryabstractMOTIVATION: Ion mobility spectrometry (IMS) has gained significant traction over the past few years for rapid, high-resolution separations of analytes based upon gas-phase ion structure, with significant potential impacts in the field of proteomic analysis. IMS coupled with mass spectrometry (MS) affords multiple improvements over traditional proteomics techniques, such as in the elucidation of secondary structure information, identification of post-translational modifications, as well as higher identification rates with reduced experiment times. The high throughput nature of this technique benefits from accurate calculation of cross sections, mobilities and associated drift times of peptides, thereby enhancing downstream data analysis. Here, we present a model that uses physicochemical properties of peptides to accurately predict a peptide's drift time directly from its amino acid sequence. This model is used in conjunction with two mathematical techniques, a partial least squares regression and a support vector regression setting. RESULTS: When tested on an experimentally created high confidence database of 8675 peptide sequences with measured drift times, both techniques statistically significantly outperform the intrinsic size parameters-based calculations, the currently held practice in the field, on all charge states (+2, +3 and +4). AVAILABILITY: The software executable, imPredict, is available for download from http:/omics.pnl.gov/software/imPredict.php CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Anuj R. Shah, Khushbu Agarwal, Erin S. Baker, Mudita Singhal, Anoop M. Mayampurath, Yehia M. Ibrahim, Lars J. Kangas, Matthew E. Monroe, Mikhail E. Belov, Gordon A. Anderson, Richard D. Smith |
Bioinform. | 5 |
| 2009 | Decon2LS: An open-source software package for automated processing and visualization of high resolution mass spectrometry dataabstractBACKGROUND: Data generated from liquid chromatography coupled to high-resolution mass spectrometry (LC-MS)-based studies of a biological sample can contain large amounts of biologically significant information in the form of proteins, peptides, and metabolites. Interpreting this data involves inferring the masses and abundances of biomolecules injected into the instrument. Because of the inherent complexity of mass spectral patterns produced by these biomolecules, the analysis is significantly enhanced by using visualization capabilities to inspect and confirm results. In this paper we describe Decon2LS, an open-source software package for automated processing and visualization of high-resolution MS data. Drawing extensively on algorithms developed over the last ten years for ICR2LS, Decon2LS packages the algorithms as a rich set of modular, reusable processing classes for performing diverse functions such as reading raw data, routine peak finding, theoretical isotope distribution modelling, and deisotoping. Because the source code is openly available, these functionalities can now be used to build derivative applications in relatively fast manner. In addition, Decon2LS provides an extensive set of visualization tools, such as high performance chart controls. RESULTS: With a variety of options that include peak processing, deisotoping, isotope composition, etc, Decon2LS supports processing of multiple raw data formats. Deisotoping can be performed on an individual scan, an individual dataset, or on multiple datasets using batch processing. Other processing options include creating a two dimensional view of mass and liquid chromatography (LC) elution time features, generating spectrum files for tandem MS data, creating total intensity chromatograms, and visualizing theoretical peptide profiles. Application of Decon2LS to deisotope different datasets obtained across different instruments yielded a high number of features that can be used to identify and quantify peptides in the biological sample. CONCLUSION: Decon2LS is an efficient software package for discovering and visualizing features in proteomics studies that require automated interpretation of mass spectra. Besides being easy to use, fast, and reliable, Decon2LS is also open-source, which allows developers in the proteomics and bioinformatics communities to reuse and refine the algorithms to meet individual needs.Decon2LS source code, installer, and tutorials may be downloaded free of charge at http://http:/ncrr.pnl.gov/software/. Navdeep Jaitly, Anoop M. Mayampurath, Kyle Littlefield, Joshua N. Adkins, Gordon A. Anderson, Richard D. Smith |
BMC Bioinform. | 2 |
| 2008 | DeconMSn: a software tool for accurate parent ion monoisotopic mass determination for tandem mass spectraabstractUNLABELLED: DeconMSn accurately determines the monoisotopic mass and charge state of parent ions from high-resolution tandem mass spectrometry data, offering significant improvement for LTQ_FT and LTQ_Orbitrap instruments over the commercially delivered Thermo Fisher Scientific's extract_msn tool. Optimal parent ion mass tolerance values can be determined using accurate mass information, thus improving peptide identifications for high-mass measurement accuracy experiments. For low-resolution data from LCQ and LTQ instruments, DeconMSn incorporates a support-vector-machine-based charge detection algorithm that identifies the most likely charge of a parent species through peak characteristics of its fragmentation pattern. AVAILABILITY: http://ncrr.pnl.gov/software/ or http://www.proteomicsresource.org/. Anoop M. Mayampurath, Navdeep Jaitly, Samuel O. Purvine, Matthew E. Monroe, Kenneth J. Auberry, Joshua N. Adkins, Richard D. Smith |
Bioinform. | 1 |