EDBT 2026 Demo / reviewers in the wild / expert
Majid Afshar
dblp:235/5494
· DBLP profile ↗
35ranked-venue papers
6as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 29 · 4 first-author · 21 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LogosKG: Hardware-Optimized Scalable and Interpretable Knowledge Graph Retrievalabstract-hop retrieval on large KGs by building on symbolic KG formulations and executing traversal as hardware-efficient operations over decomposed subject, object, and relation representations. To scale to billion-edge graphs, LogosKG integrates degree-aware partitioning, cross-graph routing, and on-demand caching. Experiments show substantial efficiency gains over CPU and GPU baselines without loss of retrieval fidelity. With proven performance in KG retrieval, a downstream two-round KG-LLM interaction demonstrates how LogosKG enables large-scale, evidence-grounded analysis of how KG topology, such as hop distribution and connectivity, shapes the alignment between structured biomedical knowledge and LLM diagnostic reasoning, thereby opening the door for next-generation KG-LLM integration. The source code is publicly available at https://github.com/LARK-NLP-Lab/LogosKG, and an online demo is available at https://lark-nlp-lab-logoskg.hf.space/. He Cheng, Yifu Wu, Saksham Khatwani, Maya Kruse, Dmitriy Dligach, Timothy A. Miller, Majid Afshar, Yanjun Gao |
ACL (1) | 7 |
| 2026 | Explainable multimodal deep learning models for variable-length sequences in critically ill patients
Jennifer Martin, Majid Afshar, Askar Safipour Afshar, John R. Caskey, Dmitriy Dligach, Yanjun Gao, Jifan Gao, Guanhua Chen 0002, Anoop M. Mayampurath, Matthew M. Churpek |
J. Biomed. Informatics | 2 |
| 2025 | Simple Yet Effective: An Information-Theoretic Approach to Multi-LLM Uncertainty QuantificationabstractLarge language models (LLMs) often behave inconsistently across inputs, indicating uncertainty and motivating the need for its quantification in high-stakes settings. Prior work on calibration and uncertainty quantification often focuses on individual models, overlooking the potential of model diversity. We hypothesize that LLMs make complementary predictions due to differences in training and the Zipfian nature of language, and that aggregating their outputs leads to more reliable uncertainty estimates. To leverage this, we propose MUSE (Multi-LLM Uncertainty via Subset Ensembles), a simple information-theoretic method that uses Jensen-Shannon Divergence to identify and aggregate well-calibrated subsets of LLMs. Experiments on binary prediction tasks demonstrate improved calibration and predictive performance compared to single-model and naïve ensemble baselines. In addition, we explore using MUSE as guided signals with chain-of-thought distillation to fine-tune LLMs for calibration. MUSE is available at:https://github.com/LARK-NLP-Lab/MUSE. Maya Kruse, Majid Afshar, Saksham Khatwani, Anoop M. Mayampurath, Yanjun Gao |
EMNLP | 2 |
| 2025 | Development and validation of the provider documentation summarization quality instrument for large language modelsabstractOBJECTIVES: As large language models (LLMs) are integrated into electronic health record (EHR) workflows, validated instruments are essential to evaluate their performance before implementation and as models and documentation practices evolve. Existing instruments for provider documentation quality are often unsuitable for the complexities of LLM-generated text and lack validation on real-world data. The Provider Documentation Summarization Quality Instrument (PDSQI-9) was developed to evaluate LLM-generated clinical summaries. This study aimed to validate the PDSQI-9 across key aspects of construct validity. MATERIALS AND METHODS: Multi-document summaries were generated from real-world EHR data across multiple specialties using several LLMs (GPT-4o, Mixtral 8x7b, and Llama 3-8b). Validation included Pearson correlation analyses for substantive validity, factor analysis and Cronbach's α for structural validity, inter-rater reliability (ICC and Krippendorff's α) for generalizability, a semi-Delphi process for content validity, and comparisons of high- versus low-quality summaries for discriminant validity. Raters underwent standardized training to ensure consistent application of the instrument. RESULTS: Seven physician raters evaluated 779 summaries and answered 8329 questions, achieving over 80% power for inter-rater reliability. The PDSQI-9 demonstrated strong internal consistency (Cronbach's α = 0.879; 95% CI, 0.867-0.891) and high inter-rater reliability (ICC = 0.867; 95% CI, 0.867-0.868), supporting structural validity and generalizability. Factor analysis identified a 4-factor model explaining 58% of the variance, representing organization, clarity, accuracy, and utility. Substantive validity was supported by correlations between note length and scores for Succinct (ρ = -0.200, P = .029) and Organized (ρ = -0.190, P = .037). The semi-Delphi process ensured clinically relevant attributes, and discriminant validity distinguished high- from low-quality summaries (P<.001). DISCUSSION: The PDSQI-9 showed high inter-rater reliability, internal consistency, and a meaningful factor structure that reliably captured key dimensions of documentation quality. It distinguished between high- and low-quality summaries, supporting its practical utility for health systems needing an evaluation instrument for LLMs. CONCLUSIONS: The PDSQI-9 demonstrates robust construct validity, supporting its use in clinical practice to evaluate LLM-generated summaries and facilitate safer, more effective integration of LLMs into healthcare workflows. Emma Croxford, Yanjun Gao, Nicholas Pellegrino, Karen K. Wong, Graham Wills, Elliot First, Miranda Schnier, Kyle Burton, Cris G. Ebby, Jillian Gorskic, Matthew Kalscheur, Samy Khalil, Marie Pisani, Tyler Rubeor, Peter D. Stetson, Frank J. Liao, Cherodeep Goswami, Brian W. Patterson, Majid Afshar |
J. Am. Medical Informatics Assoc. | 19 |
| 2025 | Toward digital twins in the intensive care unit: a medication management case studyabstractOBJECTIVE: To evaluate the efficacy of digital twins developed using a large language model (LLaMA-3), fine-tuned with Low-Rank Adapters (LoRA) on intensive care units (ICU) physician notes, and to determine whether specialty-specific training enhances treatment recommendation accuracy compared to other ICU specialties or zero-shot baselines. MATERIALS AND METHODS: Digital twins were created using LLaMA-3 fine-tuned on discharge summaries from the Medical Information Mart for Intensive Care III dataset, where medications were masked to construct training and testing datasets. The medical ICU dataset (1000 notes) was used for evaluation, and performance was assessed using Bidirectional Encoder Representations from Transformers Score (BERTScore) and ROUGE-L. A zero-shot baseline model, relying solely on contextual instructions without training, was also evaluated. While our approach moves toward digital twin capabilities, it does not incorporate real-time, patient-specific electronic health records data and can be viewed as an ICU specialty-level language model adaptation. RESULTS: Models fine-tuned on medical ICU notes achieved the highest BERTScore (0.842), outperforming models trained on other specialties or mixed datasets. Zero-shot models showed the lowest performance, highlighting the importance of training. DISCUSSION: The findings demonstrate that specialty-specific training significantly improves treatment recommendation accuracy in digital twins compared to generalized or zero-shot approaches. Tailoring models to specific ICU domains strengthens their clinical decision-support capabilities. CONCLUSION: Context-specific fine-tuning of LLMs is crucial for developing effective digital twins, offering foundational insights for personalized clinical decision support. Behnaz Eslami, Majid Afshar, Samie Tootooni, Timothy A. Miller, Matthew M. Churpek, Yanjun Gao, Dmitriy Dligach |
J. Am. Medical Informatics Assoc. | 2 |
| 2025 | Lessons learned on information retrieval in electronic health records: a comparison of embedding models and pooling strategiesabstractOBJECTIVES: Applying large language models (LLMs) to the clinical domain is challenging due to the context-heavy nature of processing medical records. Retrieval-augmented generation (RAG) offers a solution by facilitating reasoning over large text sources. However, there are many parameters to optimize in just the retrieval system alone. This paper presents an ablation study exploring how different embedding models and pooling methods affect information retrieval for the clinical domain. MATERIALS AND METHODS: Evaluating on 3 retrieval tasks on 2 electronic health record (EHR) data sources, we compared 7 models, including medical- and general-domain models, specialized encoder embedding models, and off-the-shelf decoder LLMs. We also examine the choice of embedding pooling strategy for each model, independently on the query and the text to retrieve. RESULTS: We found that the choice of embedding model significantly impacts retrieval performance, with BGE, a comparatively small general-domain model, consistently outperforming all others, including medical-specific models. However, our findings also revealed substantial variability across datasets and query text phrasings. We also determined the best pooling methods for each of these models to guide future design of retrieval systems. DISCUSSION: The choice of embedding model, pooling strategy, and query formulation can significantly impact retrieval performance and the performance of these models on other public benchmarks does not necessarily transfer to new domains. The high variability in performance across different query phrasings suggests that the choice of query may need to be tuned and validated for each task, or even for each institution's EHR. CONCLUSION: This study provides empirical evidence to guide the selection of models and pooling strategies for RAG frameworks in healthcare applications. Further studies such as this one are vital for guiding empirically-grounded development of retrieval frameworks, such as in the context of RAG, for the clinical domain. Skatje Myers, Timothy A. Miller, Yanjun Gao, Matthew M. Churpek, Anoop M. Mayampurath, Dmitriy Dligach, Majid Afshar |
J. Am. Medical Informatics Assoc. | 7 |
| 2025 | Explaining alerts from a pediatric risk prediction model using clinical textabstractOBJECTIVE: Risk prediction models are used in hospitals to identify pediatric patients at risk of clinical deterioration, enabling timely interventions and rescue. The objective of this study was to develop a new explainer algorithm that uses a patient's clinical notes to generate text-based explanations for risk prediction alerts. MATERIALS AND METHODS: We conducted a retrospective study of 39 406 patient admissions to the American Family Children's Hospital at the University of Wisconsin-Madison (2009-2020). The pediatric Calculated Assessment of Risk and Triage (pCART) validated risk prediction model was used to identify children at risk for deterioration. A transformer model was trained to use clinical notes from the 12-hour period preceding each pCART score to predict whether a patient was flagged as at risk. Then, label-aware attention highlighted text phrases most important to an at-risk alert. The study cohort was randomly split into derivation (60%) and validation (20%) data, and a separate test (20%) was used to evaluate the explainer's performance. RESULTS: Our pCART Explainer algorithm performed well in discriminating at-risk pCART alert vs no alert (c-statistic 0.805). Sample explanations from pCART Explainer revealed clinically important phrases such as "rapid breathing," "fall risk," "distension," and "grunting," thereby demonstrating excellent face validity. DISCUSSION: The pCART Explainer could quickly orient clinicians to the patient's condition by drawing attention to key phrases in notes, potentially enhancing situational awareness and guiding decision-making. CONCLUSION: We developed pCART Explainer, a novel algorithm that highlights text within clinical notes to provide medically relevant context about deterioration alerts, thereby improving the explainability of the pCART model. Samuel Nycklemoe, Sriharsha Devarapu, Yanjun Gao, Kyle A. Carey, Nicholas Kuehnel, Neil Munjal, Priti Jani, Matthew M. Churpek, Dmitriy Dligach, Majid Afshar, Anoop M. Mayampurath |
J. Am. Medical Informatics Assoc. | 10 |
| 2025 | LCD benchmark: long clinical document benchmark on mortality prediction for language modelsabstractOBJECTIVES: The application of natural language processing (NLP) in the clinical domain is important due to the rich unstructured information in clinical documents, which often remains inaccessible in structured data. When applying NLP methods to a certain domain, the role of benchmark datasets is crucial as benchmark datasets not only guide the selection of best-performing models but also enable the assessment of the reliability of the generated outputs. Despite the recent availability of language models capable of longer context, benchmark datasets targeting long clinical document classification tasks are absent. MATERIALS AND METHODS: To address this issue, we propose Long Clinical Document (LCD) benchmark, a benchmark for the task of predicting 30-day out-of-hospital mortality using discharge notes of Medical Information Mart for Intensive Care IV and statewide death data. We evaluated this benchmark dataset using baseline models, from bag-of-words and convolutional neural network to instruction-tuned large language models. Additionally, we provide a comprehensive analysis of the model outputs, including manual review and visualization of model weights, to offer insights into their predictive capabilities and limitations. RESULTS: Baseline models showed 28.9% for best-performing supervised models and 32.2% for GPT-4 in F1 metrics. Notes in our dataset have a median word count of 1687. DISCUSSION: Our analysis of the model outputs showed that our dataset is challenging for both models and human experts, but the models can find meaningful signals from the text. CONCLUSION: We expect our LCD benchmark to be a resource for the development of advanced supervised models, or prompting methods, tailored for clinical text. Wonjin Yoon, Shan Chen 0004, Yanjun Gao, Zhanzhan Zhao, Dmitriy Dligach, Danielle S. Bitterman, Majid Afshar, Timothy A. Miller |
J. Am. Medical Informatics Assoc. | 7 |
| 2024 | Development and external validation of deep learning clinical prediction models using variable-length time series dataabstractOBJECTIVES: To compare and externally validate popular deep learning model architectures and data transformation methods for variable-length time series data in 3 clinical tasks (clinical deterioration, severe acute kidney injury [AKI], and suspected infection). MATERIALS AND METHODS: This multicenter retrospective study included admissions at 2 medical centers that spanned 2007-2022. Distinct datasets were created for each clinical task, with 1 site used for training and the other for testing. Three feature engineering methods (normalization, standardization, and piece-wise linear encoding with decision trees [PLE-DTs]) and 3 architectures (long short-term memory/gated recurrent unit [LSTM/GRU], temporal convolutional network, and time-distributed wrapper with convolutional neural network [TDW-CNN]) were compared in each clinical task. Model discrimination was evaluated using the area under the precision-recall curve (AUPRC) and the area under the receiver operating characteristic curve (AUROC). RESULTS: The study comprised 373 825 admissions for training and 256 128 admissions for testing. LSTM/GRU models tied with TDW-CNN models with both obtaining the highest mean AUPRC in 2 tasks, and LSTM/GRU had the highest mean AUROC across all tasks (deterioration: 0.81, AKI: 0.92, infection: 0.87). PLE-DT with LSTM/GRU achieved the highest AUPRC in all tasks. DISCUSSION: When externally validated in 3 clinical tasks, the LSTM/GRU model architecture with PLE-DT transformed data demonstrated the highest AUPRC in all tasks. Multiple models achieved similar performance when evaluated using AUROC. CONCLUSION: The LSTM architecture performs as well or better than some newer architectures, and PLE-DT may enhance the AUPRC in variable-length time series data for predicting clinical outcomes during external validation. Fereshteh S. Bashiri, Kyle A. Carey, Jennie Martin, Jay L. Koyner, Dana P. Edelson, Emily R. Gilbert, Anoop M. Mayampurath, Majid Afshar, Matthew M. Churpek |
J. Am. Medical Informatics Assoc. | 8 |
| 2024 | Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning modelsabstractOBJECTIVE: The timely stratification of trauma injury severity can enhance the quality of trauma care but it requires intense manual annotation from certified trauma coders. The objective of this study is to develop machine learning models for the stratification of trauma injury severity across various body regions using clinical text and structured electronic health records (EHRs) data. MATERIALS AND METHODS: Our study utilized clinical documents and structured EHR variables linked with the trauma registry data to create 2 machine learning models with different approaches to representing text. The first one fuses concept unique identifiers (CUIs) extracted from free text with structured EHR variables, while the second one integrates free text with structured EHR variables. Temporal validation was undertaken to ensure the models' temporal generalizability. Additionally, analyses to assess the variable importance were conducted. RESULTS: Both models demonstrated impressive performance in categorizing leg injuries, achieving high accuracy with macro-F1 scores of over 0.8. Additionally, they showed considerable accuracy, with macro-F1 scores exceeding or near 0.7, in assessing injuries in the areas of the chest and head. We showed in our variable importance analysis that the most important features in the model have strong face validity in determining clinically relevant trauma injuries. DISCUSSION: The CUI-based model achieves comparable performance, if not higher, compared to the free-text-based model, with reduced complexity. Furthermore, integrating structured EHR data improves performance, particularly when the text modalities are insufficiently indicative. CONCLUSIONS: Our multi-modal, multiclass models can provide accurate stratification of trauma injury severity and clinically relevant interpretations. Jifan Gao, Guanhua Chen 0002, Ann P. O'Rourke, John R. Caskey, Kyle A. Carey, Madeline Oguss, Anne Stey, Dmitriy Dligach, Timothy A. Miller, Anoop M. Mayampurath, Matthew M. Churpek, Majid Afshar |
J. Am. Medical Informatics Assoc. | 12 |
| 2024 | On the role of the UMLS in supporting diagnosis generation proposed by Large Language Models
Majid Afshar, Yanjun Gao, Emma Croxford, Dina Demner-Fushman |
J. Biomed. Informatics | 1 |
| 2023 | Improving model transferability for clinical note section classification models using continued pretrainingabstractOBJECTIVE: The classification of clinical note sections is a critical step before doing more fine-grained natural language processing tasks such as social determinants of health extraction and temporal information extraction. Often, clinical note section classification models that achieve high accuracy for 1 institution experience a large drop of accuracy when transferred to another institution. The objective of this study is to develop methods that classify clinical note sections under the SOAP ("Subjective," "Object," "Assessment," and "Plan") framework with improved transferability. MATERIALS AND METHODS: We trained the baseline models by fine-tuning BERT-based models, and enhanced their transferability with continued pretraining, including domain-adaptive pretraining and task-adaptive pretraining. We added in-domain annotated samples during fine-tuning and observed model performance over a varying number of annotated sample size. Finally, we quantified the impact of continued pretraining in equivalence of the number of in-domain annotated samples added. RESULTS: We found continued pretraining improved models only when combined with in-domain annotated samples, improving the F1 score from 0.756 to 0.808, averaged across 3 datasets. This improvement was equivalent to adding 35 in-domain annotated samples. DISCUSSION: Although considered a straightforward task when performing in-domain, section classification is still a considerably difficult task when performing cross-domain, even using highly sophisticated neural network-based methods. CONCLUSION: Continued pretraining improved model transferability for cross-domain clinical note section classification in the presence of a small amount of in-domain labeled samples. Weipeng Zhou, Meliha Yetisgen, Majid Afshar, Yanjun Gao, Guergana K. Savova, Timothy A. Miller |
J. Am. Medical Informatics Assoc. | 3 |
| 2023 | DR.BENCH: Diagnostic Reasoning Benchmark for Clinical Natural Language Processing
Yanjun Gao, Dmitriy Dligach, Timothy A. Miller, John R. Caskey, Brihat Sharma, Matthew M. Churpek, Majid Afshar |
J. Biomed. Informatics | 7 |
| 2023 | Progress Note Understanding - Assessment and Plan Reasoning: Overview of the 2022 N2C2 Track 3 shared task
Yanjun Gao, Dmitriy Dligach, Timothy A. Miller, Matthew M. Churpek, Özlem Uzuner, Majid Afshar |
J. Biomed. Informatics | 6 |
| 2022 | Explaining Alerts from a Pediatric Deterioration Prediction Model Using Clinical Text
Anoop M. Mayampurath, Kyle A. Carey, Priti Jani, Majid Afshar, Matthew M. Churpek, Dmitriy Dligach |
AMIA | 4 |
| 2022 | Summarizing Patients' Problems from Hospital Progress Notes Using Pre-trained Sequence-to-Sequence ModelsabstractAutomatically summarizing patients’ main problems from daily progress notes using natural language processing methods helps to battle against information and cognitive overload in hospital settings and potentially assists providers with computerized diagnostic decision support. Problem list summarization requires a model to understand, abstract, and generate clinical documentation. In this work, we propose a new NLP task that aims to generate a list of problems in a patient’s daily care plan using input from the provider’s progress notes during hospitalization. We investigate the performance of T5 and BART, two state-of-the-art seq2seq transformer architectures, in solving this problem. We provide a corpus built on top of progress notes from publicly available electronic health record progress notes in the Medical Information Mart for Intensive Care (MIMIC)-III. T5 and BART are trained on general domain text, and we experiment with a data augmentation method and a domain adaptation pre-training method to increase exposure to medical vocabulary and knowledge. Evaluation methods include ROUGE, BERTScore, cosine similarity on sentence embedding, and F-score on medical concepts. Results show that T5 with domain adaptive pre-training achieves significant performance gains compared to a rule-based system and general domain pre-trained language models, indicating a promising direction for tackling the problem summarization task. Yanjun Gao, Dmitriy Dligach, Timothy A. Miller, Dongfang Xu, Matthew M. Churpek, Majid Afshar |
COLING | 6 |
| 2022 | Hierarchical Annotation for Building A Suite of Clinical Natural Language Processing Tasks: Progress Note UnderstandingabstractApplying methods in natural language processing on electronic health records (EHR) data has attracted rising interests. Existing corpus and annotation focus on modeling textual features and relation prediction. However, there are a paucity of annotated corpus built to model clinical diagnostic thinking, a processing involving text understanding, domain knowledge abstraction and reasoning. In this work, we introduce a hierarchical annotation schema with three stages to address clinical text understanding, clinical reasoning and summarization. We create an annotated corpus based on a large collection of publicly available daily progress notes, a type of EHR that is time-sensitive, problem-oriented, and well-documented by the format of Subjective, Objective, Assessment and Plan (SOAP). We also define a new suite of tasks, Progress Note Understanding, with three tasks utilizing the three annotation stages. This new suite aims at training and evaluating future NLP models for clinical text understanding, clinical knowledge representation, inference and summarization. Yanjun Gao, Dmitriy Dligach, Timothy A. Miller, Samuel Tesch, Ryan Laffin, Matthew M. Churpek, Majid Afshar |
LREC | 7 |
| 2022 | Optimizing feature selection methods by removing irrelevant features using sparse least squares
Majid Afshar, Hamid Usefi |
Expert Syst. Appl. | 1 |
| 2022 | Identifying infected patients using semi-supervised and transfer learningabstractOBJECTIVES: Early identification of infection improves outcomes, but developing models for early identification requires determining infection status with manual chart review, limiting sample size. Therefore, we aimed to compare semi-supervised and transfer learning algorithms with algorithms based solely on manual chart review for identifying infection in hospitalized patients. MATERIALS AND METHODS: This multicenter retrospective study of admissions to 6 hospitals included "gold-standard" labels of infection from manual chart review and "silver-standard" labels from nonchart-reviewed patients using the Sepsis-3 infection criteria based on antibiotic and culture orders. "Gold-standard" labeled admissions were randomly allocated to training (70%) and testing (30%) datasets. Using patient characteristics, vital signs, and laboratory data from the first 24 hours of admission, we derived deep learning and non-deep learning models using transfer learning and semi-supervised methods. Performance was compared in the gold-standard test set using discrimination and calibration metrics. RESULTS: The study comprised 432 965 admissions, of which 2724 underwent chart review. In the test set, deep learning and non-deep learning approaches had similar discrimination (area under the receiver operating characteristic curve of 0.82). Semi-supervised and transfer learning approaches did not improve discrimination over models fit using only silver- or gold-standard data. Transfer learning had the best calibration (unreliability index P value: .997, Brier score: 0.173), followed by self-learning gradient boosted machine (P value: .67, Brier score: 0.170). DISCUSSION: Deep learning and non-deep learning models performed similarly for identifying infection, as did models developed using Sepsis-3 and manual chart review labels. CONCLUSION: In a multicenter study of almost 3000 chart-reviewed patients, semi-supervised and transfer learning models showed similar performance for model discrimination as baseline XGBoost, while transfer learning improved calibration. Fereshteh S. Bashiri, John R. Caskey, Anoop M. Mayampurath, Nicole Dussault, Jay Dumanian, Sivasubramanium Bhavani, Kyle A. Carey, Emily R. Gilbert, Christopher J. Winslow, Nirav Shah 0004, Dana P. Edelson, Majid Afshar, Matthew M. Churpek |
J. Am. Medical Informatics Assoc. | 12 |
| 2022 | A scoping review of publicly available language tasks in clinical natural language processingabstractOBJECTIVE: To provide a scoping review of papers on clinical natural language processing (NLP) shared tasks that use publicly available electronic health record data from a cohort of patients. MATERIALS AND METHODS: We searched 6 databases, including biomedical research and computer science literature databases. A round of title/abstract screening and full-text screening were conducted by 2 reviewers. Our method followed the PRISMA-ScR guidelines. RESULTS: A total of 35 papers with 48 clinical NLP tasks met inclusion criteria between 2007 and 2021. We categorized the tasks by the type of NLP problems, including named entity recognition, summarization, and other NLP tasks. Some tasks were introduced as potential clinical decision support applications, such as substance abuse detection, and phenotyping. We summarized the tasks by publication venue and dataset type. DISCUSSION: The breadth of clinical NLP tasks continues to grow as the field of NLP evolves with advancements in language systems. However, gaps exist with divergent interests between the general domain NLP community and the clinical informatics community for task motivation and design, and in generalizability of the data sources. We also identified issues in data preparation. CONCLUSION: The existing clinical NLP tasks cover a wide range of topics and the field is expected to grow and attract more attention from both general domain NLP and clinical informatics community. We encourage future work to incorporate multidisciplinary collaboration, reporting transparency, and standardization in data preparation. We provide a listing of all the shared task papers and datasets from this review in a GitLab repository. Yanjun Gao, Dmitriy Dligach, Leslie Christensen, Samuel Tesch, Ryan Laffin, Dongfang Xu, Timothy A. Miller, Özlem Uzuner, Matthew M. Churpek, Majid Afshar |
J. Am. Medical Informatics Assoc. | 10 |
| 2021 | Bias Assessment and Correction in Machine Learning Algorithms: A Use-Case in a Natural Language Processing Algorithm to Identify Hospitalized Patients with Unhealthy Alcohol Use
Marissa Borgese, Cara Joyce, Emily E. Anderson, Matthew M. Churpek, Majid Afshar |
AMIA | 5 |
| 2021 | Sepsis Prediction Using Semi-Supervised and Transfer Learning
John R. Caskey, Fereshteh S. Bashiri, Anoop M. Mayampurath, Nicole Dussault, Jay Dumanian, Sivasubramanium Bhavani, Kyle A. Carey, Emily R. Gilbert, Christopher J. Winslow, Nirav Shah 0004, Dana P. Edelson, Majid Afshar, Matthew M. Churpek |
AMIA | 12 |
| 2021 | A multi-label classifier to screen different types of substance misuse in hospitalized patients
Brihat Sharma, Dmitriy Dligach, Hale Thomson, Matthew M. Churpek, Niranjan S. Karnik, Majid Afshar |
AMIA | 6 |
| 2021 | The Addition of United States Census-Tract Data Does Not Improve the Prediction of Substance Misuse
Daniel To, Cara Joyce, Sujay Kulshrestha, Brihat Sharma, Dmitriy Dligach, Matthew M. Churpek, Majid Afshar |
AMIA | 7 |
| 2021 | Bias and fairness assessment of a natural language processing opioid misuse classifier: detection and mitigation of electronic health record data disadvantages across racial subgroupsabstractOBJECTIVES: To assess fairness and bias of a previously validated machine learning opioid misuse classifier. MATERIALS & METHODS: Two experiments were conducted with the classifier's original (n = 1000) and external validation (n = 53 974) datasets from 2 health systems. Bias was assessed via testing for differences in type II error rates across racial/ethnic subgroups (Black, Hispanic/Latinx, White, Other) using bootstrapped 95% confidence intervals. A local surrogate model was estimated to interpret the classifier's predictions by race and averaged globally from the datasets. Subgroup analyses and post-hoc recalibrations were conducted to attempt to mitigate biased metrics. RESULTS: We identified bias in the false negative rate (FNR = 0.32) of the Black subgroup compared to the FNR (0.17) of the White subgroup. Top features included "heroin" and "substance abuse" across subgroups. Post-hoc recalibrations eliminated bias in FNR with minimal changes in other subgroup error metrics. The Black FNR subgroup had higher risk scores for readmission and mortality than the White FNR subgroup, and a higher mortality risk score than the Black true positive subgroup (P < .05). DISCUSSION: The Black FNR subgroup had the greatest severity of disease and risk for poor outcomes. Similar features were present between subgroups for predicting opioid misuse, but inequities were present. Post-hoc mitigation techniques mitigated bias in type II error rate without creating substantial type I error rates. From model design through deployment, bias and data disadvantages should be systematically addressed. CONCLUSION: Standardized, transparent bias assessments are needed to improve trustworthiness in clinical machine learning models. Hale M. Thompson, Brihat Sharma, Sameer Bhalla, Randy Boley, Connor McCluskey, Dmitriy Dligach, Matthew M. Churpek, Niranjan S. Karnik, Majid Afshar |
J. Am. Medical Informatics Assoc. | 9 |
| 2021 | Pre-training phenotyping classifiers
Dmitriy Dligach, Majid Afshar, Timothy A. Miller |
J. Biomed. Informatics | 2 |
| 2020 | Learning Hierarchical Transformer-based Representations of Clinical Notes
Xin Su 0008, Timothy A. Miller, Majid Afshar, Dmitriy Dligach |
AMIA | 3 |
| 2020 | High-dimensional feature selection for genomic datasets
Majid Afshar, Hamid Usefi |
Knowl. Based Syst. | 1 |
| 2019 | Towards a Universal Document-Level Clinical Text Encoder: Methods for Neural Network Pre-training with Applications to Substance Misuse
Dmitriy Dligach, Majid Afshar, Timothy A. Miller |
AMIA | 2 |
| 2019 | Untapped Potential of Clinical Text for Opioid Surveillance
Amy L. Olex, Tamás Gál, Majid Afshar, Dmitriy Dligach, Niranjan S. Karnik, Travis Oakes, Brihat Sharma, Meng Xie, Bridget T. McInnes, Julian Solway, Abel N. Kho, William Cramer, F. G. Moeller |
AMIA | 3 |
| 2019 | Identification of Latent Subtypes of Patients with Opioid Misuse
Brihat Sharma, Majid Afshar, Dmitriy Dligach, Robert Kanie, Elizabeth Salisbury-Afshar, Niranjan S. Karnik, Cara Joyce |
AMIA | 2 |
| 2019 | Development and application of a high throughput natural language processing architecture to convert all clinical documents in a clinical data warehouse into standardized medical vocabulariesabstractOBJECTIVE: Natural language processing (NLP) engines such as the clinical Text Analysis and Knowledge Extraction System are a solution for processing notes for research, but optimizing their performance for a clinical data warehouse remains a challenge. We aim to develop a high throughput NLP architecture using the clinical Text Analysis and Knowledge Extraction System and present a predictive model use case. MATERIALS AND METHODS: The CDW was comprised of 1 103 038 patients across 10 years. The architecture was constructed using the Hadoop data repository for source data and 3 large-scale symmetric processing servers for NLP. Each named entity mention in a clinical document was mapped to the Unified Medical Language System concept unique identifier (CUI). RESULTS: The NLP architecture processed 83 867 802 clinical documents in 13.33 days and produced 37 721 886 606 CUIs across 8 standardized medical vocabularies. Performance of the architecture exceeded 500 000 documents per hour across 30 parallel instances of the clinical Text Analysis and Knowledge Extraction System including 10 instances dedicated to documents greater than 20 000 bytes. In a use-case example for predicting 30-day hospital readmission, a CUI-based model had similar discrimination to n-grams with an area under the curve receiver operating characteristic of 0.75 (95% CI, 0.74-0.76). DISCUSSION AND CONCLUSION: Our health system's high throughput NLP architecture may serve as a benchmark for large-scale clinical research using a CUI-based approach. Majid Afshar, Dmitriy Dligach, Brihat Sharma, Xiaoyuan Cai, Jason Boyda, Steven Birch, Daniel Valdez, Susan J. Zelisko, Cara Joyce, François Modave, Ron Price |
J. Am. Medical Informatics Assoc. | 1 |
| 2019 | Natural language processing and machine learning to identify alcohol misuse from the electronic health record in trauma patients: development and internal validationabstractObjective: Alcohol misuse is present in over a quarter of trauma patients. Information in the clinical notes of the electronic health record of trauma patients may be used for phenotyping tasks with natural language processing (NLP) and supervised machine learning. The objective of this study is to train and validate an NLP classifier for identifying patients with alcohol misuse. Materials and Methods: An observational cohort of 1422 adult patients admitted to a trauma center between April 2013 and November 2016. Linguistic processing of clinical notes was performed using the clinical Text Analysis and Knowledge Extraction System. The primary analysis was the binary classification of alcohol misuse. The Alcohol Use Disorders Identification Test served as the reference standard. Results: The data corpus comprised 91 045 electronic health record notes and 16 091 features. In the final machine learning classifier, 16 features were selected from the first 24 hours of notes for identifying alcohol misuse. The classifier's performance in the validation cohort had an area under the receiver-operating characteristic curve of 0.78 (95% confidence interval [CI], 0.72 to 0.85). Sensitivity and specificity were at 56.0% (95% CI, 44.1% to 68.0%) and 88.9% (95% CI, 84.4% to 92.8%). The Hosmer-Lemeshow goodness-of-fit test demonstrates the classifier fits the data well (P = .17). A simpler rule-based keyword approach had a decrease in sensitivity when compared with the NLP classifier from 56.0% to 18.2%. Conclusions: The NLP classifier has adequate predictive validity for identifying alcohol misuse in trauma centers. External validation is needed before its application to augment screening. Majid Afshar, Andrew Phillips, Niranjan S. Karnik, Jeanne Mueller, Daniel To, Richard Gonzalez, Ron Price, Richard S. Cooper, Cara Joyce, Dmitriy Dligach |
J. Am. Medical Informatics Assoc. | 1 |
| 2019 | Toward a clinical text encoder: pretraining for clinical natural language processing with applications to substance misuseabstractOBJECTIVE: Our objective is to develop algorithms for encoding clinical text into representations that can be used for a variety of phenotyping tasks. MATERIALS AND METHODS: Obtaining large datasets to take advantage of highly expressive deep learning methods is difficult in clinical natural language processing (NLP). We address this difficulty by pretraining a clinical text encoder on billing code data, which is typically available in abundance. We explore several neural encoder architectures and deploy the text representations obtained from these encoders in the context of clinical text classification tasks. While our ultimate goal is learning a universal clinical text encoder, we also experiment with training a phenotype-specific encoder. A universal encoder would be more practical, but a phenotype-specific encoder could perform better for a specific task. RESULTS: We successfully train several clinical text encoders, establish a new state-of-the-art on comorbidity data, and observe good performance gains on substance misuse data. DISCUSSION: We find that pretraining using billing codes is a promising research direction. The representations generated by this type of pretraining have universal properties, as they are highly beneficial for many phenotyping tasks. Phenotype-specific pretraining is a viable route for trading the generality of the pretrained encoder for better performance on a specific phenotyping task. CONCLUSIONS: We successfully applied our approach to many phenotyping tasks. We conclude by discussing potential limitations of our approach. Dmitriy Dligach, Majid Afshar, Timothy A. Miller |
J. Am. Medical Informatics Assoc. | 2 |
| 2018 | A Computable Phenotype for Acute Respiratory Distress Syndrome Using Natural Language Processing and Machine Learning
Majid Afshar, Cara Joyce, Anthony Oakey, Perry Formanek, Philip Yang, Matthew M. Churpek, Richard S. Cooper, Ron Price, Susan J. Zelisko, Dmitriy Dligach |
AMIA | 1 |