Justin F. Rousseau

dblp:212/8498 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-2817-9124ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SDoH-GPT: using large language models to extract social determinants of health
abstract
OBJECTIVE: Extracting social determinants of health (SDoHs) from medical notes depends heavily on labor-intensive annotations, which are typically task-specific, hampering reusability and limiting sharing. Here, we introduce SDoH-GPT, a novel framework leveraging few-shot learning large language models (LLMs) to automate the extraction of SDoH from unstructured text, aiming to improve both efficiency and generalizability. MATERIALS AND METHODS: SDoH-GPT is a framework including the few-shot learning LLM methods to extract the SDoH from medical notes and the XGBoost classifiers which continue to classify SDoH using the annotations generated by the few-shot learning LLM methods as training datasets. The unique combination of the few-shot learning LLM methods with XGBoost utilizes the strength of LLMs as great few shot learners and the efficiency of XGBoost when the training dataset is sufficient. Therefore, SDoH-GPT can extract SDoH without relying on extensive medical annotations or costly human intervention. RESULTS: Our approach achieved tenfold and twentyfold reductions in time and cost, respectively, and superior consistency with human annotators measured by Cohen's kappa of up to 0.92. The innovative combination of LLM and XGBoost can ensure high accuracy and computational efficiency while consistently maintaining 0.90+ AUROC scores. DISCUSSION: This study has verified SDoH-GPT on three datasets and highlights the potential of leveraging LLM and XGBoost to revolutionize medical note classification, demonstrating its capability to achieve highly accurate classifications with significantly reduced time and cost. CONCLUSION: The key contribution of this study is the integration of LLM with XGBoost, which enables cost-effective and high quality annotations of SDoH. This research sets the stage for SDoH can be more accessible, scalable, and impactful in driving future healthcare solutions.
Bernardo Scapini Consoli, Xizhi Wu, Song Wang 0026, Yanshan Wang, Justin F. Rousseau, Thomas Hartvigsen, Li Shen 0001, Huanmei Wu, Yifan Peng 0002, Qi Long, Tianlong Chen 0001, Ying Ding 0001
J. Am. Medical Informatics Assoc.7
2026 CPGPrompt: translating clinical guidelines into large language model-executable decision support
abstract
OBJECTIVE: Clinical practice guidelines (CPGs) provide evidence-based recommendations for patient care; however, integrating them into artificial intelligence (AI) remains challenging. Previous approaches, such as rule-based systems or black-box AI models, face significant limitations, including poor interpretability, inconsistent adherence to guidelines, and narrow domain applicability. To address this, we develop and validate CPGPrompt, an auto-prompting system that converts narrative clinical guidelines into large language models (LLMs). MATERIALS AND METHODS: Our framework translates CPGs into structured decision trees and utilizes an LLM to dynamically navigate them for patient case evaluation. Synthetic vignettes were generated across 3 domains-headache, lower back pain, and prostate cancer-and distributed into 4 categories to test different decision scenarios. System performance was assessed on both binary specialty referral decisions and fine-grained pathway classification tasks. RESULTS: The binary specialty referral classification achieved consistently strong performance across all domains (F1: 0.85-1.00), with high recall (1.00 ± 0.00). In contrast, multiclass pathway assignment showed reduced performance, with domain-specific variations: headache (F1: 0.47), lower back pain (F1: 0.72), and prostate cancer (F1: 0.77). DISCUSSION: Domain-specific performance differences reflected the structure of each guideline. The headache guideline highlighted challenges with negation handling. The lower back pain guideline required temporal reasoning. In contrast, prostate cancer pathways benefited from quantifiable laboratory tests, resulting in more reliable decision-making. CONCLUSION: CPGPrompt demonstrates generalizability across diverse clinical domains while maintaining high sensitivity for referral decisions. Its transparent, auditable framework enables the systematic identification of failure modes and provides advantages over black-box AI approaches. However, persistent challenges with subjective clinical assessments indicate a need for targeted improvements and greater clinical robustness.
Ruiqi Deng, Geoffrey Martin, Tony Wang, Yi Liu 0059, Chunhua Weng, Yanshan Wang, Justin F. Rousseau, Yifan Peng 0002
J. Am. Medical Informatics Assoc.8
2023 Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error Detectors
abstract
Liyan Tang, Tanya Goyal, Alex Fabbri, Philippe Laban, Jiacheng Xu, Semih Yavuz, Wojciech Kryscinski, Justin Rousseau, Greg Durrett. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Liyan Tang, Tanya Goyal, Alexander R. Fabbri, Philippe Laban, Jiacheng Xu 0001, Semih Yavuz, Wojciech Kryscinski, Justin F. Rousseau, Greg Durrett
ACL (1)8
2023 Attend Who is Weak: Pruning-assisted Medical Image Localization under Sophisticated and Implicit Imbalances
abstract
Deep neural networks (DNNs) have rapidly become a de facto choice for medical image understanding tasks. However, DNNs are notoriously fragile to the class imbalance in image classification. We further point out that such imbalance fragility can be amplified when it comes to more sophisticated tasks such as pathology localization, as imbalances in such problems can have highly complex and often implicit forms of presence. For example, different pathology can have different sizes or colors (w.r.t.the background), different underlying demographic distributions, and in general different difficulty levels to recognize, even in a meticulously curated balanced distribution of training data. In this paper, we propose to use pruning to automatically and adaptively identify hard-to-learn (HTL) training samples, and improve pathology localization by attending them explicitly, during training in supervised, semi-supervised, and weakly-supervised settings. Our main inspiration is drawn from the recent finding that deep classification models have difficult-to-memorize samples and those may be effectively exposed through network pruning [15] - and we extend such observation beyond classification for the first time. We also present an interesting demographic analysis which illustrates HTLs ability to capture complex demographic imbalances. Our extensive experiments on the Skin Lesion Localization task in multiple training settings by paying additional attention to HTLs show significant improvement of localization performance by ~2-3%.
Ajay Jaiswal, Tianlong Chen 0001, Justin F. Rousseau, Yifan Peng 0002, Ying Ding 0001, Zhangyang Wang
WACV3
2022 Analyzing Impact of Socio-Economic Factors on COVID-19 Mortality Prediction Using SHAP Value
Redoan Rahman, Jooyeong Kang, Justin F. Rousseau, Ying Ding 0001
AMIA3
2022 Prompt-based Learning for Assertion Classification in Clinical Notes
Song Wang 0026, Liyan Tang, Akash Majety, Justin F. Rousseau, George Shih, Ying Ding 0001, Yifan Peng 0002
AMIA4
2022 RoS-KD: A Robust Stochastic Knowledge Distillation Approach for Noisy Medical Imaging
abstract
AI-powered Medical Imaging has recently achieved enormous attention due to its ability to provide fast-paced healthcare diagnoses. However, it usually suffers from a lack of high-quality datasets due to high annotation cost, interobserver variability, human annotator error, and errors in computer-generated labels. Deep learning models trained on noisy labelled datasets are sensitive to the noise type and lead to less generalization on the unseen samples. To address this challenge, we propose a Robust Stochastic Knowledge Distillation (RoS-KD) framework which mimics the notion of learning a topic from multiple sources to ensure deterrence in learning noisy information. More specifically, RoS-KD learns a smooth, well-informed, and robust student manifold by distilling knowledge from multiple teachers trained on overlapping subsets of training data. Our extensive experiments on popular medical imaging classification tasks (cardiopulmonary disease and lesion classification) using real-world datasets, show the performance benefit of RoS-KD, its ability to distill knowledge from many popular large networks (ResNet-50, DenseNet-121, MobileNetV2) in a comparatively small network, and its robustness to adversarial attacks (PGD, FSGM). More specifically, RoS-KD achieves >2% and > 4% improvement on F1-score for lesion classification and cardiopulmonary disease classification tasks, respectively, when the underlying student is ResNet-18 against recent competitive knowledge distillation baseline. Additionally, on cardiopulmonary disease classification task, RoS-KD outperforms most of the SOTA baselines by ~1% gain in AUC score.
Ajay Jaiswal, Kumar Ashutosh, Justin F. Rousseau, Yifan Peng 0002, Zhangyang Wang, Ying Ding 0001
ICDM3
2022 Old can be Gold: Better Gradient Flow can Make Vanilla-GCNs Great Again
abstract
Despite the enormous success of Graph Convolutional Networks (GCNs) in modeling graph-structured data, most of the current GCNs are shallow due to the notoriously challenging problems of over-smoothening and information squashing along with conventional difficulty caused by vanishing gradients and over-fitting. Previous works have been primarily focused on the study of over-smoothening and over-squashing phenomena in training deep GCNs. Surprisingly, in comparison with CNNs/RNNs, very limited attention has been given to understanding how healthy gradient flow can benefit the trainability of deep GCNs. In this paper, firstly, we provide a new perspective of gradient flow to understand the substandard performance of deep GCNs and hypothesize that by facilitating healthy gradient flow, we can significantly improve their trainability, as well as achieve state-of-the-art (SOTA) level performance from vanilla-GCNs. Next, we argue that blindly adopting the Glorot initialization for GCNs is not optimal, and derive a topology-aware isometric initialization scheme for vanilla-GCNs based on the principles of isometry. Additionally, contrary to ad-hoc addition of skip-connections, we propose to use gradient-guided dynamic rewiring of vanilla-GCNs with skip connections. Our dynamic rewiring method uses the gradient flow within each layer during training to introduce on-demand skip-connections adaptively. We provide extensive empirical evidence across multiple datasets that our methods improve gradient flow in deep vanilla-GCNs and significantly boost their performance to comfortably compete and outperform many fancy state-of-the-art methods. Codes are available at: https://github.com/VITA-Group/GradientGCN.
Ajay Jaiswal, Peihao Wang, Tianlong Chen 0001, Justin F. Rousseau, Ying Ding 0001, Zhangyang Wang
NeurIPS4
2022 Analytics to monitor local impact of the Protecting Access to Medicare Act's imaging clinical decision support requirements
abstract
OBJECTIVE: This study aimed is to: (1) extend the Integrating the Biology and the Bedside (i2b2) data and application models to include medical imaging appropriate use criteria, enabling it to serve as a platform to monitor local impact of the Protecting Access to Medicare Act's (PAMA) imaging clinical decision support (CDS) requirements, and (2) validate the i2b2 extension using data from the Medicare Imaging Demonstration (MID) CDS implementation. MATERIALS AND METHODS: This study provided a reference implementation and assessed its validity and reliability using data from the MID, the federal government's predecessor to PAMA's imaging CDS program. The Star Schema was extended to describe the interactions of imaging ordering providers with the CDS. New ontologies were added to enable mapping medical imaging appropriateness data to i2b2 schema. z-Ratio for testing the significance of the difference between 2 independent proportions was utilized. RESULTS: The reference implementation used 26 327 orders for imaging examinations which were persisted to the modified i2b2 schema. As an illustration of the analytical capabilities of the Web Client, we report that 331/1192 or 28.1% of imaging orders were deemed appropriate by the CDS system at the end of the intervention period (September 2013), an increase from 162/1223 or 13.2% for the first month of the baseline period, December 2011 (P = .0212), consistent with previous studies. CONCLUSIONS: The i2b2 platform can be extended to monitor local impact of PAMA's appropriateness of imaging ordering CDS requirements.
Vladimir I. Valtchinov, Shawn N. Murphy, Ronilda C. Lacson, Nikolay Ikonomov, Bingxue K. Zhai, Katherine P. Andriole, Justin F. Rousseau, Dick Hanson, Isaac S. Kohane, Ramin Khorasani
J. Am. Medical Informatics Assoc.7
2022 Methods for development and application of data standards in an ontology-driven information model for measuring, managing, and computing social determinants of health for individuals, households, and communities evaluated through an example of asthma
Justin F. Rousseau, Eliel Oliveira, William M. Tierney, Anjum Khurshid
J. Biomed. Informatics1
2022 Trustworthy assertion classification through prompting
Song Wang 0026, Liyan Tang, Akash Majety, Justin F. Rousseau, George Shih, Ying Ding 0001, Yifan Peng 0002
J. Biomed. Informatics4
2021 SCALP - Supervised Contrastive Learning for Cardiopulmonary Disease Classification and Localization in Chest X-rays using Patient Metadata
abstract
Computer-aided diagnosis plays a salient role in more accessible and accurate cardiopulmonary diseases classification and localization on chest radiography. Millions of people get affected and die due to these diseases without an accurate and timely diagnosis. Recently proposed contrastive learning heavily relies on data augmentation, especially positive data augmentation. However, generating clinically-accurate data augmentations for medical images is extremely difficult because the common data augmentation methods in computer vision, such as sharp, blur, and crop operations, can severely alter the clinical settings of medical images. In this paper, we proposed a novel and simple data augmentation method based on patient metadata and supervised knowledge to create clinically accurate positive and negative augmentations for chest X-rays. We introduce an end-to-end framework, SCALP, which extends the self-supervised contrastive approach to a supervised setting. Specifically, SCALP pulls together chest X-rays from the same patient (positive keys) and pushes apart chest X-rays from different patients (negative keys). In addition, it uses ResNet-50 along with the triplet-attention mechanism to identify cardiopulmonary diseases, and Grad-CAM++ to highlight the abnormal regions. Our extensive experiments demonstrate that SCALP outperforms existing baselines with significant margins in both classification and localization tasks. Specifically, the average classification AUCs improve from 82.8% (SOTA using DenseNet-121) to 83.9% (SCALP using ResNet-50), while the localization results improve on average by 3.7% over different IoU thresholds.
Ajay Jaiswal, Cyprian Zander, Yan Han 0001, Justin F. Rousseau, Yifan Peng 0002, Ying Ding 0001
ICDM5
2021 Letter to the editor in response to "Risk prediction of delirium in hospitalized patients using machine learning: an implementation and prospective evaluation study"
abstract
Dear JAMIA Editors, In their recent article, “Risk prediction of delirium in hospitalized patients using machine learning: an implementation and prospective evaluation study,” Jauk et al implemented and prospectively evaluated the performance of a machine learning algorithm to predict delirium in hospitalized patients using electronic health record (EHR) data available at admission and on the first evening after admission.1 This is an important problem in hospital medicine and neurology where delirium, a preventable condition, is under-recognized and under-treated and often leads to extended lengths of stay, increased health care costs, and acceleration in existing cognitive decline. Jauk et al demonstrated noteworthy accomplishments that will lead to exciting future investigations: (1) integrating the delirium predictive model into a clinical workflow within the EHR, (2) unobtrusively using data captured and documented in the normal care of patients, and (3) evaluating the performance of their model prospectively in the clinical setting. However, this study also illustrates a critical flaw in our approach to applying artificial intelligence and machine learning that prompts the question: what is the “ground truth” on which we are training our models? We should take pause before implementing machine learning algorithms in clinical contexts and assess the underlying classification task of the algorithms. The authors acknowledged a limitation of their study: basing the occurrence of delirium on the presence of International Classification of Disease-Tenth Revision (ICD-10 codes, terms and text © World Health Organization, Third Edition. 2007) codes F05 (“delirium due to known physiological condition” including all subcategories) and F10.4 (“alcohol withdrawal state with delirium”) assigned as diagnoses for the encounter. They recognize that a “lack of clear diagnostic criteria” for delirium “might be one reason why the incidence of delirium according to ICD codes in an administrative database (1.5% in this study) is lower than the one reported in prospective studies (ranging from 10%–40%).” Indeed, in the roadmap to advance delirium research from the Network for Investigation of Delirium: Unifying Scientists (NIDUS), Oh et al describe the need for a refined definition of and a reference standard for diagnosis of delirium.2 However, lack of a clear reference standard for delirium is not enough to explain such a deviation from prior measured incidence rates. In this case, it is apparent that the ground truth missed cases of delirium when it was present. Thus, efforts should have been made to evaluate and improve the ground truth prior to using it to train the predictive models because what algorithms are predicting might be the bias of determining the diagnosis, not the condition itself. Much like how the lack of a gold standard for diagnosis of cancer limits the utility of machine learning algorithms for diagnosing early stage cancer,3 the lack of clear diagnostic criteria to define delirium along with the dependence on the presence or absence of diagnosis codes limit the utility of the machine learning algorithm. If the algorithm performs prospectively as well as it does on the training set, it would only successfully identify cases that would have been coded with a diagnosis of delirium. Defining clinical conditions using available data, or defining “digital phenotypes,” is an art and a science in biomedical informatics. Definitions of clinical conditions have wide variability based on the data used, such as with congestive heart failure. Data beyond ICD codes are needed to improve the positive predictive value for conditions.4 Research support informatics teams have been developed at academic health centers to aid researchers in defining patient cohorts with various clinical conditions based on the best data available. Including different modalities (diagnosis codes, lab values, vital signs, reference to specific symptoms in the notes) of data in the digital phenotype definition process refines the accuracy of the cohort. We see evidence in this study of erroneously depending on diagnosis codes alone to define the digital phenotype of delirium in the comparison of expert nurses’ risk ratings for delirium compared to the calculated risk of the algorithm. There was a wide range of both predicted risk and nursing risk assessments of delirium and there was correlation between the predictive model and expert nursing assessments. However, in this study, 0/33 from the initial nursing assessment evaluation and 2/86 from the second nursing assessment had diagnoses of delirium (1 was correctly identified by the algorithm alone, and 1 was correctly identified by expert nursing alone). The authors expanded the cohort definition of delirium by searching free-text patient summaries for words related to delirium and, if positive, manually checking the cases for evidence of delirium. However, there was still a substantial difference in the incidence rate from this study and benchmark incidence rates. Additionally, the authors recognized that delirium is “not always coded in the participating hospital, and sometimes it is not even mentioned in the discharge summary.” The authors cite a lack of available data in the EHR limiting both the diagnosis of delirium as well as the performance of the prediction models. In particular, if a patient is new to the system, there is a lack of prior data. Even by including data collected during the first day, the lack of prior data posed challenges to the prediction model. This can be attributed to a dependence on structured data (demographic data, diagnosis data, laboratory data, nursing assessments, and procedures). Clinical notes, even in the emergency department setting, are a valuable source of clinically relevant data.5 Natural language processing technologies available today, and constantly improving, can extract phenotypic data from unstructured free-text notes. Such methods could make sufficient data accessible both to improve the accuracy of the digital phenotype of delirium as well as to improve the prediction model, even in those with no prior encounters. This work demonstrated that there are great opportunities to improve the defined digital phenotypes for delirium as well as other conditions to use as a more accurate ground truth in developing prediction algorithms. Natural language processing technologies can extend the search for useful data beyond those coded in the EHR to free-text reports and notes to improve both definition and prediction, but effort is needed to curate and optimize the ground truth we use to train future predictive models. This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. Both authors conceived the correspondence. JR drafted the correspondence and WT reviewed and edited the correspondence. Both authors had final approval of the correspondence and are accountable for all aspects of the work. None declared.
Justin F. Rousseau, William M. Tierney
J. Am. Medical Informatics Assoc.1
2020 Heterogeneous Graph Embeddings of Electronic Health Records Improve Critical Care Disease Predictions
Tingyi Wanyan, Martin Kang, Marcus A. Badgeley, Kipp W. Johnson, Jessica K. De Freitas, Fayzan F. Chaudhry, Akhil Vaid, Riccardo Miotto, Girish N. Nadkarni, Fei Wang 0001, Justin F. Rousseau, Ariful Azad, Ying Ding 0001, Benjamin S. Glicksberg
AIME12
2019 Managing Data Flows for Pediatric Complex-Care Patients and Their Families to Manage Care Plans
Steven B. Andrews, Mari-Ann Alexander, Justin F. Rousseau, Anjum Khurshid, Jaimie Hancock, Rahel Berhane
AMIA3
2019 Combining Forces for MRI Brain: Natural Language Processing of Radiology Reports with Structured Documentation
Justin F. Rousseau, Juan Diego Rodriguez, Jared Abrams, R. Nick Bryan
AMIA1
2018 A cross-institutional unified data request form to standardize the data request experience in Austin, TX
Justin F. Rousseau, Matt Sither, Rick Peters, Anjum Khurshid
AMIA1
2017 Can Emergency Department Provider Notes Help Achieve More Dynamic Clinical Decision Support?
Justin F. Rousseau, Ivan K. Ip, Ali S. Raja, Jeremiah D. Schuur, Ramin Khorasani
AMIA1
2017 Can Automated Retrieval of Data from Emergency Department Physician Notes Enhance Radiology Order Entry Process?
Justin F. Rousseau, Ivan K. Ip, Ali S. Raja, Vlad Valtchinov, Jeremiah D. Schuur, Ramin Khorasani
AMIA1