Rodney A. Gabriel

dblp:154/1553 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0003-4443-0021ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Engineering biomarker representations of vital signs data enhances deep learning mortality prediction
abstract
OBJECTIVES: We evaluated bidirectional long short-term memory models for predicting inpatient mortality using different approaches to processing vital signs data collected during the initial 24 h of intensive care unit (ICU) admissions. MATERIALS AND METHODS: We compared 3 vital-sign representations: (1) raw data recorded every 5 min, (2) preprocessed data averaged hourly, and (3) preprocessed data using biomarker representations that extends a digital oximetry biomarker toolbox of PhysioZoo software, applied to blood pressure, heart rate, temperature, respiratory rate, and SpO2. RESULTS: Across 2 large ICU datasets, HiRID and eICU, models trained on the frequency-normalized representation achieved higher discrimination and lower Brier scores than those trained on raw 5-min and hourly averaged data. DISCUSSION: The use of biomarker representations of vital signs yielded the largest improvements in discrimination and overall probabilistic performance reflected by lower Brier scores for predicting inpatient mortality by deep learning. CONCLUSION: Thus, we recommend using a similar approach to vital signs preprocessing for time-series predictive models.
Behrooz Mamandipoor, Isabella Shen, Chun-Nan Hsu, Rodney A. Gabriel
J. Am. Medical Informatics Assoc.4
2025 Improving postoperative length of stay forecasting with retrieval-augmented prediction
abstract
OBJECTIVE: The objective of this study is to evaluate retrieval-augmented prediction for forecasting hospital length of stay (LOS) following surgery compared to traditional machine learning (ML), standalone large language models (LLMs), and retrieval-augmented generation (RAG) approaches. MATERIALS AND METHODS: Spine surgery cases were extracted from electronic health records. Structured features and operative notes were concatenated into natural language patient representations, embedded using Sentence-Bidirectional Encoder Representations from Transformer, and stored in a vector database. Eight predictive models were implemented, including a baseline model, standalone ML with embeddings, standalone LLM (Gemma 3:27B), and combinations of these with retrieval-augmented prediction or generation. The retrieval-augmented prediction model computed a similarity-weighted average LOS from nearest neighbors. Performance was assessed using R2, mean absolute value (MAE), and root mean square error (RMSE). RESULTS: Retrieval-augmented prediction alone outperformed standalone ML and LLM models (R2 = 0.39, MAE = 4.47). Combining ML or LLM outputs with retrieval-augmented prediction further improved performance. The best performing model was a neural network blended with retrieval-augmented prediction (R2 = 0.52, MAE = 4.16). LLM-RAG alone reached R2 = 0.19, which improved to 0.47 when combined with retrieval-augmented predictions. Retrieval-augmented prediction consistently reduced MAE and RMSE by up to 32% and 38%, respectively. DISCUSSION: Retrieval-augmented prediction offers interpretable and resource-efficient forecasting by semantically leveraging prior patient cases without generative modeling. It consistently outperformed RAG and ML across metrics, approximating clinical reasoning via similarity-based inference. CONCLUSION: Retrieval-augmented prediction significantly enhances LOS prediction accuracy over standard ML and LLM models. Its interpretability and scalability make it a promising solution for integrating predictive analytics into clinical workflows.
Brian H. Park, Chun-Nan Hsu, Austin Nguyen, Ying Q. Zhou, Rodney A. Gabriel
J. Am. Medical Informatics Assoc.5
2024 Constructing synthetic datasets with generative artificial intelligence to train large language models to classify acute renal failure from clinical notes
abstract
OBJECTIVES: To compare performances of a classifier that leverages language models when trained on synthetic versus authentic clinical notes. MATERIALS AND METHODS: A classifier using language models was developed to identify acute renal failure. Four types of training data were compared: (1) notes from MIMIC-III; and (2, 3, and 4) synthetic notes generated by ChatGPT of varied text lengths of 15 (GPT-15 sentences), 30 (GPT-30 sentences), and 45 (GPT-45 sentences) sentences, respectively. The area under the receiver operating characteristics curve (AUC) was calculated from a test set from MIMIC-III. RESULTS: With RoBERTa, the AUCs were 0.84, 0.80, 0.84, and 0.76 for the MIMIC-III, GPT-15, GPT-30- and GPT-45 sentences training sets, respectively. DISCUSSION: Training language models to detect acute renal failure from clinical notes resulted in similar performances when using synthetic versus authentic training data. CONCLUSION: The use of training data derived from protected health information may not be needed.
Onkar Litake, Brian H. Park, Jeffrey L. Tully, Rodney A. Gabriel
J. Am. Medical Informatics Assoc.4
2021 A Machine Learning Approach to Predicting Time Needed for Outpatient Surgery and Patient Recovery at a Freestanding Ambulatory Surgery Center
Rodney A. Gabriel, Bhavya Harjai
AMIA1
2021 Calibrating predictive model estimates in a distributed network of patient data
Yingxiang Huang, Xiaoqian Jiang, Rodney A. Gabriel, Lucila Ohno-Machado
J. Biomed. Informatics3
2020 A tutorial on calibration measurements and calibration models for clinical prediction models
abstract
Our primary objective is to provide the clinical informatics community with an introductory tutorial on calibration measurements and calibration models for predictive models using existing R packages and custom implemented code in R on real and simulated data. Clinical predictive model performance is commonly published based on discrimination measures, but use of models for individualized predictions requires adequate model calibration. This tutorial is intended for clinical researchers who want to evaluate predictive models in terms of their applicability to a particular population. It is also for informaticians and for software engineers who want to understand the role that calibration plays in the evaluation of a clinical predictive model, and to provide them with a solid starting point to consider incorporating calibration evaluation and calibration models in their work. Covered topics include (1) an introduction to the importance of calibration in the clinical setting, (2) an illustration of the distinct roles that discrimination and calibration play in the assessment of clinical predictive models, (3) a tutorial and demonstration of selected calibration measurements, (4) a tutorial and demonstration of selected calibration models, and (5) a brief discussion of limitations of these methods and practical suggestions on how to use them in practice.
Yingxiang Huang, Fima Macheret, Rodney A. Gabriel, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.4
2020 EXpectation Propagation LOgistic REgRession on permissioned blockCHAIN (ExplorerChain): decentralized online healthcare/genomics predictive model learning
abstract
OBJECTIVE: Predicting patient outcomes using healthcare/genomics data is an increasingly popular/important area. However, some diseases are rare and require data from multiple institutions to construct generalizable models. To address institutional data protection policies, many distributed methods keep the data locally but rely on a central server for coordination, which introduces risks such as a single point of failure. We focus on providing an alternative based on a decentralized approach. We introduce the idea using blockchain technology for this purpose, with a brief description of its own potential advantages/disadvantages. MATERIALS AND METHODS: We explain how our proposed EXpectation Propagation LOgistic REgRession on Permissioned blockCHAIN (ExplorerChain) can achieve the same results when compared to a distributed model that uses a central server on 3 healthcare/genomic datasets, and what trade-offs need to be considered when using centralized/decentralized methods. We explain how the use of blockchain technology can help decrease some of the problems encountered in decentralized methods. RESULTS: We showed that the discrimination power of ExplorerChain can be statistically similar to its counterpart central server-based algorithm. While ExplorerChain inherited some benefits of blockchain, it had a small increased running time. DISCUSSION: ExplorerChain has the same prerequisites as a distributed model with a centralized server for coordination. In a manner similar to secure multi-party computation strategies, it assumes that participating institutions are honest, but "curious." CONCLUSION: When evaluated on relatively small datasets, results suggest that ExplorerChain, which combines artificial intelligence and blockchain technologies, performs as well as a central server-based method, and may avoid some risks at the cost of efficiency.
Tsung-Ting Kuo, Rodney A. Gabriel, Krishna R. Cidambi, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.2
2020 Privacy-preserving model learning on a blockchain network-of-networks
abstract
OBJECTIVE: To facilitate clinical/genomic/biomedical research, constructing generalizable predictive models using cross-institutional methods while protecting privacy is imperative. However, state-of-the-art methods assume a "flattened" topology, while real-world research networks may consist of "network-of-networks" which can imply practical issues including training on small data for rare diseases/conditions, prioritizing locally trained models, and maintaining models for each level of the hierarchy. In this study, we focus on developing a hierarchical approach to inherit the benefits of the privacy-preserving methods, retain the advantages of adopting blockchain, and address practical concerns on a research network-of-networks. MATERIALS AND METHODS: We propose a framework to combine level-wise model learning, blockchain-based model dissemination, and a novel hierarchical consensus algorithm for model ensemble. We developed an example implementation HierarchicalChain (hierarchical privacy-preserving modeling on blockchain), evaluated it on 3 healthcare/genomic datasets, as well as compared its predictive correctness, learning iteration, and execution time with a state-of-the-art method designed for flattened network topology. RESULTS: HierarchicalChain improves the predictive correctness for small training datasets and provides comparable correctness results with the competing method with higher learning iteration and similar per-iteration execution time, inherits the benefits of the privacy-preserving learning and advantages of blockchain technology, and immutable records models for each level. DISCUSSION: HierarchicalChain is independent of the core privacy-preserving learning method, as well as of the underlying blockchain platform. Further studies are warranted for various types of network topology, complex data, and privacy concerns. CONCLUSION: We demonstrated the potential of utilizing the information from the hierarchical network-of-networks topology to improve prediction.
Tsung-Ting Kuo, Jihoon Kim 0001, Rodney A. Gabriel
J. Am. Medical Informatics Assoc.3
2019 Fair compute loads enabled by blockchain: sharing models by alternating client and server roles
abstract
OBJECTIVE: Decentralized privacy-preserving predictive modeling enables multiple institutions to learn a more generalizable model on healthcare or genomic data by sharing the partially trained models instead of patient-level data, while avoiding risks such as single point of control. State-of-the-art blockchain-based methods remove the "server" role but can be less accurate than models that rely on a server. Therefore, we aim at developing a general model sharing framework to preserve predictive correctness, mitigate the risks of a centralized architecture, and compute the models in a fair way. MATERIALS AND METHODS: We propose a framework that includes both server and "client" roles to preserve correctness. We adopt a blockchain network to obtain the benefits of decentralization, by alternating the roles for each site to ensure computational fairness. Also, we developed GloreChain (Grid Binary LOgistic REgression on Permissioned BlockChain) as a concrete example, and compared it to a centralized algorithm on 3 healthcare or genomic datasets to evaluate predictive correctness, number of learning iterations and execution time. RESULTS: GloreChain performs exactly the same as the centralized method in terms of correctness and number of iterations. It inherits the advantages of blockchain, at the cost of increased time to reach a consensus model. DISCUSSION: Our framework is general or flexible and can also address intrinsic challenges of blockchain networks. Further investigations will focus on higher-dimensional datasets, additional use cases, privacy-preserving quality concerns, and ethical, legal, and social implications. CONCLUSIONS: Our framework provides a promising potential for institutions to learn a predictive model based on healthcare or genomic data in a privacy-preserving and decentralized way.
Tsung-Ting Kuo, Rodney A. Gabriel, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.2
2018 Identifying and characterizing highly similar notes in big clinical note datasets
Rodney A. Gabriel, Tsung-Ting Kuo, Julian J. McAuley, Chun-Nan Hsu
J. Biomed. Informatics1
2017 The presence of highly similar notes within the MIMIC-III dataset
Rodney A. Gabriel, Sanjeev Shenoy, Tsung-Ting Kuo, Julian J. McAuley, Chun-Nan Hsu
AMIA1
2016 Developing a predictive model for discharge delay in the Post-Anesthesia Care Unit
Rodney A. Gabriel, Jihoon Kim 0001, Lucila Ohno-Machado
AMIA1