VLDB 2026 Research / reviewers in the wild / expert
Jonathan H. Chen
dblp:35/512
· DBLP profile ↗
46ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0002-4387-8740ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 44 · 6 first-author · 24 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Quantization-aware matrix factorization for low bit rate image compressionabstractLossy image compression is essential for efficient transmission and storage. Traditional compression methods mainly rely on discrete cosine transform (DCT) or singular value decomposition (SVD), both of which represent image data in continuous domains and, therefore, necessitate carefully designed quantizers. Notably, these methods consider quantization as a separate step, which prevents quantization errors from being incorporated into the compression process and degrades the reconstruction quality, particularly in SVD-based methods. To address this issue, we introduce a quantization-aware matrix factorization (QMF) to develop a novel lossy image compression method. QMF provides a low-rank representation of the image data as a product of two smaller matrices, with elements constrained to bounded integer values, thereby effectively integrating quantization with low-rank approximation. We propose an efficient, provably convergent iterative algorithm for QMF using a block coordinate descent scheme, with subproblems having closed-form solutions. Our experiments demonstrate that our method consistently outperforms JPEG at low bit rates below 0.25 bits per pixel. We also demonstrated that our method has an improved capability to preserve visual semantics compared to JPEG at low bit rates by evaluating an ImageNet pre-trained classifier on compressed images. The project is available at https://github.com/pashtari/qmf . Pooya Ashtari, Pourya Behmandpoor, Fateme Nateghi Haredasht, Jonathan H. Chen, Panagiotis Patrinos, Sabine Van Huffel |
Inf. Sci. | 4 |
| 2026 | Extending the Fundamental Theorem of Biomedical Informatics for the AI eraabstractBACKGROUND: Charles Friedman's Fundamental Theorem of Biomedical Informatics holds that a person working in partnership with an information resource outperforms that same person unassisted. Since its publication, advances in artificial intelligence (AI), adaptive learning systems, and large-scale data infrastructures have transformed the biomedical ecosystem, extending informatics beyond clinical care into domains such as public health, consumer health, translational science, and the broader life sciences. Such expansion has further underscored the importance of the Fundamental Theorem while also elucidating ways it can be expanded to meet current needs. OBJECTIVE: To reassess and extend the Fundamental Theorem for the AI era in a manner that preserves its conceptual strength while broadening its applicability across an evolved and more complex biomedical ecosystem. METHODS: This Viewpoint synthesizes empirical evidence and sociotechnical theory related to human-AI collaboration, learning health systems (LHS), learning public health systems (LPHS), AI governance, and systems science to contextualize the Fundamental Theorem within such contemporary frameworks. RESULTS: We argue that the unit of analysis of the Fundamental Theorem should shift from individuals and tools to adaptive sociotechnical systems spanning clinical care, public health, translational research, consumer engagement, and life sciences innovation. We propose an expanded theorem: A learning biomedical ecosystem that continuously optimizes human-AI collaboration will outperform humans or AI alone. CONCLUSIONS: This evolution builds directly upon Friedman's original theorem, reaffirming its human-centered foundation, while incorporating AI-enabled computation, adaptive learning, and systems-level integration across the modern biomedical enterprise. Philip R. O. Payne, Jonathan H. Chen, Christopher A. Longhurst |
J. Am. Medical Informatics Assoc. | 2 |
| 2025 | Establishing best practices in large language model research: an application to repeat promptingabstractOBJECTIVES: We aimed to demonstrate the importance of establishing best practices in large language model research, using repeat prompting as an illustrative example. MATERIALS AND METHODS: Using data from a prior study investigating potential model bias in peer review of medical abstracts, we compared methods that ignore correlation in model outputs from repeated prompting with a random effects method that accounts for this correlation. RESULTS: High correlation within groups was found when repeatedly prompting the model, with intraclass correlation coefficient of 0.69. Ignoring the inherent correlation in the data led to over 100-fold inflation of effective sample size. After appropriately accounting for this issue, the authors' results reverse from a small but highly significant finding to no evidence of model bias. DISCUSSION: The establishment of best practices for LLM research is urgently needed, as demonstrated in this case where accounting for repeat prompting in analyses was critical for accurate study conclusions. Robert Gallo, Michael T. M. Baiocchi, Thomas Savage, Jonathan H. Chen |
J. Am. Medical Informatics Assoc. | 4 |
| 2025 | Predicting treatment retention in medication for opioid use disorder: a machine learning approach using NLP and LLM-derived clinical featuresabstractOBJECTIVE: Building upon our previous work on predicting treatment retention in medications for opioid use disorder, we aimed to improve 6-month retention prediction in buprenorphine-naloxone (BUP-NAL) therapy by incorporating features derived from large language models (LLMs) applied to unstructured clinical notes. MATERIALS AND METHODS: We used de-identified electronic health record (EHR) data from Stanford Health Care (STARR) for model development and internal validation, and the NeuroBlu behavioral health database for external validation. Structured features were supplemented with 13 clinical and psychosocial features extracted from free-text notes using the CLinical Entity Augmented Retrieval pipeline, which combines named entity recognition with LLM-based classification to provide contextual interpretation. We trained classification (Logistic Regression, Random Forest, XGBoost) and survival models (CoxPH, Random Survival Forest, Survival XGBoost), evaluated using Receiver Operating Characteristic-Area Under the Curve (ROC-AUC) and C-index. RESULTS: XGBoost achieved the highest classification performance (ROC-AUC = 0.65). Incorporating LLM-derived features improved model performance across all architectures, with the largest gains observed in simpler models such as Logistic Regression. In time-to-event analysis, Random Survival Forest and Survival XGBoost reached the highest C-index (≈0.65). SHapley Additive exPlanations analysis identified LLM-extracted features like Chronic Pain, Liver Disease, and Major Depression as key predictors. We also developed an interactive web tool for real-time clinical use. DISCUSSION: Features extracted using NLP and LLM-assisted methods improved model accuracy and interpretability, revealing valuable psychosocial risks not captured in structured EHRs. CONCLUSION: Combining structured EHR data with LLM-extracted features moderately improves BUP-NAL retention prediction, enabling personalized risk stratification and advancing AI-driven care for substance use disorders. Fateme Nateghi Haredasht, Iván López 0001, Steven Tate, Pooya Ashtari, Min Min Chan, Deepali Kulkarni, Chwen-Yuen Angie Chen, Maithri Vangala, Kira Griffith, Bryan Bunning, Adam S. Miner, Tina Hernandez-Boussard, Keith Humphreys, Anna Lembke, L. Alexander Vance, Jonathan H. Chen |
J. Am. Medical Informatics Assoc. | 16 |
| 2025 | Large language model uncertainty proxies: discrimination and calibration for medical diagnosis and treatmentabstractINTRODUCTION: The inability of large language models (LLMs) to communicate uncertainty is a significant barrier to their use in medicine. Before LLMs can be integrated into patient care, the field must assess methods to estimate uncertainty in ways that are useful to physician-users. OBJECTIVE: Evaluate the ability for uncertainty proxies to quantify LLM confidence when performing diagnosis and treatment selection tasks by assessing the properties of discrimination and calibration. METHODS: We examined confidence elicitation (CE), token-level probability (TLP), and sample consistency (SC) proxies across GPT3.5, GPT4, Llama2, and Llama3. Uncertainty proxies were evaluated against 3 datasets of open-ended patient scenarios. RESULTS: SC discrimination outperformed TLP and CE methods. SC by sentence embedding achieved the highest discriminative performance (ROC AUC 0.68-0.79), yet with poor calibration. SC by GPT annotation achieved the second-best discrimination (ROC AUC 0.66-0.74) with accurate calibration. Verbalized confidence (CE) was found to consistently overestimate model confidence. DISCUSSION AND CONCLUSIONS: SC is the most effective method for estimating LLM uncertainty of the proxies evaluated. SC by sentence embedding can effectively estimate uncertainty if the user has a set of reference cases with which to re-calibrate their results, while SC by GPT annotation is the more effective method if the user does not have reference cases and requires accurate raw calibration. Our results confirm LLMs are consistently over-confident when verbalizing their confidence (CE). Thomas Savage, Robert Gallo, Abdessalem Boukil, Vishwesh Patel, Seyed Amir Ahmad Safavi-Naini, Ali Soroush, Jonathan H. Chen |
J. Am. Medical Informatics Assoc. | 8 |
| 2025 | Monitoring strategies for continuous evaluation of deployed clinical prediction models
Grace Y. E. Kim, Conor K. Corbin, François Grolleau, Michael T. M. Baiocchi, Jonathan H. Chen |
J. Biomed. Informatics | 5 |
| 2024 | MedAlign: A Clinician-Generated Dataset for Instruction Following with Electronic Medical RecordsabstractThe ability of large language models (LLMs) to follow natural language instructions with human-level fluency suggests many opportunities in healthcare to reduce administrative burden and improve quality of care. However, evaluating LLMs on realistic text generation tasks for healthcare remains challenging. Existing question answering datasets for electronic health record (EHR) data fail to capture the complexity of information needs and documentation burdens experienced by clinicians. To address these challenges, we introduce MedAlign, a benchmark dataset of 983 natural language instructions for EHR data. MedAlign is curated by 15 clinicians (7 specialities), includes clinician-written reference responses for 303 instructions, and provides 276 longitudinal EHRs for grounding instruction-response pairs. We used MedAlign to evaluate 6 general domain LLMs, having clinicians rank the accuracy and quality of each LLM response. We found high error rates, ranging from 35% (GPT-4) to 68% (MPT-7B-Instruct), and 8.3% drop in accuracy moving from 32k to 2k context lengths for GPT-4. Finally, we report correlations between clinician rankings and automated natural language generation metrics as a way to rank LLMs without human review. MedAlign is provided under a research data use agreement to enable LLM evaluations on tasks aligned with clinician needs and preferences. Scott L. Fleming, Alejandro Lozano, William J. Haberkorn, Jenelle A. Jindal, Eduardo Pontes Reis, Rahul Thapa, Louis Blankemeier, Julian Z. Genkins, Ethan Steinberg, Ashwin Nayak 0002, Birju Patel, Chia-Chun Chiang, Alison Callahan, Zepeng Huo, Sergios Gatidis, Scott J. Adams, Oluseyi Fayanju, Shreya J. Shah, Thomas Savage, Ethan Goh, Akshay Chaudhari, Nima Aghaeepour, Christopher D. Sharp, Michael A. Pfeffer, Percy Liang, Jonathan H. Chen, Keith E. Morse, Emma Brunskill, Jason Alan Fries, Nigam H. Shah |
AAAI | 26 |
| 2023 | DEPLOYR: a technical framework for deploying custom real-time machine learning models into the electronic medical recordabstractOBJECTIVE: Heatlhcare institutions are establishing frameworks to govern and promote the implementation of accurate, actionable, and reliable machine learning models that integrate with clinical workflow. Such governance frameworks require an accompanying technical framework to deploy models in a resource efficient, safe and high-quality manner. Here we present DEPLOYR, a technical framework for enabling real-time deployment and monitoring of researcher-created models into a widely used electronic medical record system. MATERIALS AND METHODS: We discuss core functionality and design decisions, including mechanisms to trigger inference based on actions within electronic medical record software, modules that collect real-time data to make inferences, mechanisms that close-the-loop by displaying inferences back to end-users within their workflow, monitoring modules that track performance of deployed models over time, silent deployment capabilities, and mechanisms to prospectively evaluate a deployed model's impact. RESULTS: We demonstrate the use of DEPLOYR by silently deploying and prospectively evaluating 12 machine learning models trained using electronic medical record data that predict laboratory diagnostic results, triggered by clinician button-clicks in Stanford Health Care's electronic medical record. DISCUSSION: Our study highlights the need and feasibility for such silent deployment, because prospectively measured performance varies from retrospective estimates. When possible, we recommend using prospectively estimated performance measures during silent trials to make final go decisions for model deployment. CONCLUSION: Machine learning applications in healthcare are extensively researched, but successful translations to the bedside are rare. By describing DEPLOYR, we aim to inform machine learning deployment best practices and help bridge the model implementation gap. Conor K. Corbin, Rob Maclay, Aakash Acharya, Sreedevi Mony, Soumya Punnathanam, Rahul Thapa, Nikesh Kotecha, Nigam H. Shah, Jonathan H. Chen |
J. Am. Medical Informatics Assoc. | 9 |
| 2023 | Selective prediction for extracting unstructured clinical dataabstractOBJECTIVE: While there are currently approaches to handle unstructured clinical data, such as manual abstraction and structured proxy variables, these methods may be time-consuming, not scalable, and imprecise. This article aims to determine whether selective prediction, which gives a model the option to abstain from generating a prediction, can improve the accuracy and efficiency of unstructured clinical data abstraction. MATERIALS AND METHODS: We trained selective classifiers (logistic regression, random forest, support vector machine) to extract 5 variables from clinical notes: depression (n = 1563), glioblastoma (GBM, n = 659), rectal adenocarcinoma (DRA, n = 601), and abdominoperineal resection (APR, n = 601) and low anterior resection (LAR, n = 601) of adenocarcinoma. We varied the cost of false positives (FP), false negatives (FN), and abstained notes and measured total misclassification cost. RESULTS: The depression selective classifiers abstained on anywhere from 0% to 97% of notes, and the change in total misclassification cost ranged from -58% to 9%. Selective classifiers abstained on 5%-43% of notes across the GBM and colorectal cancer models. The GBM selective classifier abstained on 43% of notes, which led to improvements in sensitivity (0.94 to 0.96), specificity (0.79 to 0.96), PPV (0.89 to 0.98), and NPV (0.88 to 0.91) when compared to a non-selective classifier and when compared to structured proxy variables. DISCUSSION: We showed that selective classifiers outperformed both non-selective classifiers and structured proxy variables for extracting data from unstructured clinical notes. CONCLUSION: Selective prediction should be considered when abstaining is preferable to making an incorrect prediction. Akshay Swaminathan, Iván López 0001, Ujwal Srivastava, Edward Tran, Aarohi Bhargava-Shah, Janet Y. Wu, Alexander L. Ren, Kaitlin Caoili, Brandon Bui, Layth Alkhani, Nathan Mohit, Noel Seo, Nicholas Macedo, Winson Cheng, Charles Liu, Reena Thomas, Jonathan H. Chen, Olivier Gevaert |
J. Am. Medical Informatics Assoc. | 19 |
| 2023 | Clinical outcome prediction using observational supervision with electronic health records and audit logs
Nandita Bhaskhar, Wui Ip, Jonathan H. Chen, Daniel L. Rubin |
J. Biomed. Informatics | 3 |
| 2023 | Graph-based clinical recommender: Predicting specialists procedure orders using graph representation learning
Sajjad Fouladvand, Federico Reyes Gomez, Hamed Nilforoshan, Matthew Schwede, Morteza Noshad, Olivia Jee, Jiaxuan You, Rok Sosic, Jure Leskovec, Jonathan H. Chen |
J. Biomed. Informatics | 10 |
| 2022 | How to Avoid Incorrect Clinical Machine Learning Model Performance Estimates When Class Labels Are Only Partially Observed
Conor K. Corbin, Michael T. M. Baiocchi, Jonathan H. Chen |
AMIA | 3 |
| 2022 | Modifiable Inpatient Cost Variability at a Large Academic Medical Center
Kush Gupta, Jason Hom, Ted Ross, Jonathan Masterson, Alexander Chin, Arnold Milstein, Jonathan H. Chen |
AMIA | 7 |
| 2022 | Gaps in Nephrology Referral Care Utilization in Patients at High-Risk of Progression to Kidney Failure
Samson Peter, Maggie Wang 0002, Chi D. Chu, Delphine S. Tuot, Arhana Chattopadhyay, Jonathan H. Chen |
AMIA | 6 |
| 2022 | Leveraging EHR Audit Log Data to Unlock New Insights into Care Processes and Outcomes
Christian Rose, Robert Thombley, Morteza Noshad, Ron Li, Wendy Lu, Heather A Clancy, David Schlessinger, Vincent X. Liu, Jonathan H. Chen, Julia Adler-Milstein |
AMIA | 9 |
| 2022 | Interruptive Electronic Alerts for Choosing Wisely Recommendations: A Cluster Randomized Controlled TrialabstractOBJECTIVE: To assess the efficacy of interruptive electronic alerts in improving adherence to the American Board of Internal Medicine's Choosing Wisely recommendations to reduce unnecessary laboratory testing. MATERIALS AND METHODS: We administered 5 cluster randomized controlled trials simultaneously, using electronic medical record alerts regarding prostate-specific antigen (PSA) testing, acute sinusitis treatment, vitamin D testing, carotid artery ultrasound screening, and human papillomavirus testing. For each alert, we assigned 5 outpatient clinics to an interruptive alert and 5 were observed as a control. Primary and secondary outcomes were the number of postalert orders per 100 patients at each clinic and number of triggered alerts divided by orders, respectively. Post hoc analysis evaluated whether physicians experiencing interruptive alerts reduced their alert-triggering behaviors. RESULTS: Median postalert orders per 100 patients did not differ significantly between treatment and control groups; absolute median differences ranging from 0.04 to 0.40 for PSA testing. Median alerts per 100 orders did not differ significantly between treatment and control groups; absolute median differences ranged from 0.004 to 0.03. In post hoc analysis, providers receiving alerts regarding PSA testing in men were significantly less likely to trigger additional PSA alerts than those in the control sites (Incidence Rate Ratio 0.12, 95% CI [0.03-0.52]). DISCUSSION: Interruptive point-of-care alerts did not yield detectable changes in the overall rate of undesired orders or the order-to-alert ratio between active and silent sites. Complementary behavioral or educational interventions are likely needed to improve efforts to curb medical overuse. CONCLUSION: Implementation of interruptive alerts at the time of ordering was not associated with improved adherence to 5 Choosing Wisely guidelines. TRIAL REGISTRATION: NCT02709772. Vy T. Ho, Rachael C. Aikens, Geoffrey J. Tso, Paul Heidenreich, Christopher D. Sharp, Steven M. Asch, Jonathan H. Chen, Neil K. Shah |
J. Am. Medical Informatics Assoc. | 7 |
| 2022 | Team is brain: leveraging EHR audit log data for new insights into acute care processesabstractOBJECTIVE: To determine whether novel measures of contextual factors from multi-site electronic health record (EHR) audit log data can explain variation in clinical process outcomes. MATERIALS AND METHODS: We selected one widely-used process outcome: emergency department (ED)-based team time to deliver tissue plasminogen activator (tPA) to patients with acute ischemic stroke (AIS). We evaluated Epic audit log data (that tracks EHR user-interactions) for 3052 AIS patients aged 18+ who received tPA after presenting to an ED at three Northern California health systems (Stanford Health Care, UCSF Health, and Kaiser Permanente Northern California). Our primary outcome was door-to-needle time (DNT) and we assessed bivariate and multivariate relationships with six audit log-derived measures of treatment team busyness and prior team experience. RESULTS: Prior team experience was consistently associated with shorter DNT; teams with greater prior experience specifically on AIS cases had shorter DNT (minutes) across all sites: (Site 1: -94.73, 95% CI: -129.53 to 59.92; Site 2: -80.93, 95% CI: -130.43 to 31.43; Site 3: -42.95, 95% CI: -62.73 to 23.17). Teams with greater prior experience across all types of cases also had shorter DNT at two sites: (Site 1: -6.96, 95% CI: -14.56 to 0.65; Site 2: -19.16, 95% CI: -36.15 to 2.16; Site 3: -11.07, 95% CI: -17.39 to 4.74). Team busyness was not consistently associated with DNT across study sites. CONCLUSIONS: EHR audit log data offers a novel, scalable approach to measure key contextual factors relevant to clinical process outcomes across multiple sites. Audit log-based measures of team experience were associated with better process outcomes for AIS care, suggesting opportunities to study underlying mechanisms and improve care through deliberate training, team-building, and scheduling to maximize team experience. Christian Rose, Robert Thombley, Morteza Noshad, Heather A Clancy, David Schlessinger, Ron C. Li, Vincent X. Liu, Jonathan H. Chen, Julia Adler-Milstein |
J. Am. Medical Informatics Assoc. | 9 |
| 2022 | Signal from the noise: A mixed graphical and quantitative process mining approach to evaluate care pathways applied to emergency stroke care
Morteza Noshad, Christian Rose, Jonathan H. Chen |
J. Biomed. Informatics | 3 |
| 2021 | A Retrospective Analysis of Machine Learning Driven Antibiotic Selection in the Emergency Department
Conor K. Corbin, Arhana Chattopadhyah, Lillian Sung, Amy Chang, Stan Deresinski, Jonathan H. Chen |
AMIA | 6 |
| 2021 | A Data-Driven Algorithm to Recommend Initial Clinical Workup for Outpatient Specialty Referral
Wui Ip, Priya Prahalad, Jonathan Palma, Jonathan H. Chen |
AMIA | 4 |
| 2021 | Machine Learning Predictability of Clinical Next Generation Sequencing for Hematologic Malignancies to Guide High-Value Precision Medicine
Grace Y. E. Kim, Morteza Noshad, Henning Stehr, Rebecca Rojansky, Dita Gratzinger, Jean Oak, Rondeep Brar, David Iberri, Christina Kong, James L. Zehnder, Jonathan H. Chen |
AMIA | 11 |
| 2021 | Randomized user testing of recommender system clinical decision support
Andre Kumar, Rachael C. Aikens, Jason Horn, Lisa Shieh, Mark A. Musen, Michael T. M. Baiocchi, Russ B. Altman, Mary K. Goldstein, Steven M. Asch, Jonathan H. Chen |
AMIA | 10 |
| 2021 | Predicting Level of Care for Emergency Hospital Admissions to Optimize Triage
Nicolai P. Ostberg, Conor K. Corbin, Tiffany Eulalio, Gautam Machiraju, Ben J. Marafino, Michael T. M. Baiocchi, Christian Rose, Jonathan H. Chen |
AMIA | 9 |
| 2021 | Signal from the Noise: Quantitative Measures of Conformity and Variability From Process Mining Maps
Christian Rose, Morteza Noshad, Jonathan H. Chen |
AMIA | 3 |
| 2021 | Developing machine learning models to personalize care levels among emergency room patients for hospital admissionabstractOBJECTIVE: To develop prediction models for intensive care unit (ICU) vs non-ICU level-of-care need within 24 hours of inpatient admission for emergency department (ED) patients using electronic health record data. MATERIALS AND METHODS: Using records of 41 654 ED visits to a tertiary academic center from 2015 to 2019, we tested 4 algorithms-feed-forward neural networks, regularized regression, random forests, and gradient-boosted trees-to predict ICU vs non-ICU level-of-care within 24 hours and at the 24th hour following admission. Simple-feature models included patient demographics, Emergency Severity Index (ESI), and vital sign summary. Complex-feature models added all vital signs, lab results, and counts of diagnosis, imaging, procedures, medications, and lab orders. RESULTS: The best-performing model, a gradient-boosted tree using a full feature set, achieved an AUROC of 0.88 (95%CI: 0.87-0.89) and AUPRC of 0.65 (95%CI: 0.63-0.68) for predicting ICU care need within 24 hours of admission. The logistic regression model using ESI achieved an AUROC of 0.67 (95%CI: 0.65-0.70) and AUPRC of 0.37 (95%CI: 0.35-0.40). Using a discrimination threshold, such as 0.6, the positive predictive value, negative predictive value, sensitivity, and specificity were 85%, 89%, 30%, and 99%, respectively. Vital signs were the most important predictors. DISCUSSION AND CONCLUSIONS: Undertriaging admitted ED patients who subsequently require ICU care is common and associated with poorer outcomes. Machine learning models using readily available electronic health record data predict subsequent need for ICU admission with good discrimination, substantially better than the benchmarking ESI system. The results could be used in a multitiered clinical decision-support system to improve ED triage. Conor K. Corbin, Tiffany Eulalio, Nicolai P. Ostberg, Gautam Machiraju, Ben J. Marafino, Michael T. M. Baiocchi, Christian Rose, Jonathan H. Chen |
J. Am. Medical Informatics Assoc. | 9 |
| 2021 | Machine learning for initial insulin estimation in hospitalized patientsabstractOBJECTIVE: The study sought to determine whether machine learning can predict initial inpatient total daily dose (TDD) of insulin from electronic health records more accurately than existing guideline-based dosing recommendations. MATERIALS AND METHODS: Using electronic health records from a tertiary academic center between 2008 and 2020 of 16,848 inpatients receiving subcutaneous insulin who achieved target blood glucose control of 100-180 mg/dL on a calendar day, we trained an ensemble machine learning algorithm consisting of regularized regression, random forest, and gradient boosted tree models for 2-stage TDD prediction. We evaluated the ability to predict patients requiring more than 6 units TDD and their point-value TDDs to achieve target glucose control. RESULTS: The method achieves an area under the receiver-operating characteristic curve of 0.85 (95% confidence interval [CI], 0.84-0.87) and area under the precision-recall curve of 0.65 (95% CI, 0.64-0.67) for classifying patients who require more than 6 units TDD. For patients requiring more than 6 units TDD, the mean absolute percent error in dose prediction based on standard clinical calculators using patient weight is in the range of 136%-329%, while the regression model based on weight improves to 60% (95% CI, 57%-63%), and the full ensemble model further improves to 51% (95% CI, 48%-54%). DISCUSSION: Owingto the narrow therapeutic window and wide individual variability, insulin dosing requires adaptive and predictive approaches that can be supported through data-driven analytic tools. CONCLUSIONS: Machine learning approaches based on readily available electronic medical records can discriminate which inpatients will require more than 6 units TDD and estimate individual doses more accurately than standard guidelines and practices. Ivana Jankovic, Laurynas Kalesinskas, Michael T. M. Baiocchi, Jonathan H. Chen |
J. Am. Medical Informatics Assoc. | 5 |
| 2020 | It's not just a last mile problem: Partnering with process improvement and human factors to integrate machine learning into healthcare delivery
Ron C. Li, Naveen Muthu, Margaret Smith, Swaminathan Kandaswamy, Jonathan H. Chen |
AMIA | 5 |
| 2020 | Context is Key: Using the Audit Log to Capture Contextual Factors Affecting Stroke Care Processes
Morteza Noshad, Christian Rose, Robert Thombley, Jonathan Chiang, Conor K. Corbin, Vincent X. Liu, Julia Adler-Milstein, Jonathan H. Chen |
AMIA | 9 |
| 2020 | Data-Driven Clinical Decision Support for Computerized Physician Order Entry: Development, Evaluation, and Implementation
Yiye Zhang, Jonathan H. Chen, Adam Wright, Jessica S. Ancker, Marc Tobias |
AMIA | 2 |
| 2020 | OrderRex clinical user testing: a randomized trial of recommender system decision support on simulated casesabstractOBJECTIVE: To assess usability and usefulness of a machine learning-based order recommender system applied to simulated clinical cases. MATERIALS AND METHODS: 43 physicians entered orders for 5 simulated clinical cases using a clinical order entry interface with or without access to a previously developed automated order recommender system. Cases were randomly allocated to the recommender system in a 3:2 ratio. A panel of clinicians scored whether the orders placed were clinically appropriate. Our primary outcome included the difference in clinical appropriateness scores. Secondary outcomes included total number of orders, case time, and survey responses. RESULTS: Clinical appropriateness scores per order were comparable for cases randomized to the order recommender system (mean difference -0.11 order per score, 95% CI: [-0.41, 0.20]). Physicians using the recommender placed more orders (median 16 vs 15 orders, incidence rate ratio 1.09, 95%CI: [1.01-1.17]). Case times were comparable with the recommender system. Order suggestions generated from the recommender system were more likely to match physician needs than standard manual search options. Physicians used recommender suggestions in 98% of available cases. Approximately 95% of participants agreed the system would be useful for their workflows. DISCUSSION: User testing with a simulated electronic medical record interface can assess the value of machine learning and clinical decision support tools for clinician usability and acceptance before live deployments. CONCLUSIONS: Clinicians can use and accept machine learned clinical order recommendations integrated into an electronic order entry interface in a simulated setting. The clinical appropriateness of orders entered was comparable even when supported by automated recommendations. Andre Kumar, Rachael C. Aikens, Jason Hom, Lisa Shieh, Jonathan Chiang, David Morales, Divya Saini, Mark A. Musen, Michael T. M. Baiocchi, Russ B. Altman, Mary K. Goldstein, Steven M. Asch, Jonathan H. Chen |
J. Am. Medical Informatics Assoc. | 13 |
| 2020 | Explainable artificial intelligence models using real-world electronic health record data: a systematic scoping reviewabstractOBJECTIVE: To conduct a systematic scoping review of explainable artificial intelligence (XAI) models that use real-world electronic health record data, categorize these techniques according to different biomedical applications, identify gaps of current studies, and suggest future research directions. MATERIALS AND METHODS: We searched MEDLINE, IEEE Xplore, and the Association for Computing Machinery (ACM) Digital Library to identify relevant papers published between January 1, 2009 and May 1, 2019. We summarized these studies based on the year of publication, prediction tasks, machine learning algorithm, dataset(s) used to build the models, the scope, category, and evaluation of the XAI methods. We further assessed the reproducibility of the studies in terms of the availability of data and code and discussed open issues and challenges. RESULTS: Forty-two articles were included in this review. We reported the research trend and most-studied diseases. We grouped XAI methods into 5 categories: knowledge distillation and rule extraction (N = 13), intrinsically interpretable models (N = 9), data dimensionality reduction (N = 8), attention mechanism (N = 7), and feature interaction and importance (N = 5). DISCUSSION: XAI evaluation is an open issue that requires a deeper focus in the case of medical applications. We also discuss the importance of reproducibility of research work in this field, as well as the challenges and opportunities of XAI from 2 medical professionals' point of view. CONCLUSION: Based on our review, we found that XAI evaluation in medicine has not been adequately and formally practiced. Reproducibility remains a critical concern. Ample opportunities exist to advance XAI research in medicine. Seyedeh Neelufar Payrovnaziri, Zhaoyi Chen, Pablo Rengifo-Moreno, Tim Miller 0001, Jiang Bian 0001, Jonathan H. Chen, Xiuwen Liu 0001, Zhe He 0001 |
J. Am. Medical Informatics Assoc. | 6 |
| 2019 | Detecting unanticipated actions downstream from clinical decision support: a data mining approach
Ron C. Li, Imon Banerjee, Daniel L. Rubin, Jonathan H. Chen |
AMIA | 4 |
| 2019 | Prevalence and Predictability of Low Yield Inpatient Laboratory Diagnostic Tests
Jason Horn, Santhosh Balasubramanian, Lee F. Schroeder, Nader Najafi, Shivaal Roy, Jonathan H. Chen |
AMIA | 7 |
| 2019 | Assessing clinical heterogeneity in sepsis through treatment patterns and machine learningabstractOBJECTIVE: To use unsupervised topic modeling to evaluate heterogeneity in sepsis treatment patterns contained within granular data of electronic health records. MATERIALS AND METHODS: A multicenter, retrospective cohort study of 29 253 hospitalized adult sepsis patients between 2010 and 2013 in Northern California. We applied an unsupervised machine learning method, Latent Dirichlet Allocation, to the orders, medications, and procedures recorded in the electronic health record within the first 24 hours of each patient's hospitalization to uncover empiric treatment topics across the cohort and to develop computable clinical signatures for each patient based on proportions of these topics. We evaluated how these topics correlated with common sepsis treatment and outcome metrics including inpatient mortality, time to first antibiotic, and fluids given within 24 hours. RESULTS: Mean age was 70 ± 17 years with hospital mortality of 9.6%. We empirically identified 42 clinically recognizable treatment topics (eg, pneumonia, cellulitis, wound care, shock). Only 43.1% of hospitalizations had a single dominant topic, and a small minority (7.3%) had a single topic comprising at least 80% of their overall clinical signature. Across the entire sepsis cohort, clinical signatures were highly variable. DISCUSSION: Heterogeneity in sepsis is a major barrier to improving targeted treatments, yet existing approaches to characterizing clinical heterogeneity are narrowly defined. A machine learning approach captured substantial patient- and population-level heterogeneity in treatment during early sepsis hospitalization. CONCLUSION: Using topic modeling based on treatment patterns may enable more precise clinical characterization in sepsis and better understanding of variability in sepsis presentation and outcomes. Alison E. Fohner, John D. Greene, Brian L. Lawson, Jonathan H. Chen, Patricia Kipnis, Gabriel J. Escobar, Vincent X. Liu |
J. Am. Medical Informatics Assoc. | 4 |
| 2018 | Machine Learning Predictable Inpatient Lab Results for High-Value Care
Santhosh Balasubramanian, Jason Hom, Shivaal Roy, Jonathan H. Chen |
AMIA | 4 |
| 2018 | A "bottom up" data driven approach to curating electronic order sets
Ron C. Li, Jason K. Wang, Christopher D. Sharp, Jonathan H. Chen |
AMIA | 4 |
| 2018 | Impact of problem-based charting on the utilization and accuracy of the electronic problem listabstractObjective: Problem-based charting (PBC) is a method for clinician documentation in commercially available electronic medical record systems that integrates note writing and problem list management. We report the effect of PBC on problem list utilization and accuracy at an academic intensive care unit (ICU). Materials and Methods: An interrupted time series design was used to assess the effect of PBC on problem list utilization, which is defined as the number of new problems added to the problem list by clinicians per patient encounter, and of problem list accuracy, which was determined by calculating the recall and precision of the problem list in capturing 5 common ICU diagnoses. Results: In total, 3650 and 4344 patient records were identified before and after PBC implementation at Stanford Hospital. An increase of 2.18 problems (>50% increase) in the mean number of new problems added to the problem list per patient encounter can be attributed to the initiation of PBC. There was a significant increase in recall attributed to the initiation of PBC for sepsis (β = 0.45, P < .001) and acute renal failure (β = 0.2, P = .007), but not for acute respiratory failure, pneumonia, or venous thromboembolism. Discussion: The problem list is an underutilized component of the electronic medical record that can be a source of clinician-structured data representing the patient's clinical condition in real time. PBC is a readily available tool that can integrate problem list management into physician workflow. Conclusion: PBC improved problem list utilization and accuracy at an academic ICU. Ron C. Li, Trit Garg, Tony Cun, Lisa Shieh, Gomathi Krishnan, Daniel Z. Fang, Jonathan H. Chen |
J. Am. Medical Informatics Assoc. | 7 |
| 2018 | An evaluation of clinical order patterns machine-learned from clinician cohorts stratified by patient mortality outcomesabstractEvaluate the quality of clinical order practice patterns machine-learned from clinician cohorts stratified by patient mortality outcomes. Inpatient electronic health records from 2010 to 2013 were extracted from a tertiary academic hospital. Clinicians (n = 1822) were stratified into low-mortality (21.8%, n = 397) and high-mortality (6.0%, n = 110) extremes using a two-sided P-value score quantifying deviation of observed vs. expected 30-day patient mortality rates. Three patient cohorts were assembled: patients seen by low-mortality clinicians, high-mortality clinicians, and an unfiltered crowd of all clinicians (n = 1046, 1046, and 5230 post-propensity score matching, respectively). Predicted order lists were automatically generated from recommender system algorithms trained on each patient cohort and evaluated against (i) real-world practice patterns reflected in patient cases with better-than-expected mortality outcomes and (ii) reference standards derived from clinical practice guidelines. Across six common admission diagnoses, order lists learned from the crowd demonstrated the greatest alignment with guideline references (AUROC range = 0.86–0.91), performing on par or better than those learned from low-mortality clinicians (0.79–0.84, P < 10−5) or manually-authored hospital order sets (0.65–0.77, P < 10−3). The same trend was observed in evaluating model predictions against better-than-expected patient cases, with the crowd model (AUROC mean = 0.91) outperforming the low-mortality model (0.87, P < 10−16) and order set benchmarks (0.78, P < 10−35). Whether machine-learning models are trained on all clinicians or a subset of experts illustrates a bias-variance tradeoff in data usage. Defining robust metrics to assess quality based on internal (e.g. practice patterns from better-than-expected patient cases) or external reference standards (e.g. clinical practice guidelines) is critical to assess decision support content. Learning relevant decision support content from all clinicians is as, if not more, robust than learning from a select subgroup of clinicians favored by patient outcomes. Jason K. Wang, Jason Hom, Santhosh Balasubramanian, Alejandro Schuler, Nigam H. Shah, Mary K. Goldstein, Michael T. M. Baiocchi, Jonathan H. Chen |
J. Biomed. Informatics | 8 |
| 2017 | Impact of Clinician Experience on Machine Learned Clinical Order Patterns
Jason K. Wang, Alejandro Schuler, Nigam H. Shah, Jonathan H. Chen |
AMIA | 4 |
| 2017 | Predicting inpatient clinical order patterns with probabilistic topic models vs conventional order setsabstractOBJECTIVE: Build probabilistic topic model representations of hospital admissions processes and compare the ability of such models to predict clinical order patterns as compared to preconstructed order sets. MATERIALS AND METHODS: The authors evaluated the first 24 hours of structured electronic health record data for > 10 K inpatients. Drawing an analogy between structured items (e.g., clinical orders) to words in a text document, the authors performed latent Dirichlet allocation probabilistic topic modeling. These topic models use initial clinical information to predict clinical orders for a separate validation set of > 4 K patients. The authors evaluated these topic model-based predictions vs existing human-authored order sets by area under the receiver operating characteristic curve, precision, and recall for subsequent clinical orders. RESULTS: Existing order sets predict clinical orders used within 24 hours with area under the receiver operating characteristic curve 0.81, precision 16%, and recall 35%. This can be improved to 0.90, 24%, and 47% ( P < 10 -20 ) by using probabilistic topic models to summarize clinical data into up to 32 topics. Many of these latent topics yield natural clinical interpretations (e.g., "critical care," "pneumonia," "neurologic evaluation"). DISCUSSION: Existing order sets tend to provide nonspecific, process-oriented aid, with usability limitations impairing more precise, patient-focused support. Algorithmic summarization has the potential to breach this usability barrier by automatically inferring patient context, but with potential tradeoffs in interpretability. CONCLUSION: Probabilistic topic modeling provides an automated approach to detect thematic trends in patient care and generate decision support content. A potential use case finds related clinical orders for decision support. Jonathan H. Chen, Mary K. Goldstein, Steven M. Asch, Lester Mackey, Russ B. Altman |
J. Am. Medical Informatics Assoc. | 1 |
| 2016 | Usability of an Automated Recommender System for Clinical Order Entry
Jonathan H. Chen, Mary K. Goldstein, Steven M. Asch, Russ B. Altman |
AMIA | 1 |
| 2016 | OrderRex: clinical order decision support and outcome predictions by data-mining electronic medical recordsabstractOBJECTIVE: To answer a "grand challenge" in clinical decision support, the authors produced a recommender system that automatically data-mines inpatient decision support from electronic medical records (EMR), analogous to Netflix or Amazon.com's product recommender. MATERIALS AND METHODS: EMR data were extracted from 1 year of hospitalizations (>18K patients with >5.4M structured items including clinical orders, lab results, and diagnosis codes). Association statistics were counted for the ∼1.5K most common items to drive an order recommender. The authors assessed the recommender's ability to predict hospital admission orders and outcomes based on initial encounter data from separate validation patients. RESULTS: Compared to a reference benchmark of using the overall most common orders, the recommender using temporal relationships improves precision at 10 recommendations from 33% to 38% (P < 10(-10)) for hospital admission orders. Relative risk-based association methods improve inverse frequency weighted recall from 4% to 16% (P < 10(-16)). The framework yields a prediction receiver operating characteristic area under curve (c-statistic) of 0.84 for 30 day mortality, 0.84 for 1 week need for ICU life support, 0.80 for 1 week hospital discharge, and 0.68 for 30-day readmission. DISCUSSION: Recommender results quantitatively improve on reference benchmarks and qualitatively appear clinically reasonable. The method assumes that aggregate decision making converges appropriately, but ongoing evaluation is necessary to discern common behaviors from "correct" ones. CONCLUSIONS: Collaborative filtering recommender algorithms generate clinical decision support that is predictive of real practice patterns and clinical outcomes. Incorporating temporal relationships improves accuracy. Different evaluation metrics satisfy different goals (predicting likely events vs. "interesting" suggestions). Jonathan H. Chen, Tanya Podchiyska, Russ B. Altman |
J. Am. Medical Informatics Assoc. | 1 |
| 2014 | "Doctors who ordered this also ordered..." Automated physician order recommendations and outcome predictions by data-mining electronic medical records
Jonathan H. Chen, Russ B. Altman |
AMIA | 1 |
| 2007 | ChemDB update - full-text search and virtual chemical spaceabstractUNLABELLED: ChemDB is a chemical database containing nearly 5M commercially available small molecules, important for use as synthetic building blocks, probes in systems biology and as leads for the discovery of drugs and other useful compounds. The data is publicly available over the web for download and for targeted searches using a variety of powerful methods. The chemical data includes predicted or experimentally determined physicochemical properties, such as 3D structure, melting temperature and solubility. Recent developments include optimization of chemical structure (and substructure) retrieval algorithms, enabling full database searches in less than a second. A text-based search engine allows efficient searching of compounds based on over 65M annotations from over 150 vendors. When searching for chemicals by name, fuzzy text matching capabilities yield productive results even when the correct spelling of a chemical name is unknown, taking advantage of both systematic and common names. Finally, built in reaction models enable searches through virtual chemical space, consisting of hypothetical products readily synthesizable from the building blocks in ChemDB. AVAILABILITY: ChemDB and Supplementary Materials are available at http://cdb.ics.uci.edu. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jonathan H. Chen, Erik Linstead, Sanjay Joshua Swamidass, Dennis Wang, Pierre Baldi |
Bioinform. | 1 |
| 2006 | Functional Census of Mutation Sequence Spaces: The Example of p53 Cancer Rescue MutantsabstractMany biomedical problems relate to mutant functional properties across a sequence space of interest, e.g., flu, cancer, and HIV. Detailed knowledge of mutant properties and function improves medical treatment and prevention. A functional census of p53 cancer rescue mutants would aid the search for cancer treatments from p53 mutant rescue. We devised a general methodology for conducting a functional census of a mutation sequence space by choosing informative mutants early. The methodology was tested in a double-blind predictive test on the functional rescue property of 71 novel putative p53 cancer rescue mutants iteratively predicted in sets of three (24 iterations). The first double-blind 15-point moving accuracy was 47 percent and the last was 86 percent; r = 0.01 before an epiphanic 16th iteration and r = 0.92 afterward. Useful mutants were chosen early (overall r = 0.80). Code and data are freely available (http://www.igb.uci.edu/research/research.html, corresponding authors: R.H.L. for computation and R.K.B. for biology). Samuel A. Danziger, Sanjay Joshua Swamidass, Jue Zeng, Lawrence R. Dearth, Jonathan H. Chen, Jianlin Cheng, Vinh P. Hoang, Hiroto Saigo, Ray Luo 0001, Pierre Baldi, Rainer K. Brachmann, Richard H. Lathrop |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2005 | ChemDB: a public database of small molecules and related chemoinformatics resourcesabstractMOTIVATION: The development of chemoinformatics has been hampered by the lack of large, publicly available, comprehensive repositories of molecules, in particular of small molecules. Small molecules play a fundamental role in organic chemistry and biology. They can be used as combinatorial building blocks for chemical synthesis, as molecular probes in chemical genomics and systems biology, and for the screening and discovery of new drugs and other useful compounds. RESULTS: We describe ChemDB, a public database of small molecules available on the Web. ChemDB is built using the digital catalogs of over a hundred vendors and other public sources and is annotated with information derived from these sources as well as from computational methods, such as predicted solubility and three-dimensional structure. It supports multiple molecular formats and is periodically updated, automatically whenever possible. The current version of the database contains approximately 4.1 million commercially available compounds and 8.2 million counting isomers. The database includes a user-friendly graphical interface, chemical reactions capabilities, as well as unique search capabilities. AVAILABILITY: Database and datasets are available on http://cdb.ics.uci.edu. Jonathan H. Chen, Sanjay Joshua Swamidass, Yimeng Dou, Jocelyne Bruand, Pierre Baldi |
Bioinform. | 1 |