EDBT 2026 Demo / reviewers in the wild / expert
Colin G. Walsh
dblp:200/4319
· DBLP profile ↗
34ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-9379-2056ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 33 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A scoping review of models to identify transgender patients in electronic health recordsabstractOBJECTIVE: Electronic health records (EHRs) lack a widely adopted standard for recording transgender and gender diverse (TGD) status, complicating research on TGD health. Computational models have been developed to identify TGD individuals in EHRs; however, gaps remain in understanding which components contribute to stronger phenotyping approaches. This scoping review evaluates EHR-based models for identifying TGD individuals, focusing on identifier types, performance, external validation, and ethical reporting to guide best practices. MATERIALS AND METHODS: We searched PubMed, CINAHL, Web of Science, and Embase for peer-reviewed articles published before January 2024, following PRISMA-ScR guidelines. Included studies used EHR data to identify TGD individuals, verified TGD status, reported or allowed calculation of positive predictive value (PPV), and listed identifiers. Two authors screened and extracted data. We categorized models by data type and logic (structured, unstructured, and multimodal), summarized PPV distributions, and synthesized author-reported ethical considerations. RESULTS: Fourteen studies describing 50 models met inclusion criteria. Models using TGD-related diagnostic codes alone (n = 11) or requiring both structured and unstructured data (n = 6) showed the highest mean PPVs (85.3% and 97.1%). Models validated on larger confirmed TGD cohorts reported more stable performance, but external validation was rare. Most studies minimally addressed ethics; only 3 described protective measures or stakeholder engagement. DISCUSSION: Phenotyping of TGD individuals in EHR data remains heterogeneous in design and ethical transparency. Reported PPVs should be interpreted cautiously, as performance is influenced by study design, sample size, and verification methods. CONCLUSIONS: Our recommendations emphasize the components that strengthen phenotyping approaches-identifier choice, multimodal intersection logic, validation practices, and ethical safeguards-rather than endorsing any single model. Robert A. Becker, Jhansi U. L. Kolli, Colin G. Walsh |
J. Am. Medical Informatics Assoc. | 3 |
| 2026 | Identifying and supporting trafficked individuals: provider and community organization perspectives on existing sociotechnical approachesabstractOBJECTIVES: Trafficked persons experience adverse health consequences and seek help, but many go unrecognized by health-care professionals. This study explored professionals' perspectives on current approaches toward identifying and supporting trafficked persons in health-care settings, highlighting current technology roles, gaps, and future directions. MATERIALS AND METHODS: We developed an interview guide to investigate current human trafficking (HT) approaches, safety procedures, and HT education. Semistructured interviews were conducted via Zoom, iteratively coded in Dedoose, and analyzed using a thematic analysis approach. RESULTS: We interviewed 19 health-care and community group professionals and identified 3 themes: (1) participants described a responsibility to build trust with patients through compassionate communication, rapport, and trauma-informed approaches across different stages of care. (2) Technology played a dual role, as professionals navigated both benefits and challenges of tools such as Zoom, virtual interpreters, and cameras in trust building. (3) Safety and privacy concerns guided how participants documented patient encounters and shared community resources, ensuring confidentiality while supporting patient and community well-being. DISCUSSION: Technology can both support and hinder trust in health care, directly affecting trafficked patients and their safety. Informatics can improve care for trafficked persons, but further research is needed on technology-based interventions. We provide recommendations to strengthen trust, enhance safety, support trauma-informed care, and promote safe documentation practices. CONCLUSION: Effective sociotechnical approaches rely on trust, safety, and mindful documentation to support trafficked patients. Future research directions include refining the role of informatics in trauma-informed care to strengthen trust and mitigate unintended consequences. Michelle Gomez, Ellen Wright Clayton, Colin G. Walsh, Kim M. Unertl |
J. Am. Medical Informatics Assoc. | 3 |
| 2022 | Prospective Validation of a Suicide Attempt Risk Model in Transgender Patients
Robert A. Becker, Allison B. McCoy, Colin G. Walsh |
AMIA | 3 |
| 2022 | Modeling association between clinician interventions and outcomes in Major Depressive Disorder with observational electronic health record data
Barrett Jones, Colin G. Walsh |
AMIA | 2 |
| 2021 | Unsupervised characterization of Major Depressive Disorder medication treatment pathways
Barrett Jones, Colin G. Walsh |
AMIA | 2 |
| 2021 | Predictive Modeling of Healthcare Utilization Metrics Identifies Adult Patients at High Risk for Suicide Attempt in the Primary Care Setting
Katherine Lee, Colin G. Walsh |
AMIA | 2 |
| 2021 | Action-oriented Artificial Intelligence for Suicide Risk Prediction: Prospective EHR-based Validation in a Large Clinical System
Michael Ripperger, Drew Wilimitis, William W. Stead, Kevin B. Johnson, Colin G. Walsh |
AMIA | 5 |
| 2021 | Complementing Automated Risk Prediction with Face-to-face Screening Improves Suicide Risk Prediction
Drew Wilimitis, Robert W. Turer, Michael Ripperger, Allison B. McCoy, Sarah H. Sperry, Colin G. Walsh |
AMIA | 6 |
| 2021 | Ensemble learning to predict opioid-related overdose using statewide prescription drug monitoring program and hospital discharge data in the state of TennesseeabstractOBJECTIVE: To develop and validate algorithms for predicting 30-day fatal and nonfatal opioid-related overdose using statewide data sources including prescription drug monitoring program data, Hospital Discharge Data System data, and Tennessee (TN) vital records. Current overdose prevention efforts in TN rely on descriptive and retrospective analyses without prognostication. MATERIALS AND METHODS: Study data included 3 041 668 TN patients with 71 479 191 controlled substance prescriptions from 2012 to 2017. Statewide data and socioeconomic indicators were used to train, ensemble, and calibrate 10 nonparametric "weak learner" models. Validation was performed using area under the receiver operating curve (AUROC), area under the precision recall curve, risk concentration, and Spiegelhalter z-test statistic. RESULTS: Within 30 days, 2574 fatal overdoses occurred after 4912 prescriptions (0.0069%) and 8455 nonfatal overdoses occurred after 19 460 prescriptions (0.027%). Discrimination and calibration improved after ensembling (AUROC: 0.79-0.83; Spiegelhalter P value: 0-.12). Risk concentration captured 47-52% of cases in the top quantiles of predicted probabilities. DISCUSSION: Partitioning and ensembling enabled all study data to be used given computational limits and helped mediate case imbalance. Predicting risk at the prescription level can aggregate risk to the patient, provider, pharmacy, county, and regional levels. Implementing these models into Tennessee Department of Health systems might enable more granular risk quantification. Prospective validation with more recent data is needed. CONCLUSION: Predicting opioid-related overdose risk at statewide scales remains difficult and models like these, which required a partnership between an academic institution and state health agency to develop, may complement traditional epidemiological methods of risk identification and inform public health decisions. Michael Ripperger, Sarah C. Lotspeich, Drew Wilimitis, Carrie E. Fry, Allison Roberts, Matthew C. Lenert, Charlotte Cherry, Sanura Latham, Katelyn Robinson, Qingxia Chen, Melissa McPheeters, Ben Tyndall, Colin G. Walsh |
J. Am. Medical Informatics Assoc. | 13 |
| 2021 | Combatting human trafficking in the United States: how can medical informatics help?abstractOBJECTIVE: Human trafficking is a global problem taking many forms, including sex and labor exploitation. Trafficking victims can be any age, although most trafficking begins when victims are adolescents. Many trafficking victims have contact with health-care providers across various health-care contexts, both for emergency and routine care. MATERIALS AND METHODS: We propose 4 specific areas where medical informatics can assist with combatting trafficking: screening, clinical decision support, community-facing tools, and analytics that are both descriptive and predictive. Efforts to implement health information technology interventions focused on trafficking must be carefully integrated into existing clinical work and connected to community resources to move beyond identification to provide assistance and to support trauma-informed care. RESULTS: We lay forth a research and implementation agenda to integrate human trafficking identification and intervention into routine clinical practice, supported by health information technology. CONCLUSIONS: A sociotechnical systems approach is recommended to ensure interventions address the complex issues involved in assisting victims of human trafficking. Kim M. Unertl, Colin G. Walsh, Ellen Wright Clayton |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | A Surveillance Framework for Monitoring and Updating Clinical Prediction Models
Sharon E. Davis, Robert A. Greevy Jr., Thomas A. Lasko, Colin G. Walsh, Michael E. Matheny |
AMIA | 4 |
| 2020 | Translational Research of Machine Learning and Artificial Intelligence Advances in Clinical Settings - Experiences and Challenges
William R. Hersh, Gretchen Purcell Jackson, Marc S. Williams, Colin G. Walsh, David A. Dorr |
AMIA | 4 |
| 2020 | Detection of calibration drift in clinical prediction models to inform model updating
Sharon E. Davis, Robert A. Greevy Jr., Thomas A. Lasko, Colin G. Walsh, Michael E. Matheny |
J. Biomed. Informatics | 4 |
| 2019 | Comparison of Prediction Model Performance Updating Protocols: Using a Data-Driven Testing Procedure to Guide Updating
Sharon E. Davis, Robert A. Greevy Jr., Thomas A. Lasko, Colin G. Walsh, Michael E. Matheny |
AMIA | 4 |
| 2019 | Testing the Feasibility of an Academic-State Partnership to Combat the Opioid Epidemic in Tennessee Through Predictive Analytics
Christopher A. Puchi, Ben Tyndall, Melissa McPheeters, Colin G. Walsh |
AMIA | 4 |
| 2019 | A nonparametric updating method to correct clinical prediction model driftabstractOBJECTIVE: Clinical prediction models require updating as performance deteriorates over time. We developed a testing procedure to select updating methods that minimizes overfitting, incorporates uncertainty associated with updating sample sizes, and is applicable to both parametric and nonparametric models. MATERIALS AND METHODS: We describe a procedure to select an updating method for dichotomous outcome models by balancing simplicity against accuracy. We illustrate the test's properties on simulated scenarios of population shift and 2 models based on Department of Veterans Affairs inpatient admissions. RESULTS: In simulations, the test generally recommended no update under no population shift, no update or modest recalibration under case mix shifts, intercept correction under changing outcome rates, and refitting under shifted predictor-outcome associations. The recommended updates provided superior or similar calibration to that achieved with more complex updating. In the case study, however, small update sets lead the test to recommend simpler updates than may have been ideal based on subsequent performance. DISCUSSION: Our test's recommendations highlighted the benefits of simple updating as opposed to systematic refitting in response to performance drift. The complexity of recommended updating methods reflected sample size and magnitude of performance drift, as anticipated. The case study highlights the conservative nature of our test. CONCLUSIONS: This new test supports data-driven updating of models developed with both biostatistical and machine learning approaches, promoting the transportability and maintenance of a wide array of clinical prediction models and, in turn, a variety of applications relying on modern prediction tools. Sharon E. Davis, Robert A. Greevy Jr., Christopher Fonnesbeck, Thomas A. Lasko, Colin G. Walsh, Michael E. Matheny |
J. Am. Medical Informatics Assoc. | 5 |
| 2019 | Prognostic models will be victims of their own success, unlessabstractPredictive analytics have begun to change the workflows of healthcare by giving insight into our future health. Deploying prognostic models into clinical workflows should change behavior and motivate interventions that affect outcomes. As users respond to model predictions, downstream characteristics of the data, including the distribution of the outcome, may change. The ever-changing nature of healthcare necessitates maintenance of prognostic models to ensure their longevity. The more effective a model and intervention(s) are at improving outcomes, the faster a model will appear to degrade. Improving outcomes can disrupt the association between the model's predictors and the outcome. Model refitting may not always be the most effective response to these challenges. These problems will need to be mitigated by systematically incorporating interventions into prognostic models and by maintaining robust performance surveillance of models in clinical use. Holistically modeling the outcome and intervention(s) can lead to resilience to future compromises in performance. Matthew C. Lenert, Michael E. Matheny, Colin G. Walsh |
J. Am. Medical Informatics Assoc. | 3 |
| 2019 | Explicit causal reasoning is preferred, but not necessary for pragmatic valueabstractIn their response, researchers Sperrin, Jenkins, Martin, and Peek discuss some of the benefits of applying causal inference frameworks (CIFs) to predict treatment naïve risk in the domain of risk modeling. We agree that causality-based models using diagrams are a powerful tool and that these models can avoid the pitfalls of model-mediated changes to the outcome process.1 CIFs have also demonstrated robustness to unobserved confounders.2 There are many reasons why explicitly considering causality and estimating baseline risk in the absence of treatments are important when deploying and maintaining prognostic models in clinical operations. While these models have many desirable properties, they are not without their challenges, as Sperrin et al note. CIFs demonstrate a firm understanding of the processes one wishes to improve. Getting to the requisite level of insight to build such a diagram is a long and arduous scientific process. This is not to say many processes cannot be diagramed using current knowledge. We feel that incorporating causality where it is well understood is useful, but there are circumstances in which CIFs are likely to be incorrect and have the potential to cause error. Furthermore, causal models require data elements that reflect how a process works. Current bulwark data streams (revenue cycle-focused electronic health records) are not likely to include data relevant or sufficient for CIFs for a number (if not most) use-cases. Matthew C. Lenert, Michael E. Matheny, Colin G. Walsh |
J. Am. Medical Informatics Assoc. | 3 |
| 2019 | Maintaining automated measurement of Choosing Wisely adherence across the ICD 9 to 10 transition
John Angiolillo, S. Trent Rosenbloom, Melissa McPheeters, G. Seibert Tregoning, Russell L. Rothman, Colin G. Walsh |
J. Biomed. Informatics | 6 |
| 2019 | A method for analyzing inpatient care variability through physicians' orders
Matthew C. Lenert, Randolph A. Miller, Yevgeniy Vorobeychik, Colin G. Walsh |
J. Biomed. Informatics | 4 |
| 2018 | Balancing Performance and Interpretability: Selecting Features with Bootstrapped Ridge Regression
Matthew C. Lenert, Colin G. Walsh |
AMIA | 2 |
| 2018 | The Intersection of Data Science, People, and Organizations in Health Care: An Interactive Discussion of Challenges and Solutions
Laurie L. Novak, Rupa Valdez, Colin G. Walsh, Hojjat Salmasian, Eleanor Wynn |
AMIA | 3 |
| 2018 | Discovering hidden knowledge through auditing clinical diagnostic knowledge bases
Matthew C. Lenert, Colin G. Walsh, Randolph A. Miller |
J. Biomed. Informatics | 2 |
| 2017 | Critical Appraisal of Models for Prediction of Readmission (CAMPR): A Quality Tool to Assess Models that Predict Hospital Readmissions
Lisa Grossman Liu, Rollin R. Reeder, Colin G. Walsh, Devan Kansagara, David K. Vawdrey, Hojjat Salmasian |
AMIA | 3 |
| 2017 | X Marks the Spot: Mapping Similarity Between Clinical Trial Cohorts and US Counties
Matthew C. Lenert, Dara Eckerle Mize, Colin G. Walsh |
AMIA | 3 |
| 2017 | Beyond discrimination: A comparison of calibration methods and clinical usefulness of predictive models of readmission risk
Colin G. Walsh, Kavya Sharman, George Hripcsak |
J. Biomed. Informatics | 1 |
| 2016 | An Analysis of Readmission Events Over Time in Patients with Type 2 Diabetes
Dara Eckerle Mize, Mia A. Levy, Shubhada Jagasia, Colin G. Walsh |
AMIA | 4 |
| 2016 | Choosing Wisely Using ICD10: The impact of the ICD10 Transition on Prevalence and Cost Estimates of Low Value Healthcare Services
Colin G. Walsh, S. Trent Rosenbloom |
AMIA | 1 |
| 2015 | Outcomes Prediction via Time Intervals Related PatternsabstractThe increasing availability of multivariate temporal data in many domains, such as biomedical, security and more, provides exceptional opportunities for temporal knowledge discovery, classification and prediction, but also challenges. Temporal variables are often sparse and in many domains, such as in biomedical data, they have huge number of variables. In recent decades in the biomedical domain events, such as conditions, drugs and procedures, are stored as time intervals, which enables to discover Time Intervals Related Patterns (TIRPs) and use for classification or prediction. In this study we present a framework for outcome events prediction, called Maitreya, which includes an algorithm for TIRPs discovery called KarmaLegoD, designed to handle huge number of symbols. Three indexing strategies for pairs of symbolic time intervals are proposed and compared, showing that the use of FullyHashed indexing is only slightly slower but consumes minimal memory. We evaluated Maitreya on eight real datasets for the prediction of clinical procedures as outcome events. The use of TIRPs outperform the use of symbols, especially with horizontal support (number of instances) as TIRPs feature representation. Robert Moskovitch, Colin G. Walsh, Fei Wang 0001, George Hripcsak, Nicholas P. Tatonetti |
ICDM | 2 |
| 2014 | An Integrated Billing Application to Streamline Clinician Workflow
David K. Vawdrey, Colin G. Walsh, Peter D. Stetson |
AMIA | 2 |
| 2014 | Enabling claims-based decision support through non-interruptive capture of admission diagnoses and provider billing codes
Colin G. Walsh, David K. Vawdrey, Peter D. Stetson, Matthew R. Fred, George Hripcsak |
AMIA | 1 |
| 2014 | The effects of data sources, cohort selection, and outcome definition on a predictive model of risk of thirty-day hospital readmissionsabstractBACKGROUND: Hospital readmission risk prediction remains a motivated area of investigation and operations in light of the hospital readmissions reduction program through CMS. Multiple models of risk have been reported with variable discriminatory performances, and it remains unclear how design factors affect performance. OBJECTIVES: To study the effects of varying three factors of model development in the prediction of risk based on health record data: (1) reason for readmission (primary readmission diagnosis); (2) available data and data types (e.g. visit history, laboratory results, etc); (3) cohort selection. METHODS: Regularized regression (LASSO) to generate predictions of readmissions risk using prevalence sampling. Support Vector Machine (SVM) used for comparison in cohort selection testing. Calibration by model refitting to outcome prevalence. RESULTS: Predicting readmission risk across multiple reasons for readmission resulted in ROC areas ranging from 0.92 for readmission for congestive heart failure to 0.71 for syncope and 0.68 for all-cause readmission. Visit history and laboratory tests contributed the most predictive value; contributions varied by readmission diagnosis. Cohort definition affected performance for both parametric and nonparametric algorithms. Compared to all patients, limiting the cohort to patients whose index admission and readmission diagnoses matched resulted in a decrease in average ROC from 0.78 to 0.55 (difference in ROC 0.23, p value 0.01). Calibration plots demonstrate good calibration with low mean squared error. CONCLUSION: Targeting reason for readmission in risk prediction impacted discriminatory performance. In general, laboratory data and visit history data contributed the most to prediction; data source contributions varied by reason for readmission. Cohort selection had a large impact on model performance, and these results demonstrate the difficulty of comparing results across different studies of predictive risk modeling. Colin G. Walsh, George Hripcsak |
J. Biomed. Informatics | 1 |
| 2013 | Physician Know Thyself: EHRs and the Quantified Clinician
Daniel M. Stein, Eugenia Siegler, David K. Vawdrey, Marc Sturm, Niloo Sobhani, Soumitra Sengupta, Colin G. Walsh, Gilad J. Kuperman |
AMIA | 7 |
| 2012 | EHR on the Move: Resident Physician Perceptions of iPads and the Clinical Workflow
Colin G. Walsh, Peter D. Stetson |
AMIA | 1 |