EDBT 2026 Demo / reviewers in the wild / expert
Joshua C. Smith
dblp:124/3389
· DBLP profile ↗
10ranked-venue papers
4as first author
6since 2021 · last 2024
0000-0003-2661-3203ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Data-driven automated classification algorithms for acute health conditions: applying PheNorm to COVID-19 diseaseabstractOBJECTIVES: Automated phenotyping algorithms can reduce development time and operator dependence compared to manually developed algorithms. One such approach, PheNorm, has performed well for identifying chronic health conditions, but its performance for acute conditions is largely unknown. Herein, we implement and evaluate PheNorm applied to symptomatic COVID-19 disease to investigate its potential feasibility for rapid phenotyping of acute health conditions. MATERIALS AND METHODS: PheNorm is a general-purpose automated approach to creating computable phenotype algorithms based on natural language processing, machine learning, and (low cost) silver-standard training labels. We applied PheNorm to cohorts of potential COVID-19 patients from 2 institutions and used gold-standard manual chart review data to investigate the impact on performance of alternative feature engineering options and implementing externally trained models without local retraining. RESULTS: Models at each institution achieved AUC, sensitivity, and positive predictive value of 0.853, 0.879, 0.851 and 0.804, 0.976, and 0.885, respectively, at quantiles of model-predicted risk that maximize F1. We report performance metrics for all combinations of silver labels, feature engineering options, and models trained internally versus externally. DISCUSSION: Phenotyping algorithms developed using PheNorm performed well at both institutions. Performance varied with different silver-standard labels and feature engineering options. Models developed locally at one site also worked well when implemented externally at the other site. CONCLUSION: PheNorm models successfully identified an acute health condition, symptomatic COVID-19. The simplicity of the PheNorm approach allows it to be applied at multiple study sites with substantially reduced overhead compared to traditional approaches. Joshua C. Smith, Brian D. Williamson, David J. Cronkite, Daniel Park, Jill M. Whitaker, Michael F. McLemore, Joshua Osmanski, Robert Winter 0003, Arvind Ramaprasan, Ann Kelley, Mary Shea, Saranrat Wittayanukorn, Danijela Stojanovic, Yueqin Zhao, Sengwee Toh, Kevin B. Johnson, David Aronoff, David Carrell |
J. Am. Medical Informatics Assoc. | 1 |
| 2022 | Data-driven automated classification algorithms for acute health conditions: Applying PheNorm to COVID-19 disease
Joshua C. Smith, Daniel Park, Jill Whitaker Bey, Michael F. McLemore, Elizabeth Hanchrow, Dax Westerman, Joshua Osmanski, Robert Winter 0003, Arvind Ramaprasan, Ann Kelley, Mary Shea, David J. Cronkite, Saranrat Wittayanukorn, Danijela Stojanovic, Yueqin Zhao, Darren Toh, Kevin B. Johnson, David Aronoff, David Carrell |
AMIA | 1 |
| 2022 | Do electronic health record systems "dumb down" clinicians?abstractA panel sponsored by the American College of Medical Informatics (ACMI) at the 2021 AMIA Symposium addressed the provocative question: "Are Electronic Health Records dumbing down clinicians?" After reviewing electronic health record (EHR) development and evolution, the panel discussed how EHR use can impair care delivery. Both suboptimal functionality during EHR use and longer-term effects outside of EHR use can reduce clinicians' efficiencies, reasoning abilities, and knowledge. Panel members explored potential solutions to problems discussed. Progress will require significant engagement from clinician-users, educators, health systems, commercial vendors, regulators, and policy makers. Future EHR systems must become more user-focused and scalable and enable providers to work smarter to deliver improved care. Genevieve B. Melton, James J. Cimino, Christoph U. Lehmann, Patricia Sengstack, Joshua C. Smith, William M. Tierney, Randolph A. Miller |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | Predictive Model for Inpatient Mortality
Joshua C. Smith, Allison B. McCoy, John A. Morris, Asli Weitkamp |
AMIA | 1 |
| 2021 | Real-time clinical note monitoring to detect conditions for rapid follow-up: A case study of clinical trial enrollment in drug-induced torsades de pointes and Stevens-Johnson syndromeabstractIdentifying acute events as they occur is challenging in large hospital systems. Here, we describe an automated method to detect 2 rare adverse drug events (ADEs), drug-induced torsades de pointes and Stevens-Johnson syndrome and toxic epidermal necrolysis, in near real time for participant recruitment into prospective clinical studies. A text processing system searched clinical notes from the electronic health record (EHR) for relevant keywords and alerted study personnel via email of potential patients for chart review or in-person evaluation. Between 2016 and 2018, the automated recruitment system resulted in capture of 138 true cases of drug-induced rare events, improving recall from 43% to 93%. Our focused electronic alert system maintained 2-year enrollment, including across an EHR migration from a bespoke system to Epic. Real-time monitoring of EHR notes may accelerate research for certain conditions less amenable to conventional study recruitment paradigms. Sarah DeLozier, Peter Speltz, Jason Brito, Leigh Anne Tang, Janey Wang, Joshua C. Smith, Dario A. Giuse, Elizabeth Phillips, Kristina Williams, T. Stephen Strickland, Giovanni Davogustto, Dan M. Roden, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 6 |
| 2021 | ConceptWAS: A high-throughput method for early identification of COVID-19 presenting symptoms and characteristics from clinical notes
Juan Zhao 0003, Monika E. Grabowska, Vern Eric Kerchberger, Joshua C. Smith, H. Nur Eken, QiPing Feng, Josh F. Peterson, S. Trent Rosenbloom, Kevin B. Johnson, Wei-Qi Wei |
J. Biomed. Informatics | 4 |
| 2020 | Real-time Clinical Note Monitoring to Detect Conditions for Follow-up: a Case Study of Clinical Trial Enrollment in Drug-induced Torsades de Pointes and Stevens-Johnson Syndrome
Sarah DeLozier, Peter Speltz, Jason Brito, Leigh Anne Tang, Janey Wang, Joshua C. Smith, Dario A. Giuse, Elizabeth Phillips, Kristina Williams, Teresa Strickland, Giovanni Davogustto, Dan M. Roden, Joshua C. Denny |
AMIA | 6 |
| 2020 | Natural Language Processing and Machine Learning to Enable Clinical Decision Support for Treatment of Pediatric Pneumonia
Joshua C. Smith, Ashley Spann, Allison B. McCoy, Jakobi A. Johnson, Donald H. Arnold, Derek J. Williams, Asli Weitkamp |
AMIA | 1 |
| 2014 | Phenome-Wide Association Studies Using NLP-Derived Concepts
Pedro L. Teixeira, Robert J. Carroll, Lisa Bastarache, Peter Speltz, Joshua C. Smith, Joshua C. Denny |
AMIA | 5 |
| 2013 | Reducing patient re-identification risk for laboratory results within research datasetsabstractOBJECTIVE: To try to lower patient re-identification risks for biomedical research databases containing laboratory test results while also minimizing changes in clinical data interpretation. MATERIALS AND METHODS: In our threat model, an attacker obtains 5-7 laboratory results from one patient and uses them as a search key to discover the corresponding record in a de-identified biomedical research database. To test our models, the existing Vanderbilt TIME database of 8.5 million Safe Harbor de-identified laboratory results from 61 280 patients was used. The uniqueness of unaltered laboratory results in the dataset was examined, and then two data perturbation models were applied-simple random offsets and an expert-derived clinical meaning-preserving model. A rank-based re-identification algorithm to mimic an attack was used. The re-identification risk and the retention of clinical meaning for each model's perturbed laboratory results were assessed. RESULTS: Differences in re-identification rates between the algorithms were small despite substantial divergence in altered clinical meaning. The expert algorithm maintained the clinical meaning of laboratory results better (affecting up to 4% of test results) than simple perturbation (affecting up to 26%). DISCUSSION AND CONCLUSION: With growing impetus for sharing clinical data for research, and in view of healthcare-related federal privacy regulation, methods to mitigate risks of re-identification are important. A practical, expert-derived perturbation algorithm that demonstrated potential utility was developed. Similar approaches might enable administrators to select data protection scheme parameters that meet their preferences in the trade-off between the protection of privacy and the retention of clinical meaning of shared data. Ravi V. Atreya, Joshua C. Smith, Allison B. McCoy, Bradley A. Malin, Randolph A. Miller |
J. Am. Medical Informatics Assoc. | 2 |