Evan Sholle

dblp:212/7823 · also Evan T. Sholle · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0001-9518-4399ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 18 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From use cases to infrastructure: a cross-institutional survey of priorities in data-driven biomedical research
abstract
OBJECTIVES: Federated Ecosystems for Analytics and Standardized Technologies (FEAST) is a modular, cloud-based platform developed through the ARPA-H Biomedical Data Fabric initiative to enable secure, federated analysis of real-world biomedical data. To guide and iteratively refine its modular design, the FEAST team conducted a cross-institutional survey to systematically identify and prioritize research needs related to authorized-access data across diverse biomedical domains. This study presents a structured synthesis of submitted use cases to uncover infrastructure gaps, data integration challenges, and translational opportunities. The results from the survey inform both front-end user-facing functionality and backend data requirements, shaping how the interface supports user interactions, data types, and compliance with security and interoperability standards. MATERIALS AND METHODS: A structured survey form was distributed to researchers affiliated with participating institutions, including DNA-HIVE, The George Washington University (GW-FEAST), Weill Cornell Medicine, Vanderbilt University Medical Center, Georgetown University, European Bioinformatics Institute, and Kaiser Permanente. Respondents completed standardized fields describing the data types of interest, project goals, analytic methods, and perceived technical barriers. The collected responses were curated and analyzed to identify common needs related to privacy, interoperability, scalability, and workflow reproducibility. RESULTS: The survey compiled 61 use cases spanning genomics, imaging, clinical phenotyping, EHR-driven analytics, and precision medicine. Common themes included the need for multi-modal data integration, HL7 FHIR-based secure access, federated model training without PII retention, and containerized microservices for scalable deployment. Convergent needs across institutions emphasized consistent demand for FAIR-compliant infrastructure and readiness for real-world data analytics. CONCLUSION: The FEAST Use Cases survey provides a cross-sectional view of biomedical informatics priorities grounded in real-world data needs. The findings offer a strategic blueprint for developing federated, privacy-preserving infrastructure to support secure, collaborative, and scalable biomedical research.
Raja Mazumder, Jonathon Keeney, Luke Johnson, Lori Krammer, Patrick McNeely, Jorge Sepúlveda, Danielle Hangen, Maria Jesus Martin, Dushyanth Jyothi, Jonas De Almeida, Peter B. McGarvey, Adil Alaoui, Sarah Cha, Art Sedrakyan, Evan Sholle, Michael Matheny, Michele L. LeNoue-Newton, Robert Winter 0003, Steve Deppen, Vahan Simonyan, Anelia Horvath
J. Am. Medical Informatics Assoc.15
2023 A method to automate the discharge summary hospital course for neurology patients
abstract
OBJECTIVE: Generation of automated clinical notes has been posited as a strategy to mitigate physician burnout. In particular, an automated narrative summary of a patient's hospital stay could supplement the hospital course section of the discharge summary that inpatient physicians document in electronic health record (EHR) systems. In the current study, we developed and evaluated an automated method for summarizing the hospital course section using encoder-decoder sequence-to-sequence transformer models. MATERIALS AND METHODS: We fine-tuned BERT and BART models and optimized for factuality through constraining beam search, which we trained and tested using EHR data from patients admitted to the neurology unit of an academic medical center. RESULTS: The approach demonstrated good ROUGE scores with an R-2 of 13.76. In a blind evaluation, 2 board-certified physicians rated 62% of the automated summaries as meeting the standard of care, which suggests the method may be useful clinically. DISCUSSION AND CONCLUSION: To our knowledge, this study is among the first to demonstrate an automated method for generating a discharge summary hospital course that approaches a quality level of what a physician would write.
Vince C. Hartman, Sanika S. Bapat, Mark G. Weiner, Babak B. Navi, Evan Sholle, Thomas R. Campion Jr.
J. Am. Medical Informatics Assoc.5
2022 Design and validation of a FHIR-based EHR-driven phenotyping toolbox
abstract
OBJECTIVES: To develop and validate a standards-based phenotyping tool to author electronic health record (EHR)-based phenotype definitions and demonstrate execution of the definitions against heterogeneous clinical research data platforms. MATERIALS AND METHODS: We developed an open-source, standards-compliant phenotyping tool known as the PhEMA Workbench that enables a phenotype representation using the Fast Healthcare Interoperability Resources (FHIR) and Clinical Quality Language (CQL) standards. We then demonstrated how this tool can be used to conduct EHR-based phenotyping, including phenotype authoring, execution, and validation. We validated the performance of the tool by executing a thrombotic event phenotype definition at 3 sites, Mayo Clinic (MC), Northwestern Medicine (NM), and Weill Cornell Medicine (WCM), and used manual review to determine precision and recall. RESULTS: An initial version of the PhEMA Workbench has been released, which supports phenotype authoring, execution, and publishing to a shared phenotype definition repository. The resulting thrombotic event phenotype definition consisted of 11 CQL statements, and 24 value sets containing a total of 834 codes. Technical validation showed satisfactory performance (both NM and MC had 100% precision and recall and WCM had a precision of 95% and a recall of 84%). CONCLUSIONS: We demonstrate that the PhEMA Workbench can facilitate EHR-driven phenotype definition, execution, and phenotype sharing in heterogeneous clinical research data environments. A phenotype definition that integrates with existing standards-compliant systems, and the use of a formal representation facilitates automation and can decrease potential for human error.
Pascal S. Brandt, Jennifer A. Pacheco, Prakash Adekkanattu, Evan Sholle, Sajjad Abedian, Daniel J. Stone, David Knaack, Jie Xu 0012, Yifan Peng 0002, Natalie C. Benda, Fei Wang 0001, Yuan Luo 0001, Guoqian Jiang, Jyotishman Pathak, Luke V. Rasmussen
J. Am. Medical Informatics Assoc.4
2022 An architecture for research computing in health to support clinical and translational investigators with electronic patient data
abstract
OBJECTIVE: Obtaining electronic patient data, especially from electronic health record (EHR) systems, for clinical and translational research is difficult. Multiple research informatics systems exist but navigating the numerous applications can be challenging for scientists. This article describes Architecture for Research Computing in Health (ARCH), our institution's approach for matching investigators with tools and services for obtaining electronic patient data. MATERIALS AND METHODS: Supporting the spectrum of studies from populations to individuals, ARCH delivers a breadth of scientific functions-including but not limited to cohort discovery, electronic data capture, and multi-institutional data sharing-that manifest in specific systems-such as i2b2, REDCap, and PCORnet. Through a consultative process, ARCH staff align investigators with tools with respect to study design, data sources, and cost. Although most ARCH services are available free of charge, advanced engagements require fee for service. RESULTS: Since 2016 at Weill Cornell Medicine, ARCH has supported over 1200 unique investigators through more than 4177 consultations. Notably, ARCH infrastructure enabled critical coronavirus disease 2019 response activities for research and patient care. DISCUSSION: ARCH has provided a technical, regulatory, financial, and educational framework to support the biomedical research enterprise with electronic patient data. Collaboration among informaticians, biostatisticians, and clinicians has been critical to rapid generation and analysis of EHR data. CONCLUSION: A suite of tools and services, ARCH helps match investigators with informatics systems to reduce time to science. ARCH has facilitated research at Weill Cornell Medicine and may provide a model for informatics and research leaders to support scientists elsewhere.
Thomas R. Campion Jr., Evan Sholle, Jyotishman Pathak, Stephen B. Johnson, John P. Leonard, Curtis L. Cole
J. Am. Medical Informatics Assoc.2
2021 Impact of Social Determinants of Health on Predictive Models in 30-Day Hospital Readmission or Death for Patients with Severe Obesity
Marianne Sharko, Yongkang Zhang 0004, Yiye Zhang, Evan Sholle, Sajjad Abedian, Meghan Reading Turchioe, Jessica S. Ancker
AMIA4
2021 Comparing Automated Extraction to Manual Chart Review for COVID-Specific Research Data Abstraction: A Case Study
Andrew L. Yin, Winston L. Guo, Evan Sholle, Mangala Rajan, Laura C. Pinheiro, Parag Goyal, Justin Choi, Mark N. Alshak, Graham T. Wehmeyer, Mark G. Weiner, Monika M. Safford, Thomas R. Campion Jr., Curtis L. Cole
AMIA3
2021 Leveraging Deep Representations of Radiology Reports in Survival Analysis for Predicting Heart Failure Patient Mortality
abstract
Utilizing clinical texts in survival analysis is difficult because they are largely unstructured. Current automatic extraction models fail to capture textual information comprehensively since their labels are limited in scope. Furthermore, they typically require a large amount of data and high-quality expert annotations for training. In this work, we present a novel method of using BERT-based hidden layer representations of clinical texts as covariates for proportional hazards models to predict patient survival outcomes. We show that hidden layers yield notably more accurate predictions than predefined features, outperforming the previous baseline model by 5.7% on average across C-index and time-dependent AUC. We make our work publicly available at https://github.com/bionlplab/heart_failure_mortality.
Hyun Gi Lee, Evan Sholle, Ashley Beecy, Subhi J. Al'Aref, Yifan Peng 0002
NAACL-HLT2
2021 Critical carE Database for Advanced Research (CEDAR): An automated method to support intensive care units with electronic health record data
Edward J. Schenck, Katherine L. Hoffman, Marika M. Cusick, Joseph Kabariti, Evan Sholle, Thomas R. Campion Jr.
J. Biomed. Informatics5
2020 Feasibility of Cross-Platform EHR-Driven Phenotyping Using Clinical Quality Language
Pascal S. Brandt, Richard C. Kiefer, Jennifer A. Pacheco, Prakash Adekkanattu, Evan Sholle, Faraz S. Ahmad, Jie Xu 0012, Jessica S. Ancker, Fei Wang 0001, Yuan Luo 0001, Guoqian Jiang, Jyotishman Pathak, Luke V. Rasmussen
AMIA5
2020 Weak Supervision to Classify Unstructured Clinical Text for Current Suicidal Ideation
Marika M. Cusick, Prakash Adekkanattu, Thomas R. Campion Jr., Evan Sholle, Annie C. Myers, George Alexopoulos, Jyotishman Pathak
AMIA4
2020 Extracting and classifying diagnosis dates from clinical notes: A case study
Julia T. Fu, Evan Sholle, Spencer Krichevsky, Joseph Scandura, Thomas R. Campion Jr.
J. Biomed. Informatics2
2019 Validation of a Computable Phenotype for Site-Specific Cancer Treatment
Evan Sholle, Orrin Belden, Jaclyn Rosenzweig, Jennifer Levine, Joseph Kabariti, Thomas R. Campion Jr., Jyotishman Pathak
AMIA1
2019 Automated Information Extraction to Support Response Assessment in Myeloproliferative Neoplasms
Evan Sholle, Spencer Krichevsky, Sajjad Abedian, Prakash Adekkanattu, Diana Jaber, Niamh Savage, Joseph Scandura, Thomas R. Campion Jr.
AMIA1
2019 Informatics approaches to collecting, analyzing, and addressing social determinants of health in healthcare
Yiye Zhang, Evan Sholle, Marianne Sharko, Yongkang Zhang 0004, Jessica S. Ancker
AMIA2
2019 Underserved populations with missing race ethnicity data differ significantly from those with structured race/ethnicity documentation
abstract
OBJECTIVE: We aimed to address deficiencies in structured electronic health record (EHR) data for race and ethnicity by identifying black and Hispanic patients from unstructured clinical notes and assessing differences between patients with or without structured race/ethnicity data. MATERIALS AND METHODS: Using EHR notes for 16 665 patients with encounters at a primary care practice, we developed rule-based natural language processing (NLP) algorithms to classify patients as black/Hispanic. We evaluated performance of the method against an annotated gold standard, compared race and ethnicity between NLP-derived and structured EHR data, and compared characteristics of patients identified as black or Hispanic using only NLP vs patients identified as such only in structured EHR data. RESULTS: For the sample of 16 665 patients, NLP identified 948 additional patients as black, a 26%increase, and 665 additional patients as Hispanic, a 20% increase. Compared with the patients identified as black or Hispanic in structured EHR data, patients identified as black or Hispanic via NLP only were older, more likely to be male, less likely to have commercial insurance, and more likely to have higher comorbidity. DISCUSSION: Structured EHR data for race and ethnicity are subject to data quality issues. Supplementing structured EHR race data with NLP-derived race and ethnicity may allow researchers to better assess the demographic makeup of populations and draw more accurate conclusions about intergroup differences in health outcomes. CONCLUSIONS: Black or Hispanic patients who are not documented as such in structured EHR race/ethnicity fields differ significantly from those who are. Relatively simple NLP can help address this limitation.
Evan Sholle, Laura C. Pinheiro, Prakash Adekkanattu, Marcos Davila, Stephen B. Johnson, Jyotishman Pathak, Sanjai Sinha, Cassidie Li, Stasi A. Lubansky, Monika M. Safford, Thomas R. Campion Jr.
J. Am. Medical Informatics Assoc.1
2018 Ascertaining Depression Severity by Extracting Patient Health Questionnaire-9 (PHQ-9) Scores from Clinical Notes
Prakash Adekkanattu, Evan Sholle, Joseph DeFerio, Jyotishman Pathak, Stephen B. Johnson, Thomas R. Campion Jr.
AMIA2
2018 Methods for Integrating EHRs, Social Determinants of Health, and Built Environment Data for Patient-Centered Research
Joseph DeFerio, Evan Sholle, Kathleen Lee, Jyotishman Pathak
AMIA2
2018 A scalable method for supporting multiple patient cohort discovery projects using i2b2
Evan Sholle, Marcos Davila, Joseph Kabariti, Julian Z. Schwartz, Vinay I. Varughese, Curtis L. Cole, Thomas R. Campion Jr.
J. Biomed. Informatics1
2017 Secondary Use of Patients' Electronic Records (SUPER): An Approach for Meeting Specific Data Needs of Clinical and Translational Researchers
Evan Sholle, Joseph Kabariti, Stephen B. Johnson, John P. Leonard, Jyotishman Pathak, Vinay I. Varughese, Curtis L. Cole, Thomas R. Campion Jr.
AMIA1