EDBT 2026 Demo / reviewers in the wild / expert
John Michael Gaziano
dblp:383/9411
· DBLP profile ↗
15ranked-venue papers
0as first author
10since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAIGE-GPU: accelerating genome- and phenome-wide association studies using GPUsabstractMOTIVATION: Genome-wide association studies (GWAS) at biobank scale are computationally intensive, especially for admixed populations requiring robust statistical models. SAIGE is a widely used method for generalized linear mixed-model GWAS but is limited by its CPU-based implementation, making phenome-wide association studies impractical for many research groups. RESULTS: We developed SAIGE-GPU, a GPU-accelerated version of SAIGE that replaces CPU-intensive matrix operations with GPU-optimized kernels. The core innovation is distributing genetic relationship matrix calculations across GPUs and communication layers. Applied to 2068 phenotypes from 635 969 participants in the Million Veteran Program, including diverse and admixed populations, SAIGE-GPU achieved a 5-fold speedup in mixed model fitting on supercomputing infrastructure and cloud platforms. We further optimized the variant association testing step through multi-core and multi-trait parallelization. Deployed on Google Cloud Platform and Azure, the method provided substantial cost and time savings. AVAILABILITY AND IMPLEMENTATION: Source code and binaries are available for download at https://github.com/saigegit/SAIGE/tree/SAIGE-GPU-1.3.3. A code snapshot is archived at Zenodo for reproducibility (DOI: [10.5281/zenodo.17642591]). SAIGE-GPU is available in a containerized format for use across HPC and cloud environments and is implemented in R/C++ and runs on Linux systems. Alex Rodriguez, Youngdae Kim, Tarak Nath Nandi, Karl Keat, Rachit Kumar, Mitchell Conery, Rohan Bhukar, Molei Liu, John Hessington, Ketan Maheshwari, VA Million Veteran Program, Edmon Begoli, Georgia Tourassi, Pradeep Natarajan, Benjamin F. Voight, John Michael Gaziano, Scott M. Damrauer, Katherine P. Liao, Jennifer E. Huffman, Anurag Verma, Ravi K. Madduri |
Bioinform. | 16 |
| 2025 | Effectiveness of electronic medical record-based strategies for death and hospital admission endpoint capture in pragmatic clinical trialsabstractOBJECTIVE: Event capture in clinical trials is resource-intensive, and electronic medical records (EMRs) offer a potential solution. This study develops algorithms for EMR-based death and hospitalization capture and compares them with traditional event capture methods. MATERIALS AND METHODS: We compared the effectiveness of EMR-based event capture and site-captured events adjudicated by a clinical endpoint committee in the multi-center INfluenza Vaccine to Effectively Stop cardio Thoracic Events and Decompensated heart failure (INVESTED) trial for participants from the Veterans Affairs healthcare system. Varying time windows around event dates were used to optimize events matching. The algorithms were externally validated for heart failure hospitalizations in the Medical Information Mart for Intensive Care (MIMIC)-IV database. RESULTS: We observed 100% sensitivity for death events with a 1-day window. Sensitivity for cardiovascular, heart failure, pulmonary, and nonspecific cardiopulmonary hospitalizations using discharge diagnosis codes varied between 75% and 95%. Including Centers for Medicare & Medicaid Services data improved sensitivity with no meaningful decrease in specificity. The MIMIC-IV analysis showed 82% sensitivity and 99% specificity for heart failure hospitalizations. DISCUSSION: EMR-based method accurately identifies all-cause mortality and demonstrates high accuracy for cardiopulmonary hospitalizations. This study underscores the importance of optimal time windows, data completeness, and domain variability in EMR systems. CONCLUSION: EMR-based methods are effective strategies for capturing death and hospitalizations in clinical trials; however, their effectiveness may be influenced by the complexity of events and domain variability across different EMR systems. Nonetheless, EMR-based methods can serve as a valuable complement to traditional methods. Maryam Rahafrooz, Danne C. Elbers, Jay R. Gopal, Junling Ren, Nathan H. Chan, Cenk Yildirim, Akshay S. Desai, Abigail A. Santos, Karen Murray, Thomas Havighurst, Jacob A. Udell, Michael E. Farkouh, Lawton Cooper, John Michael Gaziano, Orly Vardeny, Lu Mao, KyungMann Kim, David R. Gagnon, Scott D. Solomon, Jacob Joseph |
J. Am. Medical Informatics Assoc. | 14 |
| 2025 | ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis
Ziming Gan, Doudou Zhou, Everett Neil Rush, Vidul Ayakulangara Panickan, Yuk-Lam Ho, George Ostrouchov, Shuting Shen, Xin Xiong 0006, Kimberly F. Greco, Chuan Hong, Clara-Lea Bonzel, Jun Wen 0001, Lauren Costa, Tianrun A. Cai, Edmon Begoli, Zongqi Xia, John Michael Gaziano, Katherine P. Liao, Kelly Cho, Tianxi Cai |
J. Biomed. Informatics | 18 |
| 2025 | DOME: Directional medical embedding vectors from Electronic Health RecordsabstractMOTIVATION: The increasing availability of Electronic Health Record (EHR) systems has created enormous potential for translational research. Recent developments in representation learning techniques have led to effective large-scale representations of EHR concepts along with knowledge graphs that empower downstream EHR studies. However, most existing methods require training with patient-level data, limiting their abilities to expand the training with multi-institutional EHR data. On the other hand, scalable approaches that only require summary-level data do not incorporate temporal dependencies between concepts. METHODS: We introduce a DirectiOnal Medical Embedding (DOME) algorithm to encode temporally directional relationships between medical concepts, using summary-level EHR data. Specifically, DOME first aggregates patient-level EHR data into an asymmetric co-occurrence matrix. Then it computes two Positive Pointwise Mutual Information (PPMI) matrices to correspondingly encode the pairwise prior and posterior dependencies between medical concepts. Following that, a joint matrix factorization is performed on the two PPMI matrices, which results in three vectors for each concept: a semantic embedding and two directional context embeddings. They collectively provide a comprehensive depiction of the temporal relationship between EHR concepts. RESULTS: We highlight the advantages and translational potential of DOME through three sets of validation studies. First, DOME consistently improves existing direction-agnostic embedding vectors for disease risk prediction in several diseases, for example achieving a relative gain of 5.5% in the area under the receiver operating characteristic (AUROC) for lung cancer. Second, DOME excels in directional drug-disease relationship inference by successfully differentiating between drug side effects and indications, correspondingly achieving relative AUROC gain over the state-of-the-art methods by 10.8% and 6.6%. Finally, DOME effectively constructs directional knowledge graphs, which distinguish disease risk factors from comorbidities, thereby revealing disease progression trajectories. The source codes are provided at https://github.com/celehs/Directional-EHR-embedding. Jun Wen 0001, Hao Xue 0005, Everett Neil Rush, Vidul Ayakulangara Panickan, Tianrun A. Cai, Doudou Zhou, Yuk-Lam Ho, Lauren Costa, Edmon Begoli, Chuan Hong, John Michael Gaziano, Kelly Cho, Katherine P. Liao, Tianxi Cai |
J. Biomed. Informatics | 11 |
| 2024 | Centralized Interactive Phenomics Resource: an integrated online phenomics knowledgebase for health data usersabstractOBJECTIVE: Development of clinical phenotypes from electronic health records (EHRs) can be resource intensive. Several phenotype libraries have been created to facilitate reuse of definitions. However, these platforms vary in target audience and utility. We describe the development of the Centralized Interactive Phenomics Resource (CIPHER) knowledgebase, a comprehensive public-facing phenotype library, which aims to facilitate clinical and health services research. MATERIALS AND METHODS: The platform was designed to collect and catalog EHR-based computable phenotype algorithms from any healthcare system, scale metadata management, facilitate phenotype discovery, and allow for integration of tools and user workflows. Phenomics experts were engaged in the development and testing of the site. RESULTS: The knowledgebase stores phenotype metadata using the CIPHER standard, and definitions are accessible through complex searching. Phenotypes are contributed to the knowledgebase via webform, allowing metadata validation. Data visualization tools linking to the knowledgebase enhance user interaction with content and accelerate phenotype development. DISCUSSION: The CIPHER knowledgebase was developed in the largest healthcare system in the United States and piloted with external partners. The design of the CIPHER website supports a variety of front-end tools and features to facilitate phenotype development and reuse. Health data users are encouraged to contribute their algorithms to the knowledgebase for wider dissemination to the research community, and to use the platform as a springboard for phenotyping. CONCLUSION: CIPHER is a public resource for all health data users available at https://phenomics.va.ornl.gov/ which facilitates phenotype reuse, development, and dissemination of phenotyping knowledge. Jacqueline Honerlaw, Yuk-Lam Ho, Francesca Fontin, Michael Murray, Ashley Galloway, David Heise, Keith Connatser, Laura Davies, Jeffrey Gosian, Monika Maripuri, John P. Russo, Rahul Sangar, Vidisha Tanukonda, Edward Zielinski, Maureen Dubreuil, Andrew J. Zimolzak, Vidul Ayakulangara Panickan, Su-Chun Cheng, Stacey B. Whitbourne, David R. Gagnon, Tianxi Cai, Katherine P. Liao, Rachel Badovinac Ramoni, John Michael Gaziano, Sumitra Muralidhar, Kelly Cho |
J. Am. Medical Informatics Assoc. | 24 |
| 2023 | Multimodal representation learning for predicting molecule-disease relationsabstractMOTIVATION: Predicting molecule-disease indications and side effects is important for drug development and pharmacovigilance. Comprehensively mining molecule-molecule, molecule-disease and disease-disease semantic dependencies can potentially improve prediction performance. METHODS: We introduce a Multi-Modal REpresentation Mapping Approach to Predicting molecular-disease relations (M2REMAP) by incorporating clinical semantics learned from electronic health records (EHR) of 12.6 million patients. Specifically, M2REMAP first learns a multimodal molecule representation that synthesizes chemical property and clinical semantic information by mapping molecule chemicals via a deep neural network onto the clinical semantic embedding space shared by drugs, diseases and other common clinical concepts. To infer molecule-disease relations, M2REMAP combines multimodal molecule representation and disease semantic embedding to jointly infer indications and side effects. RESULTS: We extensively evaluate M2REMAP on molecule indications, side effects and interactions. Results show that incorporating EHR embeddings improves performance significantly, for example, attaining an improvement over the baseline models by 23.6% in PRC-AUC on indications and 23.9% on side effects. Further, M2REMAP overcomes the limitation of existing methods and effectively predicts drugs for novel diseases and emerging pathogens. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/celehs/M2REMAP, and prediction results are provided at https://shiny.parse-health.org/drugs-diseases-dev/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jun Wen 0001, Xiang Zhang 0012, Everett Neil Rush, Vidul Ayakulangara Panickan, Tianrun A. Cai, Doudou Zhou, Yuk-Lam Ho, Lauren Costa, Edmon Begoli, Chuan Hong, John Michael Gaziano, Kelly Cho, Katherine P. Liao, Marinka Zitnik, Tianxi Cai |
Bioinform. | 12 |
| 2023 | Framework of the Centralized Interactive Phenomics Resource (CIPHER) standard for electronic health data-based phenomics knowledgebaseabstractThe development of phenotypes using electronic health records is a resource-intensive process. Therefore, the cataloging of phenotype algorithm metadata for reuse is critical to accelerate clinical research. The Department of Veterans Affairs (VA) has developed a standard for phenotype metadata collection which is currently used in the VA phenomics knowledgebase library, CIPHER (Centralized Interactive Phenomics Resource), to capture over 5000 phenotypes. The CIPHER standard improves upon existing phenotype library metadata collection by capturing the context of algorithm development, phenotyping method used, and approach to validation. While the standard was iteratively developed with VA phenomics experts, it is applicable to the capture of phenotypes across healthcare systems. We describe the framework of the CIPHER standard for phenotype metadata collection, the rationale for its development, and its current application to the largest healthcare system in the United States. Jacqueline Honerlaw, Yuk-Lam Ho, Francesca Fontin, Jeffrey Gosian, Monika Maripuri, Michael Murray, Rahul Sangar, Ashley Galloway, Andrew J. Zimolzak, Stacey B. Whitbourne, Juan P. Casas, Rachel Badovinac Ramoni, David R. Gagnon, Tianxi Cai, Katherine P. Liao, John Michael Gaziano, Sumitra Muralidhar, Kelly Cho |
J. Am. Medical Informatics Assoc. | 16 |
| 2022 | Knowledge-Driven Online Multimodal Automated Phenotyping System
Molei Liu, Sara Morini Sweet, Xin Xiong 0006, Chuan Hong, Clara-Lea Bonzel, Vidul Ayakulangara Panickan, Everett Neil Rush, Yuk-Lam Ho, Kelly Cho, John Michael Gaziano, Katherine P. Liao, Tianxi Cai, Tianrun A. Cai |
AMIA | 10 |
| 2022 | Scalable relevance ranking algorithm via semantic similarity assessment improves efficiency of medical chart reviewabstractAccurately assigning phenotype information to individual patients via computational phenotyping using Electronic Health Records (EHRs) has been seen as the first step towards enabling EHRs for precision medicine research. Chart review labels annotated by clinical experts, also known as “gold standard” labels, are essential for the development and validation of computational phenotyping algorithms. However, given the complexity of EHR systems, the process of chart review is both labor intensive and time consuming. We propose a fully automated algorithm, referred to as pGUESS, to rank EHR notes according to their relevance to a given phenotype. By identifying the most relevant notes, pGUESS can greatly improve the efficiency and accuracy of chart reviews. pGUESS uses prior guided semantic similarity to measure the informativeness of a clinical note to a given phenotype. We first select candidate clinical concepts from a pool of comprehensive medical concepts using public knowledge sources and then derive the semantic embedding vector (SEV) for a reference article (SEVref) and each note (SEVnote). The algorithm scores the relevance of a note as the cosine similarity between SEVnote and SEVref. The algorithm was validated against four sets of 200 notes that were manually annotated by clinical experts to assess their informativeness to one of three disease phenotypes. pGUESS algorithm substantially outperforms existing unsupervised approaches for classifying the relevance status with respect to both accuracy and scalability across phenotypes. Averaging over the three phenotypes, the rank correlation between the algorithm ranking and gold standard label was 0.64 for pGUESS, but only 0.47 and 0.35 for the next two best performing algorithms. pGUESS is also much more computationally scalable compared to existing algorithms. pGUESS algorithm can substantially reduce the burden of chart review and holds potential in improving the efficiency and accuracy of human annotation. Tianrun A. Cai, Zeling He, Chuan Hong, Yuk-Lam Ho, Jacqueline Honerlaw, Alon Geva, Vidul Ayakulangara Panickan, Amanda King, David R. Gagnon, John Michael Gaziano, Kelly Cho, Katherine P. Liao, Tianxi Cai |
J. Biomed. Informatics | 11 |
| 2022 | Multiview Incomplete Knowledge Graph Integration with application to cross-institutional EHR data harmonizationabstractOBJECTIVE: The growing availability of electronic health records (EHR) data opens opportunities for integrative analysis of multi-institutional EHR to produce generalizable knowledge. A key barrier to such integrative analyses is the lack of semantic interoperability across different institutions due to coding differences. We propose a Multiview Incomplete Knowledge Graph Integration (MIKGI) algorithm to integrate information from multiple sources with partially overlapping EHR concept codes to enable translations between healthcare systems. METHODS: The MIKGI algorithm combines knowledge graph information from (i) embeddings trained from the co-occurrence patterns of medical codes within each EHR system and (ii) semantic embeddings of the textual strings of all medical codes obtained from the Self-Aligning Pretrained BERT (SAPBERT) algorithm. Due to the heterogeneity in the coding across healthcare systems, each EHR source provides partial coverage of the available codes. MIKGI synthesizes the incomplete knowledge graphs derived from these multi-source embeddings by minimizing a spherical loss function that combines the pairwise directional similarities of embeddings computed from all available sources. MIKGI outputs harmonized semantic embedding vectors for all EHR codes, which improves the quality of the embeddings and enables direct assessment of both similarity and relatedness between any pair of codes from multiple healthcare systems. RESULTS: With EHR co-occurrence data from Veteran Affairs (VA) healthcare and Mass General Brigham (MGB), MIKGI algorithm produces high quality embeddings for a variety of downstream tasks including detecting known similar or related entity pairs and mapping VA local codes to the relevant EHR codes used at MGB. Based on the cosine similarity of the MIKGI trained embeddings, the AUC was 0.918 for detecting similar entity pairs and 0.809 for detecting related pairs. For cross-institutional medical code mapping, the top 1 and top 5 accuracy were 91.0% and 97.5% when mapping medication codes at VA to RxNorm medication codes at MGB; 59.1% and 75.8% when mapping VA local laboratory codes to LOINC hierarchy. When trained with 500 labels, the lab code mapping attained top 1 and 5 accuracy at 77.7% and 87.9%. MIKGI also attained best performance in selecting VA local lab codes for desired laboratory tests and COVID-19 related features for COVID EHR studies. Compared to existing methods, MIKGI attained the most robust performance with accuracy the highest or near the highest across all tasks. CONCLUSIONS: The proposed MIKGI algorithm can effectively integrate incomplete summary data from biomedical text and EHR data to generate harmonized embeddings for EHR codes for knowledge graph modeling and cross-institutional translation of EHR codes. Doudou Zhou, Ziming Gan, Alina Patwari, Everett Neil Rush, Clara-Lea Bonzel, Vidul Ayakulangara Panickan, Chuan Hong, Yuk-Lam Ho, Tianrun A. Cai, Lauren Costa, Victor M. Castro, Shawn N. Murphy, Gabriel A. Brat, Griffin M. Weber, Paul Avillach, John Michael Gaziano, Kelly Cho, Katherine P. Liao, Tianxi Cai |
J. Biomed. Informatics | 18 |
| 2019 | High-throughput multimodal automated phenotyping (MAP) with application to PheWASabstractOBJECTIVE: Electronic health records linked with biorepositories are a powerful platform for translational studies. A major bottleneck exists in the ability to phenotype patients accurately and efficiently. The objective of this study was to develop an automated high-throughput phenotyping method integrating International Classification of Diseases (ICD) codes and narrative data extracted using natural language processing (NLP). MATERIALS AND METHODS: We developed a mapping method for automatically identifying relevant ICD and NLP concepts for a specific phenotype leveraging the Unified Medical Language System. Along with health care utilization, aggregated ICD and NLP counts were jointly analyzed by fitting an ensemble of latent mixture models. The multimodal automated phenotyping (MAP) algorithm yields a predicted probability of phenotype for each patient and a threshold for classifying participants with phenotype yes/no. The algorithm was validated using labeled data for 16 phenotypes from a biorepository and further tested in an independent cohort phenome-wide association studies (PheWAS) for 2 single nucleotide polymorphisms with known associations. RESULTS: The MAP algorithm achieved higher or similar AUC and F-scores compared to the ICD code across all 16 phenotypes. The features assembled via the automated approach had comparable accuracy to those assembled via manual curation (AUCMAP 0.943, AUCmanual 0.941). The PheWAS results suggest that the MAP approach detected previously validated associations with higher power when compared to the standard PheWAS method based on ICD codes. CONCLUSION: The MAP approach increased the accuracy of phenotype definition while maintaining scalability, thereby facilitating use in studies requiring large-scale phenotyping, such as PheWAS. Katherine P. Liao, Jiehuan Sun, Tianrun A. Cai, Nicholas B. Link, Chuan Hong, Jie Huang 0030, Jennifer E. Huffman, Jessica L. Gronsbell, Yuk-Lam Ho, Victor M. Castro, Vivian S. Gainer, Shawn N. Murphy, Christopher J. O'Donnell, John Michael Gaziano, Kelly Cho, Peter Szolovits, Isaac S. Kohane, Sheng Yu 0002 |
J. Am. Medical Informatics Assoc. | 15 |
| 2018 | High-Throughput Multimodal Automated Phenotyping (MAP) Incorporating Natural Language Processing with Application to PheWAS
Katherine P. Liao, Jiehuan Sun, Tianrun A. Cai, Nicholas B. Link, Chuan Hong, Jie Huang 0030, Jennifer E. Huffman, Jessica L. Gronsbell, Lauren Costa, Victor M. Castro, Vivian S. Gainer, Shawn N. Murphy, John Michael Gaziano, Kelly Cho, Peter Szolovits, Isaac S. Kohane, Sheng Yu 0002, Tianxi Cai |
AMIA | 13 |
| 2018 | Yield and bias in defining a cohort study baseline from electronic health record data
Jason Vassy, Yuk-Lam Ho, Jacqueline Honerlaw, Kelly Cho, John Michael Gaziano, Peter W. F. Wilson, David R. Gagnon |
J. Biomed. Informatics | 5 |
| 2017 | Improving EHR Chart Review Efficiency via Semantic Similarity Assessment
Tianrun A. Cai, Andrew L. Beam, Stephanie Chan 0002, Jacqueline Honerlaw, David R. Gagnon, Kelly Cho, John Michael Gaziano, Katherine P. Liao, Tianxi Cai |
AMIA | 8 |
| 2004 | An application of conditional logistic regression and multifactor dimensionality reduction for detecting gene-gene Interactions on risk of myocardial infarction: The importance of model validationabstractBACKGROUND: To examine interactions among the angiotensin converting enzyme (ACE) insertion/deletion, plasminogen activator inhibitor-1 (PAI-1) 4G/5G, and tissue plasminogen activator (t-PA) insertion/deletion gene polymorphisms on risk of myocardial infarction using data from 343 matched case-control pairs from the Physicians Health Study. We examined the data using both conditional logistic regression and the multifactor dimensionality reduction (MDR) method. One advantage of the MDR method is that it provides an internal prediction error for validation. We summarize our use of this internal prediction error for model validation. RESULTS: The overall results for the two methods were consistent, with both suggesting an interaction between the ACE I/D and PAI-1 4G/5G polymorphisms. However, using ten-fold cross validation, the 46% prediction error for the final MDR model was not significantly lower than that expected by chance. CONCLUSIONS: The significant interaction initially observed does not validate and may represent a type I error. As data-driven analytic methods continue to be developed and used to examine complex genetic interactions, it will become increasingly important to stress model validation in order to ensure that significant effects represent true relationships rather than chance findings. Christopher S. Coffey, Patricia R. Hebert, Marylyn D. Ritchie, Harlan M. Krumholz, John Michael Gaziano, Paul M. Ridker, Nancy J. Brown, Douglas E. Vaughan, Jason H. Moore |
BMC Bioinform. | 5 |