Judong Shen

dblp:26/6418 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
5since 2021 · last 2024
0000-0001-6150-1034ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2024 Applying polygenic risk score methods to pharmacogenomics GWAS: challenges and opportunities
abstract
Polygenic risk scores (PRSs) have emerged as promising tools for the prediction of human diseases and complex traits in disease genome-wide association studies (GWAS). Applying PRSs to pharmacogenomics (PGx) studies has begun to show great potential for improving patient stratification and drug response prediction. However, there are unique challenges that arise when applying PRSs to PGx GWAS beyond those typically encountered in disease GWAS (e.g. Eurocentric or trans-ethnic bias). These challenges include: (i) the lack of knowledge about whether PGx or disease GWAS/variants should be used in the base cohort (BC); (ii) the small sample sizes in PGx GWAS with corresponding low power and (iii) the more complex PRS statistical modeling required for handling both prognostic and predictive effects simultaneously. To gain insights in this landscape about the general trends, challenges and possible solutions, we first conduct a systematic review of both PRS applications and PRS method development in PGx GWAS. To further address the challenges, we propose (i) a novel PRS application strategy by leveraging both PGx and disease GWAS summary statistics in the BC for PRS construction and (ii) a new Bayesian method (PRS-PGx-Bayesx) to reduce Eurocentric or cross-population PRS prediction bias. Extensive simulations are conducted to demonstrate their advantages over existing PRS methods applied in PGx GWAS. Our systematic review and methodology research work not only highlights current gaps and key considerations while applying PRS methods to PGx GWAS, but also provides possible solutions for better PGx PRS applications and future research.
Song Zhai, Devan V. Mehrotra, Judong Shen
Briefings Bioinform.3
2023 Integrating multiple traits for improving polygenic risk prediction in disease and pharmacogenomics GWAS
abstract
Polygenic risk score (PRS) has been recently developed for predicting complex traits and drug responses. It remains unknown whether multi-trait PRS (mtPRS) methods, by integrating information from multiple genetically correlated traits, can improve prediction accuracy and power for PRS analysis compared with single-trait PRS (stPRS) methods. In this paper, we first review commonly used mtPRS methods and find that they do not directly model the underlying genetic correlations among traits, which has been shown to be useful in guiding multi-trait association analysis in the literature. To overcome this limitation, we propose a mtPRS-PCA method to combine PRSs from multiple traits with weights obtained from performing principal component analysis (PCA) on the genetic correlation matrix. To accommodate various genetic architectures covering different effect directions, signal sparseness and across-trait correlation structures, we further propose an omnibus mtPRS method (mtPRS-O) by combining P values from mtPRS-PCA, mtPRS-ML (mtPRS based on machine learning) and stPRSs using Cauchy Combination Test. Our extensive simulation studies show that mtPRS-PCA outperforms other mtPRS methods in both disease and pharmacogenomics (PGx) genome-wide association studies (GWAS) contexts when traits are similarly correlated, with dense signal effects and in similar effect directions, and mtPRS-O is consistently superior to most other methods due to its robustness under various genetic architectures. We further apply mtPRS-PCA, mtPRS-O and other methods to PGx GWAS data from a randomized clinical trial in the cardiovascular domain and demonstrate performance improvement of mtPRS-PCA in both prediction accuracy and patient stratification as well as the robustness of mtPRS-O in PRS association test.
Song Zhai, Bin Guo 0004, Baolin Wu, Devan V. Mehrotra, Judong Shen
Briefings Bioinform.5
2023 A fast and powerful linear mixed model approach for genotype-environment interaction tests in large-scale GWAS
abstract
Genotype-by-environment interaction (GEI or GxE) plays an important role in understanding complex human traits. However, it is usually challenging to detect GEI signals efficiently and accurately while adjusting for population stratification and sample relatedness in large-scale genome-wide association studies (GWAS). Here we propose a fast and powerful linear mixed model-based approach, fastGWA-GE, to test for GEI effect and G + GxE joint effect. Our extensive simulations show that fastGWA-GE outperforms other existing GEI test methods by controlling genomic inflation better, providing larger power and running hundreds to thousands of times faster. We performed a fastGWA-GE analysis of ~7.27 million variants on 452 249 individuals of European ancestry for 13 quantitative traits and five environment variables in the UK Biobank GWAS data and identified 96 significant signals (72 variants across 57 loci) with GEI test P-values < 1 × 10-9, including 27 novel GEI associations, which highlights the effectiveness of fastGWA-GE in GEI signal discovery in large-scale GWAS.
Wujuan Zhong, Aparna Chhibber, Devan V. Mehrotra, Judong Shen
Briefings Bioinform.5
2023 Robust genetic model-based SNP-set association test using CauchyGM
abstract
MOTIVATION: Association testing on genome-wide association studies (GWAS) data is commonly performed under a single (mostly additive) genetic model framework. However, the underlying true genetic mechanisms are often unknown in practice for most complex traits. When the employed inheritance model deviates from the underlying model, statistical power may be reduced. To overcome this challenge, an integrative association test that directly infers the underlying genetic model from GWAS data has previously been proposed for single-SNP analysis. RESULTS: In this article, we propose a Cauchy combination Genetic Model-based association test (CauchyGM) under a generalized linear model framework for SNP-set level analysis. CauchyGM does not require prior knowledge on the underlying inheritance pattern of each SNP. It performs a score test that first estimates an individual P-value of each SNP in an SNP-set with both minor allele frequency (MAF) > 1% and three genotypes and further aggregates the rest SNPs using SKAT. CauchyGM then combines the correlated P-values across multiple SNPs and different genetic models within the set using Cauchy Combination Test. To further accommodate both sparse and dense signal patterns, we also propose an omnibus association test (CauchyGM-O) by combining CauchyGM with SKAT and the burden test. Our extensive simulations show that both CauchyGM and CauchyGM-O maintain the type I error well at the genome-wide significance level and provide substantial power improvement compared to existing methods. We apply our methods to a pharmacogenomic GWAS data from a large cardiovascular randomized clinical trial. Both CauchyGM and CauchyGM-O identify several novel genome-wide significant genes. AVAILABILITY AND IMPLEMENTATION: The R package CauchyGM is publicly available on github: https://github.com/ykim03517/CauchyGM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yeonil Kim, Yueh-Yun Chi, Judong Shen
Bioinform.3
2023 AWOT and CWOT for genotype and genotype-by-treatment interaction joint analysis in pharmacogenetics GWAS
abstract
MOTIVATION: Pharmacogenomics (PGx) research holds the promise for detecting association between genetic variants and drug responses in randomized clinical trials, but it is limited by small populations and thus has low power to detect signals. It is critical to increase the power of PGx genome-wide association studies (GWAS) with small sample sizes so that variant-drug-response association discoveries are not limited to common variants with extremely large effect. RESULTS: In this article, we first discuss the challenges of PGx GWAS studies and then propose the adaptively weighted joint test (AWOT) and Cauchy Weighted jOint Test (CWOT), which are two flexible and robust joint tests of the single nucleotide polymorphism main effect and genotype-by-treatment interaction effect for continuous and binary endpoints. Two analytic procedures are proposed to accurately calculate the joint test P-value. We evaluate AWOT and CWOT through extensive simulations under various scenarios. The results show that the proposed AWOT and CWOT control type I error well and outperform existing methods in detecting the most interesting signal patterns in PGx settings (i.e. with strong genotype-by-treatment interaction effects, but weak genotype main effects). We demonstrate the value of AWOT and CWOT by applying them to the PGx GWAS from the Bezlotoxumab Clostridium difficile MODIFY I/II Phase 3 trials. AVAILABILITY AND IMPLEMENTATION: The R package COWT is publicly available on CRAN https://cran.r-project.org/web/packages/cwot/index.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hong Zhang 0038, Devan V. Mehrotra, Judong Shen
Bioinform.3
2020 Composite Kernel Association Test (CKAT) for SNP-set joint assessment of genotype and genotype-by-treatment interaction in Pharmacogenetics studies
abstract
MOTIVATION: It is of substantial interest to discover novel genetic markers that influence drug response in order to develop personalized treatment strategies that maximize therapeutic efficacy and safety. To help enable such discoveries, we focus on testing the association between the cumulative effect of multiple single nucleotide polymorphisms (SNPs) in a particular genomic region and a drug response of interest. However, the currently existing methods are either computational inefficient or not able to control type I error and provide decent power for whole exome or genome analysis in Pharmacogenetics (PGx) studies with small sample sizes. RESULTS: In this article, we propose the Composite Kernel Association Test (CKAT), a flexible and robust kernel machine-based approach to jointly test the genetic main effect and SNP-treatment interaction effect for SNP-sets in Pharmacogenetics (PGx) assessments embedded within randomized clinical trials. An analytic procedure is developed to accurately calculate the P-value so that computationally extensive procedures (e.g. permutation or perturbation) can be avoided. We evaluate CKAT through extensive simulation studies and application to the gene-level association test of the reduction in Clostridium difficile infection recurrence in patients treated with bezlotoxumab. The results demonstrate that the proposed CKAT controls type I error well for PGx studies, is efficient for whole exome/genome association analysis and provides better power performance than existing methods across multiple scenarios. AVAILABILITY AND IMPLEMENTATION: The R package CKAT is publicly available on CRAN https://cran.r-project.org/web/packages/CKAT/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hong Zhang 0038, Ni Zhao, Devan V. Mehrotra, Judong Shen
Bioinform.4
2017 STOPGAP: a database for systematic target opportunity assessment by genetic association predictions
abstract
SUMMARY: We developed the STOPGAP (Systematic Target OPportunity assessment by Genetic Association Predictions) database, an extensive catalog of human genetic associations mapped to effector gene candidates. STOPGAP draws on a variety of publicly available GWAS associations, linkage disequilibrium (LD) measures, functional genomic and variant annotation sources. Algorithms were developed to merge the association data, partition associations into non-overlapping LD clusters, map variants to genes and produce a variant-to-gene score used to rank the relative confidence among potential effector genes. This database can be used for a multitude of investigations into the genes and genetic mechanisms underlying inter-individual variation in human traits, as well as supporting drug discovery applications. AVAILABILITY AND IMPLEMENTATION: Shell, R, Perl and Python scripts and STOPGAP R data files (version 2.5.1 at publication) are available at https://github.com/StatGenPRD/STOPGAP . Some of the most useful STOPGAP fields can be queried through an R Shiny web application at http://stopgapwebapp.com . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Judong Shen, Kijoung Song, Andrew J. Slater, Enrico Ferrero, Matthew R. Nelson
Bioinform.1
2013 Multi-Population Classical HLA Type Imputation
abstract
Statistical imputation of classical HLA alleles in case-control studies has become established as a valuable tool for identifying and fine-mapping signals of disease association in the MHC. Imputation into diverse populations has, however, remained challenging, mainly because of the additional haplotypic heterogeneity introduced by combining reference panels of different sources. We present an HLA type imputation model, HLA*IMP:02, designed to operate on a multi-population reference panel. HLA*IMP:02 is based on a graphical representation of haplotype structure. We present a probabilistic algorithm to build such models for the HLA region, accommodating genotyping error, haplotypic heterogeneity and the need for maximum accuracy at the HLA loci, generalizing the work of Browning and Browning (2007) and Ron et al. (1998). HLA*IMP:02 achieves an average 4-digit imputation accuracy on diverse European panels of 97% (call rate 97%). On non-European samples, 2-digit performance is over 90% for most loci and ethnicities where data available. HLA*IMP:02 supports imputation of HLA-DPB1 and HLA-DRB3-5, is highly tolerant of missing data in the imputation panel and works on standard genotype data from popular genotyping chips. It is publicly available in source code and as a user-friendly web service framework.
Alexander T. Dilthey, Stephen Leslie, Loukas Moutsianas, Judong Shen, Charles Cox, Matthew R. Nelson, Gil McVean
PLoS Comput. Biol.4