Eric R. Gamazon

dblp:27/7970 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0003-4204-8734ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Cross-ancestry information transfer framework improves protein abundance prediction and protein-trait association identification
abstract
Genetics-informed proteome-wide association studies (PWASs) provide an effective way to uncover proteomic mechanisms underlying complex diseases. PWAS relies on an ancestry-matched reference panel to model the impact of genetically determined protein expression on phenotype. However, reference panels from underrepresented populations remain relatively limited. We developed a multi-ancestry framework to enhance protein prediction in these populations by integrating diverse information-sharing strategies into a Multi-Ancestry Best-performing Model (MABM). Results indicated that MABM increased the prediction performance with higher performance observed in both cross-validation and an external dataset. Leveraging the Biobank Japan, we identified three times as many significant PWAS associations using MABM as using Lasso model. Notably, 47.5% of the MABM specific associations were reproduced in independent East Asian datasets with concordant effect sizes. Furthermore, MABM enhanced decision-making in gene/protein prioritization for functional validation for complex traits by validating well-established associations and uncovering novel trait-related candidates. The benefits of MABM were further validated in additional ancestries and demonstrated in brain tissue-based PWAS, underscoring its broad applicability. Our findings close critical gaps in multi-omics research among underrepresented populations and facilitate trait-relevant protein discovery in underrepresented populations.
Wenli Zhai, Lingyun Sun, Wenwei Fang, Yidan Dong, Chunxiao Cheng, Yuanjiao Liu, Jiadong Ji, An Pan, Eric R. Gamazon, Xiong-Fei Pan
Briefings Bioinform.11
2025 Integrating population-level and cell-based signatures for drug repositioning
abstract
MOTIVATION: Drug repositioning presents a streamlined and cost-efficient way to expand the range of therapeutic possibilities. Drugs with human genetic evidence are more likely to advance successfully through clinical trials toward Food and Drug Administration approval. Single gene-based drug repositioning methods have been implemented, but approaches leveraging a broad spectrum of molecular signatures remain underexplored. RESULTS: We propose a framework called "Transcriptome-informed Reversal Distance" (TReD) that embeds the disease signatures and drug response profiles into a high-dimensional normed space to quantify the reversal potential of candidate drugs in a disease-related cell-based screening. We applied TReD to COVID-19, type 2 diabetes, and Alzheimer's disease (AD), identifying 36, 16, and 11 candidate drugs, respectively. Among these, literature supports 69% (25/36), 31% (5/16), and 64% (7/11) of the drugs, with clinical trials conducted for seven COVID-19 candidates and three AD candidates. In summary, we propose a comprehensive genetics-anchored framework integrating population-level signatures and cell-based screening that has the potential to accelerate the search for new therapeutic strategies. AVAILABILITY AND IMPLEMENTATION: Source code and datasets considered in this study are available at Github (https://github.com/zdangm/TReD). An archived snapshot is deposited at Zenodo (https://doi.org/10.5281/zenodo.16791909).
Chunfeng He, Jiayao Fan, Chunxiao Cheng, Ran Meng, Ruiyuan Pan, Ravi V. Shah, Eric R. Gamazon
Bioinform.10
2025 Transcriptome-wide root causal inference
abstract
Root causal genes correspond to the first gene expression levels perturbed during pathogenesis by genetic or non-genetic factors. Targeting root causal genes has the potential to alleviate disease entirely by eliminating pathology near its onset. No existing algorithm has been designed to discover root causal genes from observational data alone. We therefore propose the Transcriptome-Wide Root Causal Inference (TWRCI) algorithm that identifies root causal genes and their causal graph using a combination of genetic variant and unperturbed bulk RNA sequencing data. TWRCI uses a novel competitive regression procedure to annotate cis and trans-genetic variants to the gene expression levels they directly cause. The algorithm simultaneously determines the sequence in which gene expression changes propagate through the system to pinpoint the underlying causal graph and estimate root causal effects. TWRCI outperforms alternative approaches across a diverse group of metrics by directly targeting root causal genes while accounting for distal relations, linkage disequilibrium, patient heterogeneity and widespread pleiotropy. We demonstrate the algorithm by uncovering the root causal mechanisms of two complex diseases, which we confirm by replication using independent genome-wide summary statistics.
Eric V. Strobl, Eric R. Gamazon
PLoS Comput. Biol.2
2024 NeuroimaGene: an R package for assessing the neurological correlates of genetically regulated gene expression
abstract
BACKGROUND: We present the NeuroimaGene resource as an R package designed to assist researchers in identifying genes and neurologic features relevant to psychiatric and neurological health. While recent studies have identified hundreds of genes as potential components of pathophysiology in neurologic and psychiatric disease, interpreting the physiological consequences of this variation is challenging. The integration of neuroimaging data with molecular findings is a step toward addressing this challenge. In addition to sharing associations with both molecular variation and clinical phenotypes, neuroimaging features are intrinsically informative of cognitive processes. NeuroimaGene provides a tool to understand how disease-associated genes relate to the intermediate structure of the brain. RESULTS: We created NeuroimaGene, a user-friendly, open access R package now available for public use. Its primary function is to identify neuroimaging derived brain features that are impacted by genetically regulated expression of user-provided genes or gene sets. This resource can be used to (1) characterize individual genes or gene sets as relevant to the structure and function of the brain, (2) identify the region(s) of the brain or body in which expression of target gene(s) is neurologically relevant, (3) impute the brain features most impacted by user-defined gene sets such as those produced by cohort level gene association studies, and (4) generate publication level, modifiable visual plots of significant findings. We demonstrate the utility of the resource by identifying neurologic correlates of stroke-associated genes derived from pre-existing analyses. CONCLUSIONS: Integrating neurologic data as an intermediate phenotype in the pathway from genes to brain-based diagnostic phenotypes increases the interpretability of molecular studies and enriches our understanding of disease pathophysiology. The NeuroimaGene R package is designed to assist in this process and is publicly available for use.
Xavier Bledsoe, Eric R. Gamazon
BMC Bioinform.2
2023 IMMerge: merging imputation data at scale
abstract
SUMMARY: Genomic data are often processed in batches and analyzed together to save time. However, it is challenging to combine multiple large VCFs and properly handle imputation quality and missing variants due to the limitations of available tools. To address these concerns, we developed IMMerge, a Python-based tool that takes advantage of multiprocessing to reduce running time. For the first time in a publicly available tool, imputation quality scores are correctly combined with Fisher's z transformation. AVAILABILITY AND IMPLEMENTATION: IMMerge is an open-source project under MIT license. Source code and user manual are available at https://github.com/belowlab/IMMerge.
Wanying Zhu, Hung-Hsin Chen, Alexander S. Petty, Lauren E. Petty, Hannah G. Polikowsky, Eric R. Gamazon, Jennifer E. Below, Heather M. Highland
Bioinform.6
2021 E-MAGMA: an eQTL-informed method to identify risk genes using genome-wide association study summary statistics
abstract
MOTIVATION: Genome-wide association studies have successfully identified multiple independent genetic loci that harbour variants associated with human traits and diseases, but the exact causal genes are largely unknown. Common genetic risk variants are enriched in non-protein-coding regions of the genome and often affect gene expression (expression quantitative trait loci, eQTL) in a tissue-specific manner. To address this challenge, we developed a methodological framework, E-MAGMA, which converts genome-wide association summary statistics into gene-level statistics by assigning risk variants to their putative genes based on tissue-specific eQTL information. RESULTS: We compared E-MAGMA to three eQTL informed gene-based approaches using simulated phenotype data. Phenotypes were simulated based on eQTL reference data using GCTA for all genes with at least one eQTL at chromosome 1. We performed 10 simulations per gene. The eQTL-h2 (i.e. the proportion of variation explained by the eQTLs) was set at 1%, 2% and 5%. We found E-MAGMA outperforms other gene-based approaches across a range of simulated parameters (e.g. the number of identified causal genes). When applied to genome-wide association summary statistics for five neuropsychiatric disorders, E-MAGMA identified more putative candidate causal genes compared to other eQTL-based approaches. By integrating tissue-specific eQTL information, these results show E-MAGMA will help to identify novel candidate causal genes from genome-wide association summary statistics and thereby improve the understanding of the biological basis of complex disorders. AVAILABILITY AND IMPLEMENTATION: A tutorial and input files are made available in a github repository: https://github.com/eskederks/eMAGMA-tutorial. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zachary F. Gerring, Angela Mina-Vargas, Eric R. Gamazon, Eske M. Derks
Bioinform.3
2016 STAMS: STRING-assisted module search for genome wide association studies and application to autism
abstract
Motivation: Analyzing genome wide association data in the context of biological pathways helps us understand how genetic variation influences phenotype and increases power to find associations. However, the utility of pathway-based analysis tools is hampered by undercuration and reliance on a distribution of signal across all of the genes in a pathway. Methods that combine genome wide association results with genetic networks to infer the key phenotype-modulating subnetworks combat these issues, but have primarily been limited to network definitions with yes/no labels for gene-gene interactions. A recent method (EW_dmGWAS) incorporates a biological network with weighted edge probability by requiring a secondary phenotype-specific expression dataset. In this article, we combine an algorithm for weighted-edge module searching and a probabilistic interaction network in order to develop a method, STAMS, for recovering modules of genes with strong associations to the phenotype and probable biologic coherence. Our method builds on EW_dmGWAS but does not require a secondary expression dataset and performs better in six test cases. Results: We show that our algorithm improves over EW_dmGWAS and standard gene-based analysis by measuring precision and recall of each method on separately identified associations. In the Wellcome Trust Rheumatoid Arthritis study, STAMS-identified modules were more enriched for separately identified associations than EW_dmGWAS (STAMS P-value 3.0 × 10−4; EW_dmGWAS- P-value = 0.8). We demonstrate that the area under the Precision-Recall curve is 5.9 times higher with STAMS than EW_dmGWAS run on the Wellcome Trust Type 1 Diabetes data. Availability and Implementation: STAMS is implemented as an R package and is freely available at https://simtk.org/projects/stams. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Sara Hillenmeyer, Lea K. Davis, Eric R. Gamazon, Edwin H. Cook Jr., Nancy J. Cox, Russ B. Altman
Bioinform.3
2013 Research and applications: Network models of genome-wide association studies uncover the topological centrality of protein interactions in complex diseases
abstract
BACKGROUND: While genome-wide association studies (GWAS) of complex traits have revealed thousands of reproducible genetic associations to date, these loci collectively confer very little of the heritability of their respective diseases and, in general, have contributed little to our understanding the underlying disease biology. Physical protein interactions have been utilized to increase our understanding of human Mendelian disease loci but have yet to be fully exploited for complex traits. METHODS: We hypothesized that protein interaction modeling of GWAS findings could highlight important disease-associated loci and unveil the role of their network topology in the genetic architecture of diseases with complex inheritance. RESULTS: Network modeling of proteins associated with the intragenic single nucleotide polymorphisms of the National Human Genome Research Institute catalog of complex trait GWAS revealed that complex trait associated loci are more likely to be hub and bottleneck genes in available, albeit incomplete, networks (OR=1.59, Fisher's exact test p < 2.24 × 10(-12)). Network modeling also prioritized novel type 2 diabetes (T2D) genetic variations from the Finland-USA Investigation of Non-Insulin-Dependent Diabetes Mellitus Genetics and the Wellcome Trust GWAS data, and demonstrated the enrichment of hubs and bottlenecks in prioritized T2D GWAS genes. The potential biological relevance of the T2D hub and bottleneck genes was revealed by their increased number of first degree protein interactions with known T2D genes according to several independent sources (p<0.01, probability of being first interactors of known T2D genes). CONCLUSION: Virtually all common diseases are complex human traits, and thus the topological centrality in protein networks of complex trait genes has implications in genetics, personal genomics, and therapy.
Younghee Lee, Haiquan Li, Jianrong Li, Ellen Rebman, Ikbel Achour, Kelly Regan-Fendt, Eric R. Gamazon, James L. Chen, Xinan Yang, Nancy J. Cox, Yves A. Lussier
J. Am. Medical Informatics Assoc.7
2010 SCAN: SNP and copy number annotation
abstract
MOTIVATION: Genome-wide association studies (GWAS) generate relationships between hundreds of thousands of single nucleotide polymorphisms (SNPs) and complex phenotypes. The contribution of the traditionally overlooked copy number variations (CNVs) to complex traits is also being actively studied. To facilitate the interpretation of the data and the designing of follow-up experimental validations, we have developed a database that enables the sensible prioritization of these variants by combining several approaches, involving not only publicly available physical and functional annotations but also multilocus linkage disequilibrium (LD) annotations as well as annotations of expression quantitative trait loci (eQTLs). RESULTS: For each SNP, the SCAN database provides: (i) summary information from eQTL mapping of HapMap SNPs to gene expression (evaluated by the Affymetrix exon array) in the full set of HapMap CEU (Caucasians from UT, USA) and YRI (Yoruba people from Ibadan, Nigeria) samples; (ii) LD information, in the case of a HapMap SNP, including what genes have variation in strong LD (pairwise or multilocus LD) with the variant and how well the SNP is covered by different high-throughput platforms; (iii) summary information available from public databases (e.g. physical and functional annotations); and (iv) summary information from other GWAS. For each gene, SCAN provides annotations on: (i) eQTLs for the gene (both local and distant SNPs) and (ii) the coverage of all variants in the HapMap at that gene on each high-throughput platform. For each genomic region, SCAN provides annotations on: (i) physical and functional annotations of all SNPs, genes and known CNVs within the region and (ii) all genes regulated by the eQTLs within the region. AVAILABILITY: http://www.scandb.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Eric R. Gamazon, Wei Zhang 0253, Anuar Konkashbaev, Shiwei Duan, Emily O. Kistner, Dan L. Nicolae, M. Eileen Dolan, Nancy J. Cox
Bioinform.1