Hongyu Zhao 0003

dblp:26/6554-3 · DBLP profile ↗
← Back
91ranked-venue papers
0as first author
23since 2021 · last 2026
0000-0003-1195-9607ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 88 · 21 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021
YearPublicationVenuePosition
2026 UBD: incorporating uncertainty in cell type proportion estimates from bulk samples to infer cell-type-specific profiles
abstract
Statistical deconvolution methods offer a powerful solution for estimating cell-type-specific (CTS) profiles from readily available bulk tissue data. However, a critical limitation of existing methods is that they require the knowledge of cell type proportions of individuals in the bulk data. While the ground truth of cell type proportions in bulk samples are unknown, those methods use the estimated proportions to approximate the truth, which potentially introduces additional uncertainties in the inferred CTS profiles. To address this challenge, we propose Uncertainty-aware Bayesian Deconvolution (UBD) to incorporate uncertainty in cell type proportion estimates. By explicitly modeling the uncertainty in the initial estimates, UBD refines cell type proportions and estimates sample-level CTS data simultaneously. We show that UBD can improve the estimates of CTS profiles through extensive simulations. We further demonstrate the utility of UBD to reveal more CTS signals in its applications to two real datasets.
Youshu Cheng, Hongyu Zhao 0003
Briefings Bioinform.5
2026 A unified framework for selecting and evaluating cell-type-specific gene co-expressions in single-cell data
abstract
Cell-type-specific gene co-expression networks are widely used to characterize gene relationships. Although many methods have been developed to infer such co-expression networks from single-cell data, the lack of consideration of false positive control in many evaluations and downstream analyses may lead to incorrect conclusions because higher reproducibility, higher functional coherence, and a larger overlap with known biological networks may not imply better performance if the false positives are not well controlled. In this study, we systematically compared two distinct criteria for selecting correlated gene pairs from single-cell data, p-value versus correlation strength. We found that the use of p-values instead of correlation strength is more robust for both selecting meaningful gene pairs and for the fair benchmarking of co-expression estimation methods. To make this approach universally applicable, we extended and validated a simulation method that can efficiently and reliably generate empirical p-values for co-expression estimation methods that do not have corresponding or well-controlled p-values. Furthermore, we demonstrated that a fair comparison of the estimation methods requires adjusting for the varying number of gene pairs they identified and accounting for the inherent expression-level biases within ground truth biological networks. Our study provides a practical guide for researchers to select reliable correlated gene pairs for downstream study and establishes a more rigorous standard for the evaluation and comparison of gene co-expression network estimation methods.
Xinning Shan, Yingxin Lin, Hongyu Zhao 0003
Briefings Bioinform.3
2026 A novel two-sample Mendelian randomization framework integrating common and rare variants: application to assess the effect of HDL-C on preeclampsia risk
abstract
Mendelian randomization (MR) has become an important technique for establishing causal relationships between risk factors and health outcomes. By using genetic variants as instrumental variables, it can mitigate bias due to confounding and reverse causation in observational studies. Current MR analyses have predominantly used common genetic variants as instruments, which represent only part of the genetic architecture of complex traits. Rare variants, which can have larger effect sizes and provide unique biological insights, have been understudied due to statistical and methodological challenges. We introduce MR-common and annotation-informed rare variants (MR-CARV), a novel framework integrating common and rare genetic variants in two-sample MR. This method leverages comprehensive genetic data made available by high-throughput sequencing technologies and large-scale consortia. Rare variants are aggregated into functional categories, such as gene-coding, gene-noncoding, and nongene regions, by leveraging variant annotations and biological impact as weights. The effects of rare variant sets are then estimated with STAARpipeline and combined with the estimated effects of common variants by the existing MR methods. Simulation studies demonstrate that MR-CARV maintains robust type I error and achieves higher statistical power, with up to a 66.3% relative increase compared with existing methods only based on common variants. Consistent with these findings, application to real data on high-density lipoprotein cholesterol (HDL-C) and preeclampsia showed that MR-CARV [inverse variance weighted (IVW)] yielded a more precise and statistically significant effect estimate (-0.020, SE = 0.0102, $P$ =.0470) than IVW using only common variants (-0.023, SE = 0.0123, $P$ =.0659).
David M. Haas, C. Noel Bairey Merz, Tsegaselassie Workalemahu, Kelli Ryckman, Janet M. Catov, Lisa D. Levine, Alexa Freedman, George R. Saade, Jiaqi Hu 0004, Hongyu Zhao 0003, Xihao Li, Nianjun Liu
Briefings Bioinform.12
2025 A novel prognostic framework for HBV-infected hepatocellular carcinoma: insights from ferroptosis and iron metabolism proteomics
abstract
Effective classification methods and prognostic models enable more accurate classification and treatment of hepatocellular carcinoma (HCC) patients. However, the weak correlation between RNA and protein data has limited the clinical utility of previous RNA-based prognostic models for HCC. In this work, we constructed a novel prognostic framework for HCC patients using seven differentially expressed proteins associated with ferroptosis and iron metabolism. Furthermore, this prognostic model robustly classifies HCC patients into three clinically relevant risk groups. Significant differences in overall survival, age, tumor differentiation, microvascular invasion, distant metastasis, and alpha-fetoprotein levels were observed among the risk groups. Based on the prognostic model and known biological pathways, we explored the potential mechanisms underlying the inconsistent differential expression patterns of FTH1 (Ferritin heavy chain 1) mRNA and protein. Our findings demonstrated that tumor tissues in HCC patients promote liver cancer progression by downregulating FTH1 protein expression, rather than upregulating FTH1 mRNA expression, ultimately leading to poor prognosis. Subsequently, based on risk score and tumor size, we developed a nomogram for predicting the prognosis of HCC patients, which demonstrated superior predictive performance in both the training and validation cohorts (C-index: 0.774; AUC for 1-5 years: 0.783-0.964). Additionally, our findings demonstrated that the adverse prognosis of high-risk HCC patients was closely correlated with ferroptosis in liver cancer tissues, alterations in iron metabolism, and changes in the tumor immune microenvironment. In conclusion, our prognostic model and predictive nomogram offer novel insights and tools for the effective classification of HCC patients, potentially enhancing clinical decision-making and outcomes.
Yongyong Ren, Xinbo Wang, Yuening Zhang, Yingqi Hua, Hongyu Zhao 0003, Hui Lu 0004
Briefings Bioinform.6
2025 Incorporating local ancestry information to predict genetically associated DNA methylation in admixed populations
abstract
Methylome-wide association studies (MWASs) have identified many 5'-cytosine-phosphate-guanine-3' (CpG) sites associated with complex traits. Several methods have been developed to predict CpG methylation levels from genotypes when the direct measurements of methylation are unavailable. To date, the published methods have mostly used datasets from populations of European ancestry to train prediction models for methylations, which limits the generalizability of methylome-wide association study to non-European populations. To address this gap, we proposed a new model by incorporating local ancestry (LA) information, called LA Methylation Predictor with Preselection (LAMPP), to improve the prediction accuracy of DNA methylation in admixed populations. We showed that LAMPP outperformed the conventional model and other LA models in prediction accuracy using an admixed African American population. We further applied our model to identify significant CpG sites for seven complex traits. Together, our LAMPP model is a valuable tool to reveal epigenetic underpinnings of complex traits in the admixed populations.
Youshu Cheng, Geyu Zhou, Amy Justice, Claudia Martinez-Araneda, Bradley E. Aouizerat, Hongyu Zhao 0003
Briefings Bioinform.9
2025 CosGeneGate selects multi-functional and credible biomarkers for single-cell analysis
abstract
MOTIVATION: Selecting representative genes or marker genes to distinguish cell types is an important task in single-cell sequencing analysis. Although many methods have been proposed to select marker genes, the genes selected may have redundancy and/or do not show cell-type-specific expression patterns to distinguish cell types. RESULTS: Here, we present a novel model, named CosGeneGate, to select marker genes for more effective marker selections. CosGeneGate is inspired by combining the advantages of selecting marker genes based on both cell-type classification accuracy and marker gene specific expression patterns. We demonstrate the better performance of the marker genes selected by CosGeneGate for various downstream analyses than the existing methods with both public datasets and newly sequenced datasets. The non-redundant marker genes identified by CosGeneGate for major cell types and tissues in human can be found at the website as follows: https://github.com/VivLon/CosGeneGate/blob/main/marker gene list.xlsx.
Tianyu Liu 0005, Wenxin Long, Yuge Wang, Chuan Hua He, Le Zhang 0013, Stephen M. Strittmatter, Hongyu Zhao 0003
Briefings Bioinform.8
2025 Robust pleiotropy-decomposed polygenic scores identify distinct contributions to elevated coronary artery disease polygenic risk
abstract
BACKGROUND: Polygenic risk score (PRS) have proved to offer robust risk prediction for coronary artery disease (CAD). However, the global CAD PRS summarizes the joint effects of all the markers in the genome, masking potential genetic heterogeneity that may be important for disease interpretation and targeted interventions. METHODS: Using summary-level data, we identified 43 significant CAD-related traits based on genetic correlations, and further classified them into eight pleiotropy clusters based on their biological functions. We then partitioned the genome into 2,353 near-independent regions. Variants in each region were assigned to the trait most genetically similar to CAD, and then were labeled with the corresponding pleiotropy cluster. We grouped variants without labels into a ninth, non-specific cluster. The Pleiotropy Decomposed (PD) PRSs for each of the nine clusters were calculated using variants assigned to each cluster for 407,903 samples of European ancestry from the UK Biobank (UKBB). RESULTS: We decomposed the CAD PRS into nine PD-PRSs and further stratified individuals with high CAD-PRS into nine subgroups. Each PD-PRS accounted for a higher proportion of the global CAD-PRS within its corresponding subgroup than in the remaining subjects with high CAD-PRS (e.g., 25.2% (0.07) vs. 10.06% (0.07) for lipids-PD-PRS). Additionally, these subgroups showed distinct clinical features. For example, in the lipids-related subgroup, lipoprotein(a) and LDL-cholesterol levels were 67.5% and 18.3% higher, respectively, compared to the remaining high-risk individuals. Furthermore, significant interactions were observed between blood pressure and BP PD-PRS, and between current smoking and respiratory system PD-PRS. CONCLUSION: Our findings suggest that PD-PRSs may reveal substantial genetic and phenotypic heterogeneity among individuals with high CAD-PRS. The unique PD-PRS compositions of each individual can highlight the relative importance of different pleiotropic regions.
Jiaqi Hu 0004, Yixuan Ye, Yunfeng Ruan, Pradeep Natarajan, Hongyu Zhao 0003
PLoS Comput. Biol.6
2024 Semi-supervised Knowledge Transfer Across Multi-omic Single-cell Data
abstract
Knowledge transfer between multi-omic single-cell data aims to effectively transfer cell types from scRNA-seq data to unannotated scATAC-seq data. Several approaches aim to reduce the heterogeneity of multi-omic data while maintaining the discriminability of cell types with extensive annotated data. However, in reality, the cost of collecting both a large amount of labeled scRNA-seq data and scATAC-seq data is expensive. Therefore, this paper explores a practical yet underexplored problem of knowledge transfer across multi-omic single-cell data under cell type scarcity. To address this problem, we propose a semi-supervised knowledge transfer framework named Dual label scArcity elimiNation with Cross-omic multi-samplE Mixup (DANCE). To overcome the label scarcity in scRNA-seq data, we generate pseudo-labels based on optimal transport and merge them into the labeled scRNA-seq data. Moreover, we adopt a divide-and-conquer strategy which divides the scATAC-seq data into source-like and target-specific data. For source-like samples, we employ consistency regularization with random perturbations while for target-specific samples, we select a few candidate labels and progressively eliminate incorrect cell types from the label set for additional supervision. Next, we generate virtual scRNA-seq samples with multi-sample Mixup based on the class-wise similarity to reduce cell heterogeneity. Extensive experiments on many benchmark datasets suggest the superiority of our DANCE over a series of state-of-the-art methods.
Fan Zhang 0111, Tianyu Liu 0005, Xiaojiang Peng, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001, Hongyu Zhao 0003
NeurIPS8
2024 LDER-GE estimates phenotypic variance component of gene-environment interactions in human complex traits accurately with GE interaction summary statistics and full LD information
abstract
Gene-environment (GE) interactions are essential in understanding human complex traits. Identifying these interactions is necessary for deciphering the biological basis of such traits. In this study, we review state-of-art methods for estimating the proportion of phenotypic variance explained by genome-wide GE interactions and introduce a novel statistical method Linkage-Disequilibrium Eigenvalue Regression for Gene-Environment interactions (LDER-GE). LDER-GE improves the accuracy of estimating the phenotypic variance component explained by genome-wide GE interactions using large-scale biobank association summary statistics. LDER-GE leverages the complete Linkage Disequilibrium (LD) matrix, as opposed to only the diagonal squared LD matrix utilized by LDSC (Linkage Disequilibrium Score)-based methods. Our extensive simulation studies demonstrate that LDER-GE performs better than LDSC-based approaches by enhancing statistical efficiency by ~23%. This improvement is equivalent to a sample size increase of around 51%. Additionally, LDER-GE effectively controls type-I error rate and produces unbiased results. We conducted an analysis using UK Biobank data, comprising 307 259 unrelated European-Ancestry subjects and 966 766 variants, across 217 environmental covariate-phenotype (E-Y) pairs. LDER-GE identified 34 significant E-Y pairs while LDSC-based method only identified 23 significant E-Y pairs with 22 overlapped with LDER-GE. Furthermore, we employed LDER-GE to estimate the aggregated variance component attributed to multiple GE interactions, leading to an increase in the explained phenotypic variance with GE interactions compared to considering main genetic effects only. Our results suggest the importance of impacts of GE interactions on human complex traits.
Zihan Dong, Wei Jiang 0019, Andrew T. Dewan, Hongyu Zhao 0003
Briefings Bioinform.5
2023 MuSe-GNN: Learning Unified Gene Representation From Multimodal Biological Graph Data
abstract
Discovering genes with similar functions across diverse biomedical contexts poses a significant challenge in gene representation learning due to data heterogeneity. In this study, we resolve this problem by introducing a novel model called Multimodal Similarity Learning Graph Neural Network, which combines Multimodal Machine Learning and Deep Graph Neural Networks to learn gene representations from single-cell sequencing and spatial transcriptomic data. Leveraging 82 training datasets from 10 tissues, three sequencing techniques, and three species, we create informative graph structures for model training and gene representations generation, while incorporating regularization with weighted similarity learning and contrastive learning to learn cross-data gene-gene relationships. This novel design ensures that we can offer gene representations containing functional similarity across different contexts in a joint space. Comprehensive benchmarking analysis shows our model's capacity to effectively capture gene function similarity across multiple modalities, outperforming state-of-the-art methods in gene representation learning by up to $\textbf{100.4}$%. Moreover, we employ bioinformatics tools in conjunction with gene representations to uncover pathway enrichment, regulation causal networks, and functions of disease-associated genes. Therefore, our model efficiently produces unified gene representations for the analysis of gene functions, tissue functions, diseases, and species evolution.
Tianyu Liu 0005, Yuge Wang, Rex Ying, Hongyu Zhao 0003
NeurIPS4
2023 HBV-infected hepatocellular carcinoma can be robustly classified into three clinically relevant subgroups by a novel analytical protocol
abstract
Liver cancer is the third leading cause of cancer-related death worldwide, and hepatocellular carcinoma (HCC) accounts for a relatively large proportion of all primary liver malignancies. Among the several known risk factors, hepatitis B virus (HBV) infection is one of the important causes of HCC. In this study, we demonstrated that the HBV-infected HCC patients could be robustly classified into three clinically relevant subgroups, i.e. Cluster1, Cluster2 and Cluster3, based on consistent differentially expressed mRNAs and proteins, which showed better generalization. The proposed three subgroups showed different molecular characteristics, immune microenvironment and prognostic survival characteristics. The Cluster1 subgroup had near-normal levels of metabolism-related proteins, low proliferation activity and good immune infiltration, which were associated with its good liver function, smaller tumor size, good prognosis, low alpha-fetoprotein (AFP) levels and lower clinical stage. In contrast, the Cluster3 subgroup had the lowest levels of metabolism-related proteins, which corresponded with its severe liver dysfunction. Also, high proliferation activity and poor immune microenvironment in Cluster3 subgroup were associated with its poor prognosis, larger tumor size, high AFP levels, high incidence of tumor thrombus and higher clinical stage. The characteristics of the Cluster2 subgroup were between the Cluster1 and Cluster3 groups. In addition, MCM2-7, RFC2-5, MSH2, MSH6, SMC2, SMC4, NCPAG and TOP2A proteins were significantly upregulated in the Cluster3 subgroup. Meanwhile, abnormally high phosphorylation levels of these proteins were associated with high levels of DNA repair, telomere maintenance and proliferative features. Therefore, these proteins could be identified as potential diagnostic and prognostic markers. In general, our research provided a novel analytical protocol and insights for the robust classification, treatment and prevention of HBV-infected HCC.
Leijie Li, Yuening Zhang, Yongyong Ren, Jianlei Gu, Xinbo Wang, Hongyu Zhao 0003, Hui Lu 0004
Briefings Bioinform.7
2023 A novel Bayesian framework for harmonizing information across tissues and studies to increase cell type deconvolution accuracy
abstract
Computational cell type deconvolution on bulk transcriptomics data can reveal cell type proportion heterogeneity across samples. One critical factor for accurate deconvolution is the reference signature matrix for different cell types. Compared with inferring reference signature matrices from cell lines, rapidly accumulating single-cell RNA-sequencing (scRNA-seq) data provide a richer and less biased resource. However, deriving cell type signature from scRNA-seq data is challenging due to high biological and technical noises. In this article, we introduce a novel Bayesian framework, tranSig, to improve signature matrix inference from scRNA-seq by leveraging shared cell type-specific expression patterns across different tissues and studies. Our simulations show that tranSig is robust to the number of signature genes and tissues specified in the model. Applications of tranSig to bulk RNA sequencing data from peripheral blood, bronchoalveolar lavage and aorta demonstrate its accuracy and power to characterize biological heterogeneity across groups. In summary, tranSig offers an accurate and robust approach to defining gene expression signatures of different cell types, facilitating improved in silico cell type deconvolutions.
Wenxuan Deng, Wei Jiang 0019, Xiting Yan, Ningshan Li, Milica Vukmirovic, Naftali Kaminski, Jing Wang 0003, Hongyu Zhao 0003
Briefings Bioinform.10
2023 Benchmarking of local genetic correlation estimation methods using summary statistics from genome-wide association studies
abstract
Local genetic correlation evaluates the correlation of additive genetic effects between different traits across the same genetic variants at a genomic locus. It has been proven informative for understanding the genetic similarities of complex traits beyond that captured by global genetic correlation calculated across the whole genome. Several summary-statistics-based approaches have been developed for estimating local genetic correlation, including $\rho$-hess, SUPERGNOVA and LAVA. However, there has not been a comprehensive evaluation of these methods to offer practical guidelines on the choices of these methods. In this study, we conduct benchmark comparisons of the performance of these three methods through extensive simulation and real data analyses. We focus on two technical difficulties in estimating local genetic correlation: sample overlaps across traits and local linkage disequilibrium (LD) estimates when only the external reference panels are available. Our simulations suggest the likelihood of incorrectly identifying correlated regions and local correlation estimation accuracy are highly dependent on the estimation of the local LD matrix. These observations are corroborated by real data analyses of 31 complex traits. Overall, our findings illuminate the distinct results yielded by different methods applied in post-genome-wide association studies (post-GWAS) local correlation studies. We underscore the sensitivity of local genetic correlation estimates and inferences to the precision of local LD estimation. These observations accentuate the vital need for ongoing refinement in methodologies.
Yiliang Zhang, Yunxuan Zhang, Hongyu Zhao 0003
Briefings Bioinform.4
2022 MZINBVA: variational approximation for multilevel zero-inflated negative-binomial models for association analysis in microbiome surveys
abstract
As our understanding of the microbiome has expanded, so has the recognition of its critical role in human health and disease, thereby emphasizing the importance of testing whether microbes are associated with environmental factors or clinical outcomes. However, many of the fundamental challenges that concern microbiome surveys arise from statistical and experimental design issues, such as the sparse and overdispersed nature of microbiome count data and the complex correlation structure among samples. For example, in the human microbiome project (HMP) dataset, the repeated observations across time points (level 1) are nested within body sites (level 2), which are further nested within subjects (level 3). Therefore, there is a great need for the development of specialized and sophisticated statistical tests. In this paper, we propose multilevel zero-inflated negative-binomial models for association analysis in microbiome surveys. We develop a variational approximation method for maximum likelihood estimation and inference. It uses optimization, rather than sampling, to approximate the log-likelihood and compute parameter estimates, provides a robust estimate of the covariance of parameter estimates and constructs a Wald-type test statistic for association testing. We evaluate and demonstrate the performance of our method using extensive simulation studies and an application to the HMP dataset. We have developed an R package MZINBVA to implement the proposed method, which is available from the GitHub repository https://github.com/liudoubletian/MZINBVA.
Peirong Xu, Yueyao Du, Hui Lu 0004, Hongyu Zhao 0003, Tao Wang 0067
Briefings Bioinform.5
2022 A Markov random field model-based approach for differentially expressed gene detection from single-cell RNA-seq data
abstract
The development of single-cell RNA-sequencing (scRNA-seq) technologies has offered insights into complex biological systems at the single-cell resolution. In particular, these techniques facilitate the identifications of genes showing cell-type-specific differential expressions (DE). In this paper, we introduce MARBLES, a novel statistical model for cross-condition DE gene detection from scRNA-seq data. MARBLES employs a Markov Random Field model to borrow information across similar cell types and utilizes cell-type-specific pseudobulk count to account for sample-level variability. Our simulation results showed that MARBLES is more powerful than existing methods to detect DE genes with an appropriate control of false positive rate. Applications of MARBLES to real data identified novel disease-related DE genes and biological pathways from both a single-cell lipopolysaccharide mouse dataset with 24 381 cells and 11 076 genes and a Parkinson's disease human data set with 76 212 cells and 15 891 genes. Overall, MARBLES is a powerful tool to identify cell-type-specific DE genes across conditions from scRNA-seq data.
Biqing Zhu, Le Zhang 0013, Sreeganga S. Chandra, Hongyu Zhao 0003
Briefings Bioinform.5
2022 ResPAN: a powerful batch correction model for scRNA-seq data through residual adversarial networks
abstract
MOTIVATION: With the advancement of technology, we can generate and access large-scale, high dimensional and diverse genomics data, especially through single-cell RNA sequencing (scRNA-seq). However, integrative downstream analysis from multiple scRNA-seq datasets remains challenging due to batch effects. RESULTS: In this article, we propose a light-structured deep learning framework called ResPAN for scRNA-seq data integration. ResPAN is based on Wasserstein Generative Adversarial Network (WGAN) combined with random walk mutual nearest neighbor pairing and fully skip-connected autoencoders to reduce the differences among batches. We also discuss the limitations of existing methods and demonstrate the advantages of our model over seven other methods through extensive benchmarking studies on both simulated data under various scenarios and real datasets across different scales. Our model achieves leading performance on both batch correction and biological information conservation and maintains scalable to datasets with over half a million cells. AVAILABILITY AND IMPLEMENTATION: An open-source implementation of ResPAN and scripts to reproduce the results can be downloaded from: https://github.com/AprilYuge/ResPAN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yuge Wang, Tianyu Liu 0005, Hongyu Zhao 0003
Bioinform.3
2022 fastANCOM: a fast method for analysis of compositions of microbiomes
abstract
SUMMARY: Analysis of compositions of microbiomes (ANCOM) compares the absolute abundances of microbes between two or more ecosystems using relative abundances in specimens derived from these ecosystems. Despite its impressive performance, there are two drawbacks to ANCOM. First, with K microbes it requires fitting K(K-1)/2 models for log-ratios of counts, and so can be computationally intensive. Second, it does not output P-values for microbes detected as differentially abundant. We propose a fast implementation of ANCOM, fastANCOM, that fits only K models for log-transformed counts. fastANCOM provides P-values to declare statistical significance and outputs log fold changes of abundance between groups. We demonstrate that fastANCOM compares favorably with existing differential abundance testing methods in terms of running time, false discovery rate and power. AVAILABILITY AND IMPLEMENTATION: fastANCOM is available at https://github.com/ZRChao/fastANCOM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hongyu Zhao 0003, Tao Wang 0067
Bioinform.3
2022 An unbiased kinship estimation method for genetic data analysis
abstract
Accurate estimate of relatedness is important for genetic data analyses, such as heritability estimation and association mapping based on data collected from genome-wide association studies. Inaccurate relatedness estimates may lead to biased heritability estimations and spurious associations. Individual-level genotype data are often used to estimate kinship coefficient between individuals. The commonly used sample correlation-based genomic relationship matrix (scGRM) method estimates kinship coefficient by calculating the average sample correlation coefficient among all single nucleotide polymorphisms (SNPs), where the observed allele frequencies are used to calculate both the expectations and variances of genotypes. Although this method is widely used, a substantial proportion of estimated kinship coefficients are negative, which are difficult to interpret. In this paper, through mathematical derivation, we show that there indeed exists bias in the estimated kinship coefficient using the scGRM method when the observed allele frequencies are regarded as true frequencies. This leads to negative bias for the average estimate of kinship among all individuals, which explains the estimated negative kinship coefficients. Based on this observation, we propose an unbiased estimation method, UKin, which can reduce kinship estimation bias. We justify our improved method with rigorous mathematical proof. We have conducted simulations as well as two real data analyses to compare UKin with scGRM and three other kinship estimating methods: rGRM, tsGRM, and KING. Our results demonstrate that both bias and root mean square error in kinship coefficient estimation could be reduced by using UKin. We further investigated the performance of UKin, KING, and three GRM-based methods in calculating the SNP-based heritability, and show that UKin can improve estimation accuracy for heritability regardless of the scale of SNP panel.
Wei Jiang 0019, Siting Li, Shuang Song 0006, Hongyu Zhao 0003
BMC Bioinform.5
2022 phyloMDA: an R package for phylogeny-aware microbiome data analysis
abstract
BACKGROUND: Modern sequencing technologies have generated low-cost microbiome survey datasets, across sample sites, conditions, and treatments, on an unprecedented scale and throughput. These datasets often come with a phylogenetic tree that provides a unique opportunity to examine how shared evolutionary history affects the different patterns in host-associated microbial communities. RESULTS: In this paper, we describe an R package, phyloMDA, for phylogeny-aware microbiome data analysis. It includes the Dirichlet-tree multinomial model for multivariate abundance data, tree-guided empirical Bayes estimation of microbial compositions, and tree-based multiscale regression methods with relative abundances as predictors. CONCLUSION: phyloMDA is a versatile and user-friendly tool to analyze microbiome data while incorporating the phylogenetic information and addressing some of the challenges posed by the data.
Hongyu Zhao 0003, Tao Wang 0067
BMC Bioinform.4
2022 Non-linear archetypal analysis of single-cell RNA-seq data by deep autoencoders
abstract
Advances in single-cell RNA sequencing (scRNA-seq) have led to successes in discovering novel cell types and understanding cellular heterogeneity among complex cell populations through cluster analysis. However, cluster analysis is not able to reveal continuous spectrum of states and underlying gene expression programs (GEPs) shared across cell types. We introduce scAAnet, an autoencoder for single-cell non-linear archetypal analysis, to identify GEPs and infer the relative activity of each GEP across cells. We use a count distribution-based loss term to account for the sparsity and overdispersion of the raw count data and add an archetypal constraint to the loss function of scAAnet. We first show that scAAnet outperforms existing methods for archetypal analysis across different metrics through simulations. We then demonstrate the ability of scAAnet to extract biologically meaningful GEPs using publicly available scRNA-seq datasets including a pancreatic islet dataset, a lung idiopathic pulmonary fibrosis dataset and a prefrontal cortex dataset.
Yuge Wang, Hongyu Zhao 0003
PLoS Comput. Biol.2
2021 Comparison of methods for estimating genetic correlation between complex traits using GWAS summary statistics
abstract
Genetic correlation is the correlation of phenotypic effects by genetic variants across the genome on two phenotypes. It is an informative metric to quantify the overall genetic similarity between complex traits, which provides insights into their polygenic genetic architecture. Several methods have been proposed to estimate genetic correlation based on data collected from genome-wide association studies (GWAS). Due to the easy access of GWAS summary statistics and computational efficiency, methods only requiring GWAS summary statistics as input have become more popular than methods utilizing individual-level genotype data. Here, we present a benchmark study for different summary-statistics-based genetic correlation estimation methods through simulation and real data applications. We focus on two major technical challenges in estimating genetic correlation: marker dependency caused by linkage disequilibrium (LD) and sample overlap between different studies. To assess the performance of different methods in the presence of these two challenges, we first conducted comprehensive simulations with diverse LD patterns and sample overlaps. Then we applied these methods to real GWAS summary statistics for a wide spectrum of complex traits. Based on these experiments, we conclude that methods relying on accurate LD estimation are less robust in real data applications due to the imprecision of LD obtained from reference panels. Our findings offer guidance on how to choose appropriate methods for genetic correlation estimation in post-GWAS analysis.
Yiliang Zhang, Youshu Cheng, Wei Jiang 0019, Yixuan Ye, Qiongshi Lu, Hongyu Zhao 0003
Briefings Bioinform.6
2021 Transformation and differential abundance analysis of microbiome data incorporating phylogeny
abstract
MOTIVATION: Microbiome data have proven extremely useful for understanding microbial communities and their impacts in health and disease. Although microbiome analysis methods and standards are evolving rapidly, obtaining meaningful and interpretable results from microbiome studies still requires careful statistical treatment. In particular, many existing and emerging methods for differential abundance (DA) analysis fail to account for the fact that microbiome data are high-dimensional and sparse, compositional, negatively and positively correlated and phylogenetically structured. To better describe microbiome data and improve the power of DA testing, there is still a great need for the continued development of appropriate statistical methodology. RESULTS: In this article, we propose a model-based approach for microbiome data transformation, and a phylogenetically informed procedure for DA testing based on the transformed data. First, we extend the Dirichlet-tree multinomial (DTM) to zero-inflated DTM for multivariate modeling of microbial counts, addressing data sparsity and correlation and phylogeny among bacterial taxa. Then, within this framework and using a Bayesian formulation, we introduce posterior mean transformation to convert raw counts into non-zero relative abundances that sum to one, accounting for the compositionality nature of microbiome data. Second, using the transformed data, we propose adaptive analysis of composition of microbiomes (adaANCOM) for DA testing by constructing log-ratios adaptively on the tree for each taxon, greatly reducing the computational complexity of ANCOM in high dimensions. Finally, we present extensive simulation studies, an analysis of HMP data across 18 body sites and 2 visits, and an application to a gut microbiome and malnutrition study, to investigate the performance of posterior mean transformation and adaANCOM. Comparisons with ANCOM and other DA testing procedures show that adaANCOM controls the false discovery rate well, allows for easy interpretation of the results, and is computationally efficient for high-dimensional problems. AVAILABILITY AND IMPLEMENTATION: The developed R package is available at https://github.com/ZRChao/adaANCOM. For replicability purposes, scripts for our simulations and data analysis are available at https://github.com/ZRChao/Papers_supplementary. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hongyu Zhao 0003, Tao Wang 0067
Bioinform.2
2021 A Markov random field model for network-based differential expression analysis of single-cell RNA-seq data
abstract
BACKGROUND: Recent development of single cell sequencing technologies has made it possible to identify genes with different expression (DE) levels at the cell type level between different groups of samples. In this article, we propose to borrow information through known biological networks to increase statistical power to identify differentially expressed genes (DEGs). RESULTS: We develop MRFscRNAseq, which is based on a Markov random field (MRF) model to appropriately accommodate gene network information as well as dependencies among cell types to identify cell-type specific DEGs. We implement an Expectation-Maximization (EM) algorithm with mean field-like approximation to estimate model parameters and a Gibbs sampler to infer DE status. Simulation study shows that our method has better power to detect cell-type specific DEGs than conventional methods while appropriately controlling type I error rate. The usefulness of our method is demonstrated through its application to study the pathogenesis and biological processes of idiopathic pulmonary fibrosis (IPF) using a single-cell RNA-sequencing (scRNA-seq) data set, which contains 18,150 protein-coding genes across 38 cell types on lung tissues from 32 IPF patients and 28 normal controls. CONCLUSIONS: The proposed MRF model is implemented in the R package MRFscRNAseq available on GitHub. By utilizing gene-gene and cell-cell networks, our method increases statistical power to detect differentially expressed genes from scRNA-seq data.
Biqing Zhu, Taylor Adams, Naftali Kaminski, Hongyu Zhao 0003
BMC Bioinform.6
2020 Inference of Dynamic Graph Changes for Functional Connectome
abstract
Dynamic functional connectivity is an effective measure for the brain’s responses to continuous stimuli. We propose an inferential method to detect the dynamic changes of brain networks based on time-varying graphical models. Whereas most existing methods focus on testing the existence of change points, the dynamics in the brain network offer more signals in many neuroscience studies. We propose a novel method to conduct hypothesis testing on changes in dynamic brain networks. We introduce a bootstrap statistic to approximate the supreme of the high-dimensional empirical processes over dynamically changing edges. Our simulations show that this framework can capture the change points with changed connectivity. Finally, we apply our method to a brain imaging dataset under a natural audio-video stimulus and illustrate that we are able to detect temporal changes in brain networks. The functions of the identified regions are consistent with specific emotional annotations, which are closely associated with changes inferred by our method.
Dingjue Ji, Yiliang Zhang, Hongyu Zhao 0003
AISTATS5
2020 NITUMID: Nonnegative matrix factorization-based Immune-TUmor MIcroenvironment Deconvolution
abstract
MOTIVATION: A number of computational methods have been proposed recently to profile tumor microenvironment (TME) from bulk RNA data, and they have proved useful for understanding microenvironment differences among therapeutic response groups. However, these methods are not able to account for tumor proportion nor variable mRNA levels across cell types. RESULTS: In this article, we propose a Nonnegative Matrix Factorization-based Immune-TUmor MIcroenvironment Deconvolution (NITUMID) framework for TME profiling that addresses these limitations. It is designed to provide robust estimates of tumor and immune cells proportions simultaneously, while accommodating mRNA level differences across cell types. Through comprehensive simulations and real data analyses, we demonstrate that NITUMID not only can accurately estimate tumor fractions and cell types' mRNA levels, which are currently unavailable in other methods; it also outperforms most existing deconvolution methods in regular cell type profiling accuracy. Moreover, we show that NITUMID can more effectively detect clinical and prognostic signals from gene expression profiles in tumor than other methods. AVAILABILITY AND IMPLEMENTATION: The algorithm is implemented in R. The source code can be downloaded at https://github.com/tdw1221/NITUMID. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Daiwei Tang, Hongyu Zhao 0003
Bioinform.3
2020 Predicting viral exposure response from modeling the changes of co-expression networks using time series gene expression data
abstract
BACKGROUND: Deciphering the relationship between clinical responses and gene expression profiles may shed light on the mechanisms underlying diseases. Most existing literature has focused on exploring such relationship from cross-sectional gene expression data. It is likely that the dynamic nature of time-series gene expression data is more informative in predicting clinical response and revealing the physiological process of disease development. However, it remains challenging to extract useful dynamic information from time-series gene expression data. RESULTS: We propose a statistical framework built on considering co-expression network changes across time from time series gene expression data. It first detects change point for co-expression networks and then employs a Bayesian multiple kernel learning method to predict exposure response. There are two main novelties in our method: the use of change point detection to characterize the co-expression network dynamics, and the use of kernel function to measure the similarity between subjects. Our algorithm allows exposure response prediction using dynamic network information across a collection of informative gene sets. Through parameter estimations, our model has clear biological interpretations. The performance of our method on the simulated data under different scenarios demonstrates that the proposed algorithm has better explanatory power and classification accuracy than commonly used machine learning algorithms. The application of our method to time series gene expression profiles measured in peripheral blood from a group of subjects with respiratory viral exposure shows that our method can predict exposure response at early stage (within 24 h) and the informative gene sets are enriched for pathways related to respiratory and influenza virus infection. CONCLUSIONS: The biological hypothesis in this paper is that the dynamic changes of the biological system are related to the clinical response. Our results suggest that when the relationship between the clinical response and a single gene or a gene set is not significant, we may benefit from studying the relationships among genes in gene sets that may lead to novel biological insights.
Fangli Dong, Tao Wang 0067, Hui Lu 0004, Hongyu Zhao 0003
BMC Bioinform.6
2020 An empirical Bayes approach to normalization and differential abundance testing for microbiome data
abstract
BACKGROUND: Advances in DNA sequencing have offered researchers an unprecedented opportunity to better study the variety of species living in and on the human body. However, the analysis of microbiome data is complicated by several challenges. First, the sequencing depth may vary by orders of magnitude across samples. Second, species are rare and the data often contain many zeros. Third, the specimen is a fraction of the microbial ecosystem, and so the data are compositional carrying only relative information. Other characteristics of microbiome data include pronounced over-dispersion in taxon abundances, and the existence of a phylogenetic tree that relates all bacterial species. To address some of these challenges, microbiome analysis workflows often normalize the read counts prior to downstream analysis. However, there are limitations in the current literature on the normalization of microbiome data. RESULTS: Under the multinomial distribution for the read counts and a prior for the unknown proportions, we propose an empirical Bayes approach to microbiome data normalization. Using a tree-based extension of the Dirichlet prior, we further extend our method by incorporating the phylogenetic tree into the normalization process. We study the impact of normalization on differential abundance analysis. In the presence of tree structure, we propose a phylogeny-aware detection procedure. CONCLUSIONS: Extensive simulations and gut microbiome data applications are conducted to demonstrate the superior performance of our empirical Bayes method over other normalization methods, and over commonly-used methods for differential abundance testing. Original R scripts are available at GitHub (https://github.com/liudoubletian/eBay).
Hongyu Zhao 0003, Tao Wang 0067
BMC Bioinform.2
2020 Leveraging functional annotation to identify genes associated with complex diseases
abstract
To increase statistical power to identify genes associated with complex traits, a number of transcriptome-wide association study (TWAS) methods have been proposed using gene expression as a mediating trait linking genetic variations and diseases. These methods first predict expression levels based on inferred expression quantitative trait loci (eQTLs) and then identify expression-mediated genetic effects on diseases by associating phenotypes with predicted expression levels. The success of these methods critically depends on the identification of eQTLs, which may not be functional in the corresponding tissue, due to linkage disequilibrium (LD) and the correlation of gene expression between tissues. Here, we introduce a new method called T-GEN (Transcriptome-mediated identification of disease-associated Genes with Epigenetic aNnotation) to identify disease-associated genes leveraging epigenetic information. Through prioritizing SNPs with tissue-specific epigenetic annotation, T-GEN can better identify SNPs that are both statistically predictive and biologically functional. We found that a significantly higher percentage (an increase of 18.7% to 47.2%) of eQTLs identified by T-GEN are inferred to be functional by ChromHMM and more are deleterious based on their Combined Annotation Dependent Depletion (CADD) scores. Applying T-GEN to 207 complex traits, we were able to identify more trait-associated genes (ranging from 7.7% to 102%) than those from existing methods. Among the identified genes associated with these traits, T-GEN can better identify genes with high (>0.99) pLI scores compared to other methods. When T-GEN was applied to late-onset Alzheimer's disease, we identified 96 genes located at 15 loci, including two novel loci not implicated in previous GWAS. We further replicated 50 genes in an independent GWAS, including one of the two novel loci.
Wenfeng Zhang, Geyu Zhou, Qiongshi Lu, Hongyu Zhao 0003
PLoS Comput. Biol.8
2020 Leveraging effect size distributions to improve polygenic risk scores derived from summary statistics of genome-wide association studies
abstract
Genetic risk prediction is an important problem in human genetics, and accurate prediction can facilitate disease prevention and treatment. Calculating polygenic risk score (PRS) has become widely used due to its simplicity and effectiveness, where only summary statistics from genome-wide association studies are needed in the standard method. Recently, several methods have been proposed to improve standard PRS by utilizing external information, such as linkage disequilibrium and functional annotations. In this paper, we introduce EB-PRS, a novel method that leverages information for effect sizes across all the markers to improve prediction accuracy. Compared to most existing genetic risk prediction methods, our method does not need to tune parameters nor external information. Real data applications on six diseases, including asthma, breast cancer, celiac disease, Crohn's disease, Parkinson's disease and type 2 diabetes show that EB-PRS achieved 307.1%, 42.8%, 25.5%, 3.1%, 74.3% and 49.6% relative improvements in terms of predictive r2 over standard PRS method with optimally tuned parameters. Besides, compared to LDpred that makes use of LD information, EB-PRS also achieved 37.9%, 33.6%, 8.6%, 36.2%, 40.6% and 10.8% relative improvements. We note that our method is not the first method leveraging effect size distributions. Here we first justify our method by presenting theoretical optimal property over existing methods in this class of methods, and substantiate our theoretical result with extensive simulation results. The R-package EBPRS that implements our method is available on CRAN.
Shuang Song 0006, Wei Jiang 0019, Lin Hou 0003, Hongyu Zhao 0003
PLoS Comput. Biol.4
2019 An evaluation of noncoding genome annotation tools through enrichment analysis of 15 genome-wide association studies
abstract
The overwhelming list of new bacterial genomes becoming available on a daily basis makes accurate genome annotation an essential step that ultimately determines the relevance of thousands of genomes stored in public databanks. The MicroScope platform (http://www.genoscope.cns.fr/agc/microscope) is an integrative resource that supports systematic and efficient revision of microbial genome annotation, data management and comparative analysis. Starting from the results of our syntactic, functional and relational annotation pipelines, MicroScope provides an integrated environment for the expert annotation and comparative analysis of prokaryotic genomes. It combines tools and graphical interfaces to analyze genomes and to perform the manual curation of gene function in a comparative genomics and metabolic context. In this article, we describe the free-of-charge MicroScope services for the annotation and analysis of microbial (meta)genomes, transcriptomic and re-sequencing data. Then, the functionalities of the platform are presented in a way providing practical guidance and help to the nonspecialists in bioinformatics. Newly integrated analysis tools (i.e. prediction of virulence and resistance genes in bacterial genomes) and original method recently developed (the pan-genome graph representation) are also described. Integrated environments such as MicroScope clearly contribute, through the user community, to help maintaining accurate resources.
Boyang Li 0011, Qiongshi Lu, Hongyu Zhao 0003
Briefings Bioinform.3
2019 ProteomicsBrowser: MS/proteomics data visualization and investigation
abstract
SUMMARY: Large-scale, quantitative proteomics data are being generated at ever increasing rates by high-throughput, mass spectrometry technologies. However, due to the complexity of these large datasets as well as the increasing numbers of post-translational modifications (PTMs) that are being identified, developing effective methods for proteomic visualization has been challenging. ProteomicsBrowser was designed to meet this need for comprehensive data visualization. Using peptide information files exported from mass spectrometry search engines or quantitative tools as input, the peptide sequences are aligned to an internal protein database such as UniProtKB. Each identified peptide ion including those with PTMs is then visualized along the parent protein in the Browser. A unique property of ProteomicsBrowser is the ability to combine overlapping peptides in different ways to focus analysis of sequence coverage, charge state or PTMs. ProteomicsBrowser includes other useful functions, such as a data filtering tool and basic statistical analyses to qualify quantitative data. AVAILABILITY AND IMPLEMENTATION: ProteomicsBrowser is implemented in Java8 and is available at https://medicine.yale.edu/keck/nida/proteomicsbrowser.aspx and https://github.com/peng-gang/ProteomicsBrowser. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Gang Peng 0004, Rashaun Wilson, Yishuo Tang, TuKiet T. Lam, Angus C. Nairn, Kenneth R. Williams, Hongyu Zhao 0003
Bioinform.7
2019 Network Clustering Analysis Using Mixture Exponential-Family Random Graph Models and Its Application in Genetic Interaction Data
abstract
MOTIVATION: Epistatic miniarrary profile (EMAP) studies have enabled the mapping of large-scale genetic interaction networks and generated large amounts of data in model organisms. It provides an incredible set of molecular tools and advanced technologies that should be efficiently understanding the relationship between the genotypes and phenotypes of individuals. However, the network information gained from EMAP cannot be fully exploited using the traditional statistical network models. Because the genetic network is always heterogeneous, for example, the network structure features for one subset of nodes are different from those of the left nodes. Exponential-family random graph models (ERGMs) are a family of statistical models, which provide a principled and flexible way to describe the structural features (e.g., the density, centrality, and assortativity) of an observed network. However, the single ERGM is not enough to capture this heterogeneity of networks. In this paper, we consider a mixture ERGM (MixtureEGRM) networks, which model a network with several communities, where each community is described by a single EGRM. RESULTS: EM algorithm is a classical method to solve the mixture problem, however, it will be very slow when the data size is huge in the numerous applications. We adopt an efficient novel online graph clustering algorithm to classify the graph nodes and estimate the ERGM parameters for the MixtureERGM. In comparison studies, the MixtureERGM outperforms the role analysis for the network cluster in which the mixture of exponential-family random graph model is developed for many ego-network according to their roles. One genetic interaction network of yeast and two real social networks (provided as supplemental materials, which can be found on the Computer Society Digital Library at http://doi.ieeecomputersociety.org/10.1109/TCBB.2017.2743711) show the wide potential application of the MixtureERGM.
Yishu Wang 0003, Huaying Fang, Dejie Yang, Hongyu Zhao 0003, Minghua Deng
IEEE ACM Trans. Comput. Biol. Bioinform.4
2018 Improving SNP prioritization and pleiotropic architecture estimation by incorporating prior knowledge using graph-GPA
abstract
Summary: Integration of genetic studies for multiple phenotypes is a powerful approach to improving the identification of genetic variants associated with complex traits. Although it has been shown that leveraging shared genetic basis among phenotypes, namely pleiotropy, can increase statistical power to identify risk variants, it remains challenging to effectively integrate genome-wide association study (GWAS) datasets for a large number of phenotypes. We previously developed graph-GPA, a Bayesian hierarchical model that integrates multiple GWAS datasets to boost statistical power for the identification of risk variants and to estimate pleiotropic architecture within a unified framework. Here we propose a novel improvement of graph-GPA which incorporates external knowledge about phenotype-phenotype relationship to guide the estimation of genetic correlation and the association mapping. The application of graph-GPA to GWAS datasets for 12 complex diseases with a prior disease graph obtained from a text mining of biomedical literature illustrates its power to improve the identification of risk genetic variants and to facilitate understanding of genetic relationship among complex diseases. Availability and implementation: graph-GPA is implemented as an R package 'GGPA', which is publicly available at http://dongjunchung.github.io/GGPA/. DDNet, a web interface to query diseases of interest and download a prior disease graph obtained from a text mining of biomedical literature, is publicly available at http://www.chunglab.io/ddnet/. Supplementary information: Supplementary data are available at Bioinformatics online.
Hang Joon Kim, Zhenning Yu, Andrew Lawson, Hongyu Zhao 0003, Dongjun Chung
Bioinform.4
2018 Spectral clustering based on learning similarity matrix
abstract
Motivation: Single-cell RNA-sequencing (scRNA-seq) technology can generate genome-wide expression data at the single-cell levels. One important objective in scRNA-seq analysis is to cluster cells where each cluster consists of cells belonging to the same cell type based on gene expression patterns. Results: We introduce a novel spectral clustering framework that imposes sparse structures on a target matrix. Specifically, we utilize multiple doubly stochastic similarity matrices to learn a similarity matrix, motivated by the observation that each similarity matrix can be a different informative representation of the data. We impose a sparse structure on the target matrix followed by shrinking pairwise differences of the rows in the target matrix, motivated by the fact that the target matrix should have these structures in the ideal case. We solve the proposed non-convex problem iteratively using the ADMM algorithm and show the convergence of the algorithm. We evaluate the performance of the proposed clustering method on various simulated as well as real scRNA-seq data, and show that it can identify clusters accurately and robustly. Availability and implementation: The algorithm is implemented in MATLAB. The source code can be downloaded at https://github.com/ishspsy/project/tree/master/MPSSC. Supplementary information: Supplementary data are available at Bioinformatics online.
Hongyu Zhao 0003
Bioinform.2
2017 GRAPE: a pathway template method to characterize tissue-specific functionality from gene expression profiles
abstract
BACKGROUND: Personalizing treatment regimes based on gene expression profiles of individual tumors will facilitate management of cancer. Although many methods have been developed to identify pathways perturbed in tumors, the results are often not generalizable across independent datasets due to the presence of platform/batch effects. There is a need to develop methods that are robust to platform/batch effects and able to identify perturbed pathways in individual samples. RESULTS: We present Gene-Ranking Analysis of Pathway Expression (GRAPE) as a novel method to identify abnormal pathways in individual samples that is robust to platform/batch effects in gene expression profiles generated by multiple platforms. GRAPE first defines a template consisting of an ordered set of pathway genes to characterize the normative state of a pathway based on the relative rankings of gene expression levels across a set of reference samples. This template can be used to assess whether a sample conforms to or deviates from the typical behavior of the reference samples for this pathway. We demonstrate that GRAPE performs well versus existing methods in classifying tissue types within a single dataset, and that GRAPE achieves superior robustness and generalizability across different datasets. A powerful feature of GRAPE is the ability to represent individual gene expression profiles as a vector of pathways scores. We present applications to the analyses of breast cancer subtypes and different colonic diseases. We perform survival analysis of several TCGA subtypes and find that GRAPE pathway scores perform well in comparison to other methods. CONCLUSIONS: GRAPE templates offer a novel approach for summarizing the behavior of gene-sets across a collection of gene expression profiles. These templates offer superior robustness across distinct experimental batches compared to existing methods. GRAPE pathway scores enable identification of abnormal gene-set behavior in individual samples using a non-competitive approach that is fundamentally distinct from popular enrichment-based methods. GRAPE may be an appropriate tool for researchers seeking to identify individual samples displaying abnormal gene-set behavior as well as to explore differences in the consensus gene-set behavior of groups of samples. GRAPE is available in R for download at https://CRAN.R-project.org/package=GRAPE .
Michael I. Klein, David F. Stern, Hongyu Zhao 0003
BMC Bioinform.3
2017 A novel pathway-based distance score enhances assessment of disease heterogeneity in gene expression
abstract
BACKGROUND: Distance based unsupervised clustering of gene expression data is commonly used to identify heterogeneity in biologic samples. However, high noise levels in gene expression data and relatively high correlation between genes are often encountered, so traditional distances such as Euclidean distance may not be effective at discriminating the biological differences between samples. An alternative method to examine disease phenotypes is to use pre-defined biological pathways. These pathways have been shown to be perturbed in different ways in different subjects who have similar clinical features. We hypothesize that differences in the expressions of genes in a given pathway are more predictive of differences in biological differences compared to standard approaches and if integrated into clustering analysis will enhance the robustness and accuracy of the clustering method. To examine this hypothesis, we developed a novel computational method to assess the biological differences between samples using gene expression data by assuming that ontologically defined biological pathways in biologically similar samples have similar behavior. RESULTS: Pre-defined biological pathways were downloaded and genes in each pathway were used to cluster samples using the Gaussian mixture model. The clustering results across different pathways were then summarized to calculate the pathway-based distance score between samples. This method was applied to both simulated and real data sets and compared to the traditional Euclidean distance and another pathway-based clustering method, Pathifier. The results show that the pathway-based distance score performs significantly better than the Euclidean distance, especially when the heterogeneity is low and genes in the same pathways are correlated. Compared to Pathifier, we demonstrated that our approach achieves higher accuracy and robustness for small pathways. When the pathway size is large, by downsampling the pathways into smaller pathways, our approach was able to achieve comparable performance. CONCLUSIONS: We have developed a novel distance score that represents the biological differences between samples using gene expression data and pre-defined biological pathway information. Application of this distance score results in more accurate, robust, and biologically meaningful clustering results in both simulated data and real data when compared to traditional methods. It also has comparable or better performance compared to Pathifier.
Xiting Yan, Anqi Liang, Lauren Cohn, Hongyu Zhao 0003, Geoffrey Lowell Chupp
BMC Bioinform.5
2017 A comparison of graph- and kernel-based -omics data integration algorithms for classifying complex traits
abstract
BACKGROUND: High-throughput sequencing data are widely collected and analyzed in the study of complex diseases in quest of improving human health. Well-studied algorithms mostly deal with single data source, and cannot fully utilize the potential of these multi-omics data sources. In order to provide a holistic understanding of human health and diseases, it is necessary to integrate multiple data sources. Several algorithms have been proposed so far, however, a comprehensive comparison of data integration algorithms for classification of binary traits is currently lacking. RESULTS: In this paper, we focus on two common classes of integration algorithms, graph-based that depict relationships with subjects denoted by nodes and relationships denoted by edges, and kernel-based that can generate a classifier in feature space. Our paper provides a comprehensive comparison of their performance in terms of various measurements of classification accuracy and computation time. Seven different integration algorithms, including graph-based semi-supervised learning, graph sharpening integration, composite association network, Bayesian network, semi-definite programming-support vector machine (SDP-SVM), relevance vector machine (RVM) and Ada-boost relevance vector machine are compared and evaluated with hypertension and two cancer data sets in our study. In general, kernel-based algorithms create more complex models and require longer computation time, but they tend to perform better than graph-based algorithms. The performance of graph-based algorithms has the advantage of being faster computationally. CONCLUSIONS: The empirical results demonstrate that composite association network, relevance vector machine, and Ada-boost RVM are the better performers. We provide recommendations on how to choose an appropriate algorithm for integrating data from multiple sources.
Kang K. Yan, Hongyu Zhao 0003, Herbert Pang
BMC Bioinform.2
2017 graph-GPA: A graphical model for prioritizing GWAS results and investigating pleiotropic architecture
abstract
Genome-wide association studies (GWAS) have identified tens of thousands of genetic variants associated with hundreds of phenotypes and diseases, which have provided clinical and medical benefits to patients with novel biomarkers and therapeutic targets. However, identification of risk variants associated with complex diseases remains challenging as they are often affected by many genetic variants with small or moderate effects. There has been accumulating evidence suggesting that different complex traits share common risk basis, namely pleiotropy. Recently, several statistical methods have been developed to improve statistical power to identify risk variants for complex traits through a joint analysis of multiple GWAS datasets by leveraging pleiotropy. While these methods were shown to improve statistical power for association mapping compared to separate analyses, they are still limited in the number of phenotypes that can be integrated. In order to address this challenge, in this paper, we propose a novel statistical framework, graph-GPA, to integrate a large number of GWAS datasets for multiple phenotypes using a hidden Markov random field approach. Application of graph-GPA to a joint analysis of GWAS datasets for 12 phenotypes shows that graph-GPA improves statistical power to identify risk variants compared to statistical methods based on smaller number of GWAS datasets. In addition, graph-GPA also promotes better understanding of genetic mechanisms shared among phenotypes, which can potentially be useful for the development of improved diagnosis and therapeutics. The R implementation of graph-GPA is currently available at https://dongjunchung.github.io/GGPA/.
Dongjun Chung, Hang Joon Kim, Hongyu Zhao 0003
PLoS Comput. Biol.3
2017 Leveraging functional annotations in genetic risk prediction for human complex diseases
abstract
Genetic risk prediction is an important goal in human genetics research and precision medicine. Accurate prediction models will have great impacts on both disease prevention and early treatment strategies. Despite the identification of thousands of disease-associated genetic variants through genome wide association studies (GWAS), genetic risk prediction accuracy remains moderate for most diseases, which is largely due to the challenges in both identifying all the functionally relevant variants and accurately estimating their effect sizes in the presence of linkage disequilibrium. In this paper, we introduce AnnoPred, a principled framework that leverages diverse types of genomic and epigenomic functional annotations in genetic risk prediction for complex diseases. AnnoPred is trained using GWAS summary statistics in a Bayesian framework in which we explicitly model various functional annotations and allow for linkage disequilibrium estimated from reference genotype data. Compared with state-of-the-art risk prediction methods, AnnoPred achieves consistently improved prediction accuracy in both extensive simulations and real data.
Yiming Hu, Qiongshi Lu, Ryan Powles, Xinwei Yao 0002, Can Yang 0002, Xinran Xu, Hongyu Zhao 0003
PLoS Comput. Biol.8
2016 Predicting synergistic effects between compounds through their structural similarity and effects on transcriptomes
abstract
MOTIVATION: Combinatorial therapies have been under intensive research for cancer treatment. However, due to the large number of possible combinations among candidate compounds, exhaustive screening is prohibitive. Hence, it is important to develop computational tools that can predict compound combination effects, prioritize combinations and limit the search space to facilitate and accelerate the development of combinatorial therapies. RESULTS: In this manuscript we consider the NCI-DREAM Drug Synergy Prediction Challenge dataset to identify features informative about combination effects. Through systematic exploration of differential expression profiles after single compound treatments and comparison of molecular structures of compounds, we found that synergistic levels of combinations are statistically significantly associated with compounds' dissimilarity in structure and similarity in induced gene expression changes. These two types of features offer complementary information in predicting experimentally measured combination effects of compound pairs. Our findings offer insights on the mechanisms underlying different combination effects and may help prioritize promising combinations in the very large search space. AVAILABILITY AND IMPLEMENTATION: The R code for the analysis is available on https://github.com/YiyiLiu1/DrugCombination CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online.
Yiyi Liu, Hongyu Zhao 0003
Bioinform.2
2016 GenoWAP: GWAS signal prioritization through integrated analysis of genomic functional annotation
abstract
MOTIVATION: Genome-wide association study (GWAS) has been a great success in the past decade. However, significant challenges still remain in both identifying new risk loci and interpreting results. Bonferroni-corrected significance level is known to be conservative, leading to insufficient statistical power when the effect size is moderate at risk locus. Complex structure of linkage disequilibrium also makes it challenging to separate causal variants from nonfunctional ones in large haplotype blocks. Under such circumstances, a computational approach that may increase signal replication rate and identify potential functional sites among correlated markers is urgently needed. RESULTS: We describe GenoWAP, a GWAS signal prioritization method that integrates genomic functional annotation and GWAS test statistics. The effectiveness of GenoWAP is demonstrated through its applications to Crohn's disease and schizophrenia using the largest studies available, where highly ranked loci show substantially stronger signals in the whole dataset after prioritization based on a subset of samples. At the single nucleotide polymorphism (SNP) level, top ranked SNPs after prioritization have both higher replication rates and consistently stronger enrichment of eQTLs. Within each risk locus, GenoWAP may be able to distinguish functional sites from groups of correlated SNPs. AVAILABILITY AND IMPLEMENTATION: GenoWAP is freely available on the web at http://genocanyon.med.yale.edu/GenoWAP.
Qiongshi Lu, Xinwei Yao 0002, Yiming Hu, Hongyu Zhao 0003
Bioinform.4
2016 NBLDA: negative binomial linear discriminant analysis for RNA-Seq data
abstract
BACKGROUND: RNA-sequencing (RNA-Seq) has become a powerful technology to characterize gene expression profiles because it is more accurate and comprehensive than microarrays. Although statistical methods that have been developed for microarray data can be applied to RNA-Seq data, they are not ideal due to the discrete nature of RNA-Seq data. The Poisson distribution and negative binomial distribution are commonly used to model count data. Recently, Witten (Annals Appl Stat 5:2493-2518, 2011) proposed a Poisson linear discriminant analysis for RNA-Seq data. The Poisson assumption may not be as appropriate as the negative binomial distribution when biological replicates are available and in the presence of overdispersion (i.e., when the variance is larger than or equal to the mean). However, it is more complicated to model negative binomial variables because they involve a dispersion parameter that needs to be estimated. RESULTS: In this paper, we propose a negative binomial linear discriminant analysis for RNA-Seq data. By Bayes' rule, we construct the classifier by fitting a negative binomial model, and propose some plug-in rules to estimate the unknown parameters in the classifier. The relationship between the negative binomial classifier and the Poisson classifier is explored, with a numerical investigation of the impact of dispersion on the discriminant score. Simulation results show the superiority of our proposed method. We also analyze two real RNA-Seq data sets to demonstrate the advantages of our method in real-world applications. CONCLUSIONS: We have developed a new classifier using the negative binomial model for RNA-seq data classification. Our simulation results show that our proposed classifier has a better performance than existing works. The proposed classifier can serve as an effective tool for classifying RNA-seq data. Based on the comparison results, we have provided some guidelines for scientists to decide which method should be used in the discriminant analysis of RNA-Seq data. R code is available at http://www.comp.hkbu.edu.hk/~xwan/NBLDA.R or https://github.com/yangchadam/NBLDA.
Hongyu Zhao 0003, Tiejun Tong
BMC Bioinform.2
2016 Leveraging protein quaternary structure to identify oncogenic driver mutations
abstract
BACKGROUND: Identifying key "driver" mutations which are responsible for tumorigenesis is critical in the development of new oncology drugs. Due to multiple pharmacological successes in treating cancers that are caused by such driver mutations, a large body of methods have been developed to differentiate these mutations from the benign "passenger" mutations which occur in the tumor but do not further progress the disease. Under the hypothesis that driver mutations tend to cluster in key regions of the protein, the development of algorithms that identify these clusters has become a critical area of research. RESULTS: We have developed a novel methodology, QuartPAC (Quaternary Protein Amino acid Clustering), that identifies non-random mutational clustering while utilizing the protein quaternary structure in 3D space. By integrating the spatial information in the Protein Data Bank (PDB) and the mutational data in the Catalogue of Somatic Mutations in Cancer (COSMIC), QuartPAC is able to identify clusters which are otherwise missed in a variety of proteins. The R package is available on Bioconductor at: http://bioconductor.jp/packages/3.1/bioc/html/QuartPAC.html . CONCLUSION: QuartPAC provides a unique tool to identify mutational clustering while accounting for the complete folded protein quaternary structure.
Gregory A. Ryslik, Yuwei Cheng, Yorgo Modis, Hongyu Zhao 0003
BMC Bioinform.4
2016 Efficient Drug-Pathway Association Analysis via Integrative Penalized Matrix Decomposition
abstract
Traditional drug discovery practice usually follows the "one drug - one target" approach, seeking to identify drug molecules that act on individual targets, which ignores the systemic nature of human diseases. Pathway-based drug discovery recently emerged as an appealing approach to overcome this limitation. An important first step of such pathway-based drug discovery is to identify associations between drug molecules and biological pathways. This task has been made feasible by the accumulating data from high-throughput transcription and drug sensitivity profiling. In this paper, we developed "iPaD", an integrative Penalized Matrix Decomposition method to identify drug-pathway associations through jointly modeling of such high-throughput transcription and drug sensitivity data. A scalable bi-convex optimization algorithm was implemented and gave iPaD tremendous advantage in computational efficiency over current state-of-the-art method, which allows it to handle the ever-growing large-scale data sets that current method cannot afford to. On two widely used real data sets, iPaD also significantly outperformed the current method in terms of the number of validated drug-pathway associations that were identified. The Matlab code of our algorithm publicly available at http://licong-jason.github.io/iPaD/.
Can Yang 0002, Greg Hather, Ray Liu, Hongyu Zhao 0003
IEEE ACM Trans. Comput. Biol. Bioinform.5
2015 CCLasso: correlation inference for compositional data through Lasso
abstract
MOTIVATION: Direct analysis of microbial communities in the environment and human body has become more convenient and reliable owing to the advancements of high-throughput sequencing techniques for 16S rRNA gene profiling. Inferring the correlation relationship among members of microbial communities is of fundamental importance for genomic survey study. Traditional Pearson correlation analysis treating the observed data as absolute abundances of the microbes may lead to spurious results because the data only represent relative abundances. Special care and appropriate methods are required prior to correlation analysis for these compositional data. RESULTS: In this article, we first discuss the correlation definition of latent variables for compositional data. We then propose a novel method called CCLasso based on least squares with [Formula: see text] penalty to infer the correlation network for latent variables of compositional data from metagenomic data. An effective alternating direction algorithm from augmented Lagrangian method is used to solve the optimization problem. The simulation results show that CCLasso outperforms existing methods, e.g. SparCC, in edge recovery for compositional data. It also compares well with SparCC in estimating correlation network of microbe species from the Human Microbiome Project. AVAILABILITY AND IMPLEMENTATION: CCLasso is open source and freely available from https://github.com/huayingfang/CCLasso under GNU LGPL v3. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Huaying Fang, Chengcheng Huang, Hongyu Zhao 0003, Minghua Deng
Bioinform.3
2015 M3-S: a genotype calling method incorporating information from samples with known genotypes
abstract
BACKGROUND: A key challenge in analyzing high throughput Single Nucleotide Polymorphism (SNP) arrays is the accurate inference of genotypes for SNPs with low minor allele frequencies. A number of calling algorithms have been developed to infer genotypes for common SNPs, but they are limited in their performance in calling rare SNPs. The existing algorithms can be broadly classified into three categories, including: population-based methods, SNP-based methods, and a hybrid of the two approaches. Despite the relatively better performance of the hybrid approach, it is still challenging to analyze rare SNPs. RESULTS: We propose to utilize information from samples with known genotypes to develop a two stage genotyping procedure, namely M(3)-S, for rare SNP calling. This new approach can improve genotyping accuracy through clearly defining the boundaries of genotype clusters from samples with known genotypes, and enlarge the call rate by combining the simulated data based on the inferred genotype clusters information with the study population. CONCLUSIONS: Applications to real data demonstrates that this new approach M(3)-S outperforms existing methods in calling rare SNPs.
Gengxin Li, Hongyu Zhao 0003
BMC Bioinform.2
2015 The application of sparse estimation of covariance matrix to quadratic discriminant analysis
abstract
BACKGROUND: Although Linear Discriminant Analysis (LDA) is commonly used for classification, it may not be directly applied in genomics studies due to the large p, small n problem in these studies. Different versions of sparse LDA have been proposed to address this significant challenge. One implicit assumption of various LDA-based methods is that the covariance matrices are the same across different classes. However, rewiring of genetic networks (therefore different covariance matrices) across different diseases has been observed in many genomics studies, which suggests that LDA and its variations may be suboptimal for disease classifications. However, it is not clear whether considering differing genetic networks across diseases can improve classification in genomics studies. RESULTS: We propose a sparse version of Quadratic Discriminant Analysis (SQDA) to explicitly consider the differences of the genetic networks across diseases. Both simulation and real data analysis are performed to compare the performance of SQDA with six commonly used classification methods. CONCLUSIONS: SQDA provides more accurate classification results than other methods for both simulated and real data. Our method should prove useful for classification in genomics studies and other research settings, where covariances differ among classes.
Jiehuan Sun, Hongyu Zhao 0003
BMC Bioinform.2
2015 Functional Module Analysis for Gene Coexpression Networks with Network Integration
abstract
Network has been a general tool for studying the complex interactions between different genes, proteins, and other small molecules. Module as a fundamental property of many biological networks has been widely studied and many computational methods have been proposed to identify the modules in an individual network. However, in many cases, a single network is insufficient for module analysis due to the noise in the data or the tuning of parameters when building the biological network. The availability of a large amount of biological networks makes network integration study possible. By integrating such networks, more informative modules for some specific disease can be derived from the networks constructed from different tissues, and consistent factors for different diseases can be inferred. In this paper, we have developed an effective method for module identification from multiple networks under different conditions. The problem is formulated as an optimization model, which combines the module identification in each individual network and alignment of the modules from different networks together. An approximation algorithm based on eigenvector computation is proposed. Our method outperforms the existing methods, especially when the underlying modules in multiple networks are different in simulation studies. We also applied our method to two groups of gene coexpression networks for humans, which include one for three different cancers, and one for three tissues from the morbidly obese patients. We identified 13 modules with three complete subgraphs, and 11 modules with two complete subgraphs, respectively. The modules were validated through Gene Ontology enrichment and KEGG pathway enrichment analysis. We also showed that the main functions of most modules for the corresponding disease have been addressed by other researchers, which may provide the theoretical basis for further studying the modules experimentally.
Shuqin Zhang, Hongyu Zhao 0003, Michael Kwok-Po Ng
IEEE ACM Trans. Comput. Biol. Bioinform.2
2014 A spatial simulation approach to account for protein structure when identifying non-random somatic mutations
abstract
BACKGROUND: Current research suggests that a small set of "driver" mutations are responsible for tumorigenesis while a larger body of "passenger" mutations occur in the tumor but do not progress the disease. Due to recent pharmacological successes in treating cancers caused by driver mutations, a variety of methodologies that attempt to identify such mutations have been developed. Based on the hypothesis that driver mutations tend to cluster in key regions of the protein, the development of cluster identification algorithms has become critical. RESULTS: We have developed a novel methodology, SpacePAC (Spatial Protein Amino acid Clustering), that identifies mutational clustering by considering the protein tertiary structure directly in 3D space. By combining the mutational data in the Catalogue of Somatic Mutations in Cancer (COSMIC) and the spatial information in the Protein Data Bank (PDB), SpacePAC is able to identify novel mutation clusters in many proteins such as FGFR3 and CHRM2. In addition, SpacePAC is better able to localize the most significant mutational hotspots as demonstrated in the cases of BRAF and ALK. The R package is available on Bioconductor at: http://www.bioconductor.org/packages/release/bioc/html/SpacePAC.html. CONCLUSION: SpacePAC adds a valuable tool to the identification of mutational clusters while considering protein tertiary structure.
Gregory A. Ryslik, Yuwei Cheng, Kei-Hoi Cheung, Robert D. Bjornson, Daniel Zelterman, Yorgo Modis, Hongyu Zhao 0003
BMC Bioinform.7
2014 A graph theoretic approach to utilizing protein structure to identify non-random somatic mutations
abstract
BACKGROUND: It is well known that the development of cancer is caused by the accumulation of somatic mutations within the genome. For oncogenes specifically, current research suggests that there is a small set of "driver" mutations that are primarily responsible for tumorigenesis. Further, due to recent pharmacological successes in treating these driver mutations and their resulting tumors, a variety of approaches have been developed to identify potential driver mutations using methods such as machine learning and mutational clustering. We propose a novel methodology that increases our power to identify mutational clusters by taking into account protein tertiary structure via a graph theoretical approach. RESULTS: We have designed and implemented GraphPAC (Graph Protein Amino acid Clustering) to identify mutational clustering while considering protein spatial structure. Using GraphPAC, we are able to detect novel clusters in proteins that are known to exhibit mutation clustering as well as identify clusters in proteins without evidence of prior clustering based on current methods. Specifically, by utilizing the spatial information available in the Protein Data Bank (PDB) along with the mutational data in the Catalogue of Somatic Mutations in Cancer (COSMIC), GraphPAC identifies new mutational clusters in well known oncogenes such as EGFR and KRAS. Further, by utilizing graph theory to account for the tertiary structure, GraphPAC discovers clusters in DPP4, NRP1 and other proteins not identified by existing methods. The R package is available at: http://bioconductor.org/packages/release/bioc/html/GraphPAC.html. CONCLUSION: GraphPAC provides an alternative to iPAC and an extension to current methodology when identifying potential activating driver mutations by utilizing a graph theoretic approach when considering protein tertiary structure.
Gregory A. Ryslik, Yuwei Cheng, Kei-Hoi Cheung, Yorgo Modis, Hongyu Zhao 0003
BMC Bioinform.5
2014 Ttn as a likely causal gene for QTL of alcohol preference on mouse chromosome 2
abstract
Background Many quantitative trait loci (QTL) influencing mouse model phenotypes for alcoholism have been mapped genetically. However, the gene(s) comprising the QTL (QTG) are largely unknown. In previous work, Bennett and colleagues created congenic strains carrying the DBA/2IBG (D2) region for alcohol preference (AP) on chromosome 2, on a C57BL/6IBG (B6) background [1]. Subsequently, interval specific congenic recombinant strains (ISCRS), in which the full D2 QTL region was broken into smaller, partially overlapping regions of introgression, were generated and tested. With information from two ISCRS, the QTL has been mapped onto mouse chromosome 2 (Chr2) in a region of 3.4Mb by using C57BL/6J (B6) x DBA/2J (D2) recombinant inbred (RI) strains as well as by using F2 populations. Several candidate genes, Gad1, Atp5g3, Atf2, Sp3 and Sp9, have been evaluated but none of them has been confirmed for a definitive role in the regulation of the QTL of AP on Chr2 [2,3].
Lishi Wang, Yan Jiao, Beth Bennett, Robert W. Williams, Hongyu Zhao 0003, Joel Gelernter, Henry R. Kranzler, Lindsay A. Farrer, Weikuan Gu
BMC Bioinform.7
2013 Joint analysis of expression profiles from multiple cancers improves the identification of microRNA-gene interactions
abstract
MOTIVATION: MicroRNAs (miRNAs) play a crucial role in tumorigenesis and development through their effects on target genes. The characterization of miRNA-gene interactions will lead to a better understanding of cancer mechanisms. Many computational methods have been developed to infer miRNA targets with/without expression data. Because expression datasets are in general limited in size, most existing methods concatenate datasets from multiple studies to form one aggregated dataset to increase sample size and power. However, such simple aggregation analysis results in identifying miRNA-gene interactions that are mostly common across datasets, whereas specific interactions may be missed by these methods. Recent releases of The Cancer Genome Atlas data provide paired expression profiling of miRNAs and genes in multiple tumors with sufficiently large sample size. To study both common and cancer-specific interactions, it is desirable to develop a method that can jointly analyze multiple cancers to study miRNA-gene interactions without combining all the data into one single dataset. RESULTS: We developed a novel statistical method to jointly analyze expression profiles from multiple cancers to identify miRNA-gene interactions that are both common across cancers and specific to certain cancers. The benefit of this joint analysis approach is demonstrated by both simulation studies and real data analysis of The Cancer Genome Atlas datasets. Compared with simple aggregate analysis or single sample analysis, our method can effectively use the shared information among different but related cancers to improve the identification of miRNA-gene interactions. Another useful property of our method is that it can estimate similarity among cancers through their shared miRNA-gene interactions. AVAILABILITY AND IMPLEMENTATION: The program, MCMG, implemented in R is available at http://bioinformatics.med.yale.edu/group/.
Frank J. Slack, Hongyu Zhao 0003
Bioinform.3
2013 Accounting for non-genetic factors by low-rank representation and sparse regression for eQTL mapping
abstract
MOTIVATION: Expression quantitative trait loci (eQTL) studies investigate how gene expression levels are affected by DNA variants. A major challenge in inferring eQTL is that a number of factors, such as unobserved covariates, experimental artifacts and unknown environmental perturbations, may confound the observed expression levels. This may both mask real associations and lead to spurious association findings. RESULTS: In this article, we introduce a LOw-Rank representation to account for confounding factors and make use of Sparse regression for eQTL mapping (LORS). We integrate the low-rank representation and sparse regression into a unified framework, in which single-nucleotide polymorphisms and gene probes can be jointly analyzed. Given the two model parameters, our formulation is a convex optimization problem. We have developed an efficient algorithm to solve this problem and its convergence is guaranteed. We demonstrate its ability to account for non-genetic effects using simulation, and then apply it to two independent real datasets. Our results indicate that LORS is an effective tool to account for non-genetic effects. First, our detected associations show higher consistency between studies than recently proposed methods. Second, we have identified some new hotspots that can not be identified without accounting for non-genetic effects. AVAILABILITY: The software is available at: http://bioinformatics.med.yale.edu/software.aspx. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Can Yang 0002, Shuqin Zhang, Hongyu Zhao 0003
Bioinform.4
2013 Differential expression analysis for paired RNA-seq data
abstract
BACKGROUND: RNA-Seq technology measures the transcript abundance by generating sequence reads and counting their frequencies across different biological conditions. To identify differentially expressed genes between two conditions, it is important to consider the experimental design as well as the distributional property of the data. In many RNA-Seq studies, the expression data are obtained as multiple pairs, e.g., pre- vs. post-treatment samples from the same individual. We seek to incorporate paired structure into analysis. RESULTS: We present a Bayesian hierarchical mixture model for RNA-Seq data to separately account for the variability within and between individuals from a paired data structure. The method assumes a Poisson distribution for the data mixed with a gamma distribution to account variability between pairs. The effect of differential expression is modeled by two-component mixture model. The performance of this approach is examined by simulated and real data. CONCLUSIONS: In this setting, our proposed model provides higher sensitivity than existing methods to detect differential expression. Application to real RNA-Seq data demonstrates the usefulness of this method for detecting expression alteration for genes with low average expression levels or shorter transcript length.
Lisa M. Chung, John P. Ferguson, Vincent Bruno, Ruth R. Montgomery, Hongyu Zhao 0003
BMC Bioinform.7
2013 Utilizing protein structure to identify non-random somatic mutations
abstract
BACKGROUND: Human cancer is caused by the accumulation of somatic mutations in tumor suppressors and oncogenes within the genome. In the case of oncogenes, recent theory suggests that there are only a few key "driver" mutations responsible for tumorigenesis. As there have been significant pharmacological successes in developing drugs that treat cancers that carry these driver mutations, several methods that rely on mutational clustering have been developed to identify them. However, these methods consider proteins as a single strand without taking their spatial structures into account. We propose an extension to current methodology that incorporates protein tertiary structure in order to increase our power when identifying mutation clustering. RESULTS: We have developed iPAC (identification of Protein Amino acid Clustering), an algorithm that identifies non-random somatic mutations in proteins while taking into account the three dimensional protein structure. By using the tertiary information, we are able to detect both novel clusters in proteins that are known to exhibit mutation clustering as well as identify clusters in proteins without evidence of clustering based on existing methods. For example, by combining the data in the Protein Data Bank (PDB) and the Catalogue of Somatic Mutations in Cancer, our algorithm identifies new mutational clusters in well known cancer proteins such as KRAS and PI3KC α. Further, by utilizing the tertiary structure, our algorithm also identifies clusters in EGFR, EIF2AK2, and other proteins that are not identified by current methodology. The R package is available at: http://www.bioconductor.org/packages/2.12/bioc/html/iPAC.html. CONCLUSION: Our algorithm extends the current methodology to identify oncogenic activating driver mutations by utilizing tertiary protein structure when identifying nonrandom somatic residue mutation clusters.
Gregory A. Ryslik, Yuwei Cheng, Kei-Hoi Cheung, Yorgo Modis, Hongyu Zhao 0003
BMC Bioinform.5
2013 HapBoost: A Fast Approach to Boosting Haplotype Association Analyses in Genome-Wide Association Studies
abstract
Genome-wide association study (GWAS) has been successful in identifying genetic variants that are associated with complex human diseases. In GWAS, multilocus association analyses through linkage disequilibrium (LD), named haplotype-based analyses, may have greater power than single-locus analyses for detecting disease susceptibility loci. However, the large number of SNPs genotyped in GWAS poses great computational challenges in the detection of haplotype associations. We present a fast method named HapBoost for finding haplotype associations, which can be applied to quickly screen the whole genome. The effectiveness of HapBoost is demonstrated by using both synthetic and real data sets. The experimental results show that the proposed approach can achieve comparably accurate results while it performs much faster than existing methods.
Can Yang 0002, Qiang Yang 0001, Hongyu Zhao 0003, Weichuan Yu
IEEE ACM Trans. Comput. Biol. Bioinform.4
2013 Multisample aCGH Data Analysis via Total Variation and Spectral Regularization
abstract
DNA copy number variation (CNV) accounts for a large proportion of genetic variation. One commonly used approach to detecting CNVs is array-based comparative genomic hybridization (aCGH). Although many methods have been proposed to analyze aCGH data, it is not clear how to combine information from multiple samples to improve CNV detection. In this paper, we propose to use a matrix to approximate the multisample aCGH data and minimize the total variation of each sample as well as the nuclear norm of the whole matrix. In this way, we can make use of the smoothness property of each sample and the correlation among multiple samples simultaneously in a convex optimization framework. We also developed an efficient and scalable algorithm to handle large-scale data. Experiments demonstrate that the proposed method outperforms the state-of-the-art techniques under a wide range of scenarios and it is capable of processing large data sets with millions of probes.
Xiaowei Zhou 0001, Can Yang 0002, Hongyu Zhao 0003, Weichuan Yu
IEEE ACM Trans. Comput. Biol. Bioinform.4
2012 M3: an improved SNP calling algorithm for Illumina BeadArray data
abstract
SUMMARY: Genotype calling from high-throughput platforms such as Illumina and Affymetrix is a critical step in data processing, so that accurate information on genetic variants can be obtained for phenotype-genotype association studies. A number of algorithms have been developed to infer genotypes from data generated through the Illumina BeadStation platform, including GenCall, GenoSNP, Illuminus and CRLMM. Most of these algorithms are built on population-based statistical models to genotype every SNP in turn, such as GenCall with the GenTrain clustering algorithm, and require a large reference population to perform well. These approaches may not work well for rare variants where only a small proportion of the individuals carry the variant. A fundamentally different approach, implemented in GenoSNP, adopts a single nucleotide polymorphism (SNP)-based model to infer genotypes of all the SNPs in one individual, making it an appealing alternative to call rare variants. However, compared to the population-based strategies, more SNPs in GenoSNP may fail the Hardy-Weinberg Equilibrium test. To take advantage of both strategies, we propose a two-stage SNP calling procedure, named the modified mixture model (M(3)), to improve call accuracy for both common and rare variants. The effectiveness of our approach is demonstrated through applications to genotype calling on a set of HapMap samples used for quality control purpose in a large case-control study of cocaine dependence. The increase in power with M(3) is greater for rare variants than for common variants depending on the model. AVAILABILITY: M(3) algorithm: http://bioinformatics.med.yale.edu/group. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Gengxin Li, Joel Gelernter, Henry R. Kranzler, Hongyu Zhao 0003
Bioinform.4
2012 iFad: an integrative factor analysis model for drug-pathway association inference
abstract
MOTIVATION: Pathway-based drug discovery considers the therapeutic effects of compounds in the global physiological environment. This approach has been gaining popularity in recent years because the target pathways and mechanism of action for many compounds are still unknown, and there are also some unexpected off-target effects. Therefore, the inference of drug-pathway associations is a crucial step to fully realize the potential of system-based pharmacological research. Transcriptome data offer valuable information on drug-pathway targets because the pathway activities may be reflected through gene expression levels. Hence, it is of great interest to jointly analyze the drug sensitivity and gene expression data from the same set of samples to investigate the gene-pathway-drug-pathway associations. RESULTS: We have developed iFad, a Bayesian sparse factor analysis model to jointly analyze the paired gene expression and drug sensitivity datasets measured across the same panel of samples. The model enables direct incorporation of prior knowledge regarding gene-pathway and/or drug-pathway associations to aid the discovery of new association relationships. We use a collapsed Gibbs sampling algorithm for inference. Satisfactory performance of the proposed model was found for both simulated datasets and real data collected on the NCI-60 cell lines. Our results suggest that iFad is a promising approach for the identification of drug targets. This model also provides a general statistical framework for pathway-based integrative analysis of other types of -omics data. AVAILABILITY: The R package 'iFad' and real NCI-60 dataset used are available at http://bioinformatics.med.yale.edu/group.
Haisu Ma, Hongyu Zhao 0003
Bioinform.2
2012 FacPad: Bayesian sparse factor modeling for the inference of pathways responsive to drug treatment
abstract
MOTIVATION: It is well recognized that the effects of drugs are far beyond targeting individual proteins, but rather influencing the complex interactions among many relevant biological pathways. Genome-wide expression profiling before and after drug treatment has become a powerful approach for capturing a global snapshot of cellular response to drugs, as well as to understand drugs' mechanism of action. Therefore, it is of great interest to analyze this type of transcriptomic profiling data for the identification of pathways responsive to different drugs. However, few computational tools exist for this task. RESULTS: We have developed FacPad, a Bayesian sparse factor model, for the inference of pathways responsive to drug treatments. This model represents biological pathways as latent factors and aims to describe the variation among drug-induced gene expression alternations in terms of a much smaller number of latent factors. We applied this model to the Connectivity Map data set (build 02) and demonstrated that FacPad is able to identify many drug-pathway associations, some of which have been validated in the literature. Although this method was originally designed for the analysis of drug-induced transcriptional alternation data, it can be naturally applied to many other settings beyond polypharmacology. AVAILABILITY AND IMPLEMENTATION: The R package 'FacPad' is publically available at: http://cran.open-source-solution.org/web/packages/FacPad/.
Haisu Ma, Hongyu Zhao 0003
Bioinform.2
2012 Improved mean estimation and its application to diagonal discriminant analysis
abstract
MOTIVATION: High-dimensional data such as microarrays have created new challenges to traditional statistical methods. One such example is on class prediction with high-dimension, low-sample size data. Due to the small sample size, the sample mean estimates are usually unreliable. As a consequence, the performance of the class prediction methods using the sample mean may also be unsatisfactory. To obtain more accurate estimation of parameters some statistical methods, such as regularizations through shrinkage, are often desired. RESULTS: In this article, we investigate the family of shrinkage estimators for the mean value under the quadratic loss function. The optimal shrinkage parameter is proposed under the scenario when the sample size is fixed and the dimension is large. We then construct a shrinkage-based diagonal discriminant rule by replacing the sample mean by the proposed shrinkage mean. Finally, we demonstrate via simulation studies and real data analysis that the proposed shrinkage-based rule outperforms its original competitor in a wide range of settings.
Tiejun Tong, Hongyu Zhao 0003
Bioinform.3
2011 COSINE: COndition-SpecIfic sub-NEtwork identification using a global optimization method
abstract
MOTIVATION: The identification of condition specific sub-networks from gene expression profiles has important biological applications, ranging from the selection of disease-related biomarkers to the discovery of pathway alterations across different phenotypes. Although many methods exist for extracting these sub-networks, very few existing approaches simultaneously consider both the differential expression of individual genes and the differential correlation of gene pairs, losing potentially valuable information in the data. RESULTS: In this article, we propose a new method, COSINE (COndition SpecIfic sub-NEtwork), which employs a scoring function that jointly measures the condition-specific changes of both 'nodes' (individual genes) and 'edges' (gene-gene co-expression). It uses the genetic algorithm to search for the single optimal sub-network which maximizes the scoring function. We applied COSINE to both simulated datasets with various differential expression patterns, and three real datasets, one prostate cancer dataset, a second one from the across-tissue comparison of morbidly obese patients and the other from the across-population comparison of the HapMap samples. Compared with previous methods, COSINE is more powerful in identifying truly significant sub-networks of appropriate size and meaningful biological relevance. AVAILABILITY: The R code is available as the COSINE package on CRAN: http://cran.r-project.org/web/packages/COSINE/index.html.
Haisu Ma, Eric E. Schadt, Lee M. Kaplan, Hongyu Zhao 0003
Bioinform.4
2011 Score regularization for peptide identification
abstract
Peptide identification from tandem mass spectrometry (MS/MS) data is one of the most important problems in computational proteomics. This technique relies heavily on the accurate assessment of the quality of peptide-spectrum matches (PSMs). However, current MS technology and PSM scoring algorithm are far from perfect, leading to the generation of incorrect peptide-spectrum pairs. Thus, it is critical to develop new post-processing techniques that can distinguish true identifications from false identifications effectively. In this paper, we present a consistency-based PSM re-ranking method to improve the initial identification results. This method uses one additional assumption that two peptides belonging to the same protein should be correlated to each other. We formulate an optimization problem that embraces two objectives through regularization: the smoothing consistency among scores of correlated peptides and the fitting consistency between new scores and initial scores. This optimization problem can be solved analytically. The experimental study on several real MS/MS data sets shows that this re-ranking method improves the identification performance. The score regularization method can be used as a general post-processing step for improving peptide identifications. Source codes and data sets are available at: http://bioinformatics.ust.hk/SRPI.rar .
Zengyou He, Hongyu Zhao 0003, Weichuan Yu
BMC Bioinform.2
2011 Bias Detection and Correction in RNA-Sequencing Data
abstract
BACKGROUND: High throughput sequencing technology provides us unprecedented opportunities to study transcriptome dynamics. Compared to microarray-based gene expression profiling, RNA-Seq has many advantages, such as high resolution, low background, and ability to identify novel transcripts. Moreover, for genes with multiple isoforms, expression of each isoform may be estimated from RNA-Seq data. Despite these advantages, recent work revealed that base level read counts from RNA-Seq data may not be randomly distributed and can be affected by local nucleotide composition. It was not clear though how the base level read count bias may affect gene level expression estimates. RESULTS: In this paper, by using five published RNA-Seq data sets from different biological sources and with different data preprocessing schemes, we showed that commonly used estimates of gene expression levels from RNA-Seq data, such as reads per kilobase of gene length per million reads (RPKM), are biased in terms of gene length, GC content and dinucleotide frequencies. We directly examined the biases at the gene-level, and proposed a simple generalized-additive-model based approach to correct different sources of biases simultaneously. Compared to previously proposed base level correction methods, our method reduces bias in gene-level expression estimates more effectively. CONCLUSIONS: Our method identifies and corrects different sources of biases in gene-level expression measures from RNA-Seq data, and provides more accurate estimates of gene expression levels from RNA-Seq. This method should prove useful in meta-analysis of gene expression levels using different platforms or experimental protocols.
Lisa M. Chung, Hongyu Zhao 0003
BMC Bioinform.3
2010 Pathway analysis using random forests with bivariate node-split for survival outcomes
abstract
MOTIVATION: There is great interest in pathway-based methods for genomics data analysis in the research community. Although machine learning methods, such as random forests, have been developed to correlate survival outcomes with a set of genes, no study has assessed the abilities of these methods in incorporating pathway information for analyzing microarray data. In general, genes that are identified without incorporating biological knowledge are more difficult to interpret. Correlating pathway-based gene expression with survival outcomes may lead to biologically more meaningful prognosis biomarkers. Thus, a comprehensive study on how these methods perform in a pathway-based setting is warranted. RESULTS: In this article, we describe a pathway-based method using random forests to correlate gene expression data with survival outcomes and introduce a novel bivariate node-splitting random survival forests. The proposed method allows researchers to identify important pathways for predicting patient prognosis and time to disease progression, and discover important genes within those pathways. We compared different implementations of random forests with different split criteria and found that bivariate node-splitting random survival forests with log-rank test is among the best. We also performed simulation studies that showed random forests outperforms several other machine learning algorithms and has comparable results with a newly developed component-wise Cox boosting model. Thus, pathway-based survival analysis using machine learning tools represents a promising approach in dissecting pathways and for generating new biological hypothesis from microarray studies. AVAILABILITY: R package Pwayrfsurvival is available from URL: http://www.duke.edu/~hp44/pwayrfsurvival.htm. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Herbert Pang, Debayan Datta, Hongyu Zhao 0003
Bioinform.3
2009 Effect of false positive and false negative rates on inference of binding target conservation across different conditions and species from ChIP-chip data
abstract
BACKGROUND: ChIP-chip data are routinely used to identify transcription factor binding targets. However, the presence of false positives and false negatives in ChIP-chip data complicates and hinders analyses, especially when the binding targets for a specific transcription factor are compared across conditions or species. RESULTS: We propose an Expectation Maximization based approach to infer the underlying true counts of "positives" and "negatives" from the observed counts. Based on this approach, we study the effect of false positives and false negatives on inferences related to transcription regulation. CONCLUSION: Our results indicate that if there is a significant degree of association among the binding targets across conditions/species (log odds ratio > 4), moderate values of false positive and false negative rates (0.005 and 0.4 respectively) would not change our inference qualitatively (i.e. the presence or absence of conservation) based on the observed experimental data despite a significant change in the observed counts. However, if the underlying association is marginal, with odds ratios close to 1, moderate to large values of false positive and false negative rates (0.01 and 0.2 respectively) could mask the underlying association.
Debayan Datta, Hongyu Zhao 0003
BMC Bioinform.2
2009 Genotyping and inflated type I error rate in genome-wide association case/control studies
abstract
BACKGROUND: One common goal of a case/control genome wide association study (GWAS) is to find SNPs associated with a disease. Traditionally, the first step in such studies is to assign a genotype to each SNP in each subject, based on a statistic summarizing fluorescence measurements. When the distributions of the summary statistics are not well separated by genotype, the act of genotype assignment can lead to more potential problems than acknowledged by the literature. RESULTS: Specifically, we show that the proportions of each called genotype need not equal the true proportions in the population, even as the number of subjects grows infinitely large. The called genotypes for two subjects need not be independent, even when their true genotypes are independent. Consequently, p-values from tests of association can be anti-conservative, even when the distributions of the summary statistic for the cases and controls are identical. To address these problems, we propose two new tests designed to reduce the inflation in the type I error rate caused by these problems. The first algorithm, logiCALL, measures call quality by fully exploring the likelihood profile of intensity measurements, and the second algorithm avoids genotyping by using a likelihood ratio statistic. CONCLUSION: Genotyping can introduce avoidable false positives in GWAS.
Joshua N. Sampson, Hongyu Zhao 0003
BMC Bioinform.2
2008 Estimating dynamic models for gene regulation networks
abstract
MOTIVATION: Transcription regulation is a fundamental process in biology, and it is important to model the dynamic behavior of gene regulation networks. Many approaches have been proposed to specify the network structure. However, finding the network connectivity is not sufficient to understand the network dynamics. Instead, one needs to model the regulation reactions, usually with a set of ordinary differential equations (ODEs). Because some of the parameters involved in these ODEs are unknown, their values need to be inferred from the observed data. RESULTS: In this article, we introduce the generalized profiling method to estimate ODE parameters in a gene regulation network from microarray gene expression data which can be rather noisy. Because numerically solving ODEs is computationally expensive, we apply the penalized smoothing technique, a fast and stable computational method to approximate ODE solutions. The ODE solutions with our parameter estimates fit the data well. A goodness-of-fit test of dynamic models is developed to identify gene regulation networks.
Jiguo Cao, Hongyu Zhao 0003
Bioinform.2
2008 Considering dependence among genes and markers for false discovery control in eQTL mapping
abstract
MOTIVATION: Multiple comparison adjustment is a significant and challenging statistical issue in large-scale biological studies. In previous studies, dependence among genes is largely ignored. However, such dependence may be strong for some genomic-scale studies such as genetical genomics [also called expression quantitative trait loci (eQTL) mapping] in which thousands of genes are treated as quantitative traits and mapped to different genetical markers. Besides the dependence among markers, the dependence among the expression levels of genes can also have a significant impact on data analysis and interpretation. RESULTS: In this article, we propose to consider both the mean as well as the variance of false discovery number for multiple comparison adjustment to handle dependence among hypotheses. This is achieved by developing a variance estimator for false discovery number, and using the upper bound of false discovery proportion (uFDP) for false discovery control. More importantly, we introduce a weighted version of uFDP (wuFDP) control to improve the statistical power of eQTL identification. In addition, the wuFDP approach can better control false positives than false discovery rate (FDR) and uFDP approaches when markers are in linkage disequilibrium. The relative performance of uFDP control and wuFDP control is illustrated through simulation studies and real data analysis. SUPPLEMENTARY INFORMATION: Supplementary figures, tables and appendices are available at Bioinformatics online.
Tiejun Tong, Hongyu Zhao 0003
Bioinform.3
2008 Statistical methods to infer cooperative binding among transcription factors in Saccharomyces cerevisiae
abstract
MOTIVATION: Transcription factors regulate transcription in prokaryotes and eukaryotes by binding to specific DNA sequences in the regulatory regions of the genes. This regulation usually occurs in a coordinated manner involving multiple transcription factors. Genome-wide location data, also called ChIP-chip data, have enabled researchers to infer the binding sites for individual regulatory proteins. However, current methods to infer binding sites, such as simple thresholding based on p-values, are not optimal for a number of study objectives like combinatorial regulation, leading to potential loss of information. Hence, there is a need to develop more efficient statistical methods for analyzing such data. RESULTS: We propose to use log-linear models to study cooperative binding among transcription factors and have developed an Expectation-Maximization algorithm for statistical inferences. Our method is advantageous over simple thresholding methods both based on simulation and real data studies. We apply our method to infer the cooperative network of 204 regulators in Rich Medium and a subset of them in four different environmental conditions. Our results indicate that the cooperative network is condition specific; for a set of regulators, the network structure changes under different environmental conditions. AVAILABILITY: Our program is available at http://bioinformatics.med.yale.edu/TFcooperativity.
Debayan Datta, Hongyu Zhao 0003
Bioinform.2
2008 Building pathway clusters from Random Forests classification using class votes
abstract
BACKGROUND: Recent years have seen the development of various pathway-based methods for the analysis of microarray gene expression data. These approaches have the potential to bring biological insights into microarray studies. A variety of methods have been proposed to construct networks using gene expression data. Because individual pathways do not act in isolation, it is important to understand how different pathways coordinate to perform cellular functions. However, there are no published methods describing how to build pathway clusters that are closely related to traits of interest. RESULTS: We propose to build pathway clusters from pathway-based classification methods. The proposed methods allow researchers to identify clusters of pathways sharing similar functions. These pathways may or may not share genes. As an illustration, our approach is applied to three human breast cancer microarray data sets. We found that our methods yielded consistent and interpretable results for these three data sets. We further investigated one of the pathway clusters found using PubMatrix. We found that informative genes in the pathway clusters do have more publications with keywords, like estrogen receptor, compared with informative genes in other top pathways. In addition, using the shortest path analysis in GeneGo's MetaCore and Human Protein Reference Database, we were able to identify the links which connect the pathways without shared genes within the pathway cluster. CONCLUSION: Our proposed pathway clustering methods allow bioinformaticians and biologists to investigate how informative genes within pathways are related to each other and understand possible crosstalk between pathways in a cluster. Therefore, building pathway clusters may lead to a better understanding of molecular mechanisms affecting a trait of interest, and help generate further biological hypotheses from gene expression data.
Herbert Pang, Hongyu Zhao 0003
BMC Bioinform.2
2007 Integrating domain knowledge with statistical and data mining methods for high-density genomic SNP disease association analysis
Valentin Dinu, Hongyu Zhao 0003, Perry L. Miller
J. Biomed. Informatics2
2006 Pathway analysis using random forests classification and regression
abstract
MOTIVATION: Although numerous methods have been developed to better capture biological information from microarray data, commonly used single gene-based methods neglect interactions among genes and leave room for other novel approaches. For example, most classification and regression methods for microarray data are based on the whole set of genes and have not made use of pathway information. Pathway-based analysis in microarray studies may lead to more informative and relevant knowledge for biological researchers. RESULTS: In this paper, we describe a pathway-based classification and regression method using Random Forests to analyze gene expression data. The proposed methods allow researchers to rank important pathways from externally available databases, discover important genes, find pathway-based outlying cases and make full use of a continuous outcome variable in the regression setting. We also compared Random Forests with other machine learning methods using several datasets and found that Random Forests classification error rates were either the lowest or the second-lowest. By combining pathway information and novel statistical methods, this procedure represents a promising computational strategy in dissecting pathways and can provide biological insight into the study of microarray data. AVAILABILITY: Source code written in R is available from http://bioinformatics.med.yale.edu/pathway-analysis/rf.htm.
Herbert Pang, Aiping Lin, Matthew Holford, Bradley E. Enerson, Michael P. Lawton, Eugenia Floyd, Hongyu Zhao 0003
Bioinform.8
2006 PSMIX: an R package for population structure inference via maximum likelihood method
abstract
BACKGROUND: Inference of population stratification and individual admixture from genetic markers is an integrative part of a study in diverse situations, such as association mapping and evolutionary studies. Bayesian methods have been proposed for population stratification and admixture inference using multilocus genotypes and widely used in practice. However, these Bayesian methods demand intensive computation resources and may run into convergence problem in Markov Chain Monte Carlo based posterior samplings. RESULTS: We have developed PSMIX, an R package based on maximum likelihood method using expectation-maximization algorithm, for inference of population stratification and individual admixture. CONCLUSION: Compared with software based on Bayesian methods (e.g., STRUCTURE), PSMIX has similar accuracy, but more efficient computations.PSMIX and its supplemental documents are freely available at http://bioinformatics.med.yale.edu/PSMIX.
Baolin Wu, Nianjun Liu, Hongyu Zhao 0003
BMC Bioinform.3
2006 Multiple Peak Alignment in Sequential Data Analysis: A Scale-Space-Based Approach
abstract
In this paper, we address the multiple peak alignment problem in sequential data analysis with an approach based on the Gaussian scale-space theory. We assume that multiple sets of detected peaks are the observed samples of a set of common peaks. We also assume that the locations of the observed peaks follow unimodal distributions (e.g., normal distribution) with their means equal to the corresponding locations of the common peaks and variances reflecting the extension of their variations. Under these assumptions, we convert the problem of estimating locations of the unknown number of common peaks from multiple sets of detected peaks into a much simpler problem of searching for local maxima in the scale-space representation. The optimization of the scale parameter is achieved using an energy minimization approach. We compare our approach with a hierarchical clustering method using both simulated data and real mass spectrometry data. We also demonstrate the merit of extending the binary peak detection method (i.e., a candidate is considered either as a peak or as a nonpeak) with a quantitative scoring measure-based approach (i.e., we assign to each candidate a possibility of being a peak).
Weichuan Yu, Xiaoye Li, Baolin Wu, Kenneth R. Williams, Hongyu Zhao 0003
IEEE ACM Trans. Comput. Biol. Bioinform.6
2005 Negative correlation between compositional symmetries and local recombination rates
abstract
Although still not much understood, the universal reverse complement symmetry in genomes may contain much information about the genome. In this article, under the hypothesis that recombination rate variations may be related to the high order DNA structure, we studied the association between local recombination rates and local symmetry levels in mouse, rat and human. We found significant negative correlations between recombination rates and reverse complement compositional symmetries in these three organisms. This negative correlation pattern also held at individual chromosome levels when data only from each individual chromosome was analyzed.
Hongyu Zhao 0003
Bioinform.2
2005 A semiparametric approach for marker gene selection based on gene expression data
abstract
MOTIVATION: Identification of differentially expressed genes is a major issue in gene expression data analysis and selection of marker genes is critical in tumor classification using gene expression data. In this paper, we propose a semiparametric two-sample test to identify both differentially expressed genes and select marker genes for sample classification. RESULTS: A simulation study shows that the proposed method is more robust and powerful than the methods, generally used such as t-tests and non-parametric rank-sum tests, when the sample size is small. Cross-validation shows that the sample classification based on genes selected using this semiparametric method has lower misclassification rates. CONTACT: [email protected].
Hongyu Zhao 0003
Bioinform.2
2005 VitaPad: visualization tools for the analysis of pathway data
abstract
MOTIVATION: Packages that support the creation of pathway diagrams are limited by their inability to be readily extended to new classes of pathway-related data. RESULTS: VitaPad is a cross-platform application that enables users to create and modify biological pathway diagrams and incorporate microarray data with them. It improves on existing software in the following areas: (i) It can create diagrams dynamically through graph layout algorithms. (ii) It is open-source and uses an open XML format to store data, allowing for easy extension or integration with other tools. (iii) It features a cutting-edge user interface with intuitive controls, high-resolution graphics and fully customizable appearance. AVAILABILITY: http://bioinformatics.med.yale.edu CONTACTS: [email protected]; [email protected].
Matthew Holford, Naixin Li, Prakash M. Nadkarni, Hongyu Zhao 0003
Bioinform.4
2005 Detection of DNA copy number alterations using penalized least squares regression
abstract
MOTIVATION: Genomic DNA copy number alterations are characteristic of many human diseases including cancer. Various techniques and platforms have been proposed to allow researchers to partition the whole genome into segments where copy numbers change between contiguous segments, and subsequently to quantify DNA copy number alterations. In this paper, we incorporate the spatial dependence of DNA copy number data into a regression model and formalize the detection of DNA copy number alterations as a penalized least squares regression problem. In addition, we use a stationary bootstrap approach to estimate the statistical significance and false discovery rate. RESULTS: The proposed method is studied by simulations and illustrated by an application to an extensively analyzed dataset in the literature. The results show that the proposed method can correctly detect the numbers and locations of the true breakpoints while appropriately controlling the false positives. AVAILABILITY: http://bioinformatics.med.yale.edu/DNACopyNumber CONTACT: [email protected] SUPPLEMENTARY INFORMATION: http://bioinformatics.med.yale.edu/DNACopyNumber.
Baolin Wu, Paul Lizardi, Hongyu Zhao 0003
Bioinform.4
2005 Inferring protein-protein interactions through high-throughput interaction data from diverse organisms
abstract
MOTIVATION: Identifying protein-protein interactions is critical for understanding cellular processes. Because protein domains represent binding modules and are responsible for the interactions between proteins, computational approaches have been proposed to predict protein interactions at the domain level. The fact that protein domains are likely evolutionarily conserved allows us to pool information from data across multiple organisms for the inference of domain-domain and protein-protein interaction probabilities. RESULTS: We use a likelihood approach to estimating domain-domain interaction probabilities by integrating large-scale protein interaction data from three organisms, Saccharomyces cerevisiae, Caenorhabditis elegans and Drosophila melanogaster. The estimated domain-domain interaction probabilities are then used to predict protein-protein interactions in S.cerevisiae. Based on a thorough comparison of sensitivity and specificity, Gene Ontology term enrichment and gene expression profiles, we have demonstrated that it may be far more informative to predict protein-protein interactions from diverse organisms than from a single organism. AVAILABILITY: The program for computing the protein-protein interaction probabilities and supplementary material are available at http://bioinformatics.med.yale.edu/interaction.
Nianjun Liu, Hongyu Zhao 0003
Bioinform.3
2005 HAPLORE: a program for haplotype reconstruction in general pedigrees without recombination
abstract
MOTIVATION: Haplotype reconstruction is an essential step in genetic linkage and association studies. Although many methods have been developed to estimate haplotype frequencies and reconstruct haplotypes for a sample of unrelated individuals, haplotype reconstruction in large pedigrees with a large number of genetic markers remains a challenging problem. METHODS: We have developed an efficient computer program, HAPLORE (HAPLOtype REconstruction), to identify all haplotype sets that are compatible with the observed genotypes in a pedigree for tightly linked genetic markers. HAPLORE consists of three steps that can serve different needs in applications. In the first step, a set of logic rules is used to reduce the number of compatible haplotypes of each individual in the pedigree as much as possible. After this step, the haplotypes of all individuals in the pedigree can be completely or partially determined. These logic rules are applicable to completely linked markers and they can be used to impute missing data and check genotyping errors. In the second step, a haplotype-elimination algorithm similar to the genotype-elimination algorithms used in linkage analysis is applied to delete incompatible haplotypes derived from the first step. All superfluous haplotypes of the pedigree members will be excluded after this step. In the third step, the expectation-maximization (EM) algorithm combined with the partition and ligation technique is used to estimate haplotype frequencies based on the inferred haplotype configurations through the first two steps. Only compatible haplotype configurations with haplotypes having frequencies greater than a threshold are retained. RESULTS: We test the effectiveness and the efficiency of HAPLORE using both simulated and real datasets. Our results show that, the rule-based algorithm is very efficient for completely genotyped pedigree. In this case, almost all of the families have one unique haplotype configuration. In the presence of missing data, the number of compatible haplotypes can be substantially reduced by HAPLORE, and the program will provide all possible haplotype configurations of a pedigree under different circumstances, if such multiple configurations exist. These inferred haplotype configurations, as well as the haplotype frequencies estimated by the EM algorithm, can be used in genetic linkage and association studies. AVAILABILITY: The program can be downloaded from http://bioinformatics.med.yale.edu.
Fengzhu Sun, Hongyu Zhao 0003
Bioinform.3
2005 Are scale-free networks robust to measurement errors?
abstract
BACKGROUND: Many complex random networks have been found to be scale-free. Existing literature on scale-free networks has rarely considered potential false positive and false negative links in the observed networks, especially in biological networks inferred from high-throughput experiments. Therefore, it is important to study the impact of these measurement errors on the topology of the observed networks. RESULTS: This article addresses the impact of erroneous links on network topological inference and explores possible error mechanisms for scale-free networks with an emphasis on Saccharomyces cerevisiae protein interaction networks. We study this issue by both theoretical derivations and simulations. We show that the ignorance of erroneous links in network analysis may lead to biased estimates of the scale parameter and recommend robust estimators in such scenarios. Possible error mechanisms of yeast protein interaction networks are explored by comparisons between real data and simulated data. CONCLUSION: Our studies show that, in the presence of erroneous links, the connectivity distribution of scale-free networks is still scale-free for the middle range connectivities, but can be greatly distorted for low and high connecitivities. It is more appropriate to use robust estimators such as the least trimmed mean squares estimator to estimate the scale parameter gamma under such circumstances. Moreover, we show by simulation studies that the scale-free property is robust to some error mechanisms but untenable to others. The simulation results also suggest that different error mechanisms may be operating in the yeast protein interaction networks produced from different data sources. In the MIPS gold standard protein interaction data, there appears to be a high rate of false negative links, and the false negative and false positive rates are more or less constant across proteins with different connectivities. However, the error mechanism of yeast two-hybrid data may be very different, where the overall false negative rate is low and the false negative rates tend to be higher for links involving proteins with more interacting partners.
Hongyu Zhao 0003
BMC Bioinform.2
2005 Case Report: A High Productivity/Low Maintenance Approach to High-performance Computation for Biomedicine: Four Case Studies
abstract
The rapid advances in high-throughput biotechnologies such as DNA microarrays and mass spectrometry have generated vast amounts of data ranging from gene expression to proteomics data. The large size and complexity involved in analyzing such data demand a significant amount of computing power. High-performance computation (HPC) is an attractive and increasingly affordable approach to help meet this challenge. There is a spectrum of techniques that can be used to achieve computational speedup with varying degrees of impact in terms of how drastic a change is required to allow the software to run on an HPC platform. This paper describes a high- productivity/low-maintenance (HP/LM) approach to HPC that is based on establishing a collaborative relationship between the bioinformaticist and HPC expert that respects the former's codes and minimizes the latter's efforts. The goal of this approach is to make it easy for bioinformatics researchers to continue to make iterative refinements to their programs, while still being able to take advantage of HPC. The paper describes our experience applying these HP/LM techniques in four bioinformatics case studies: (1) genome-wide sequence comparison using Blast, (2) identification of biomarkers based on statistical analysis of large mass spectrometry data sets, (3) complex genetic analysis involving ordinal phenotypes, (4) large-scale assessment of the effect of possible errors in analyzing microarray data. The case studies illustrate how the HP/LM approach can be applied to a range of representative bioinformatics applications and how the approach can lead to significant speedup of computationally intensive bioinformatics applications, while making only modest modifications to the programs themselves.
Nicholas Carriero, Michael V. Osier, Kei-Hoi Cheung, Perry L. Miller, Mark Gerstein, Hongyu Zhao 0003, Baolin Wu, Scott A. Rifkin, Joseph T. Chang, Heping Zhang, Kevin P. White, Kenneth R. Williams, Martin H. Schultz
J. Am. Medical Informatics Assoc.6
2005 Integrating mRNA Decay Information into Co-Regulation Study
Hongyu Zhao 0003
J. Comput. Sci. Technol.2
2004 A statistical method for identifying differential gene-gene co-expression patterns
abstract
MOTIVATION: To understand cancer etiology, it is important to explore molecular changes in cellular processes from normal state to cancerous state. Because genes interact with each other during cellular processes, carcinogenesis related genes may form differential co-expression patterns with other genes in different cell states. In this study, we develop a statistical method for identifying differential gene-gene co-expression patterns in different cell states. RESULTS: For efficient pattern recognition, we extend the traditional F-statistic and obtain an Expected Conditional F-statistic (ECF-statistic), which incorporates statistical information of location and correlation. We also propose a statistical method for data transformation. Our approach is applied to a microarray gene expression dataset for prostate cancer study. For a gene of interest, our method can select other genes that have differential gene-gene co-expression patterns with this gene in different cell states. The 10 most frequently selected genes, include hepsin, GSTP1 and AMACR, which have recently been proposed to be associated with prostate carcinogenesis. However, genes GSTP1 and AMACR cannot be identified by studying differential gene expression alone. By using tumor suppressor genes TP53, PTEN and RB1, we identify seven genes that also include hepsin, GSTP1 and AMACR. We show that genes associated with cancer may have differential gene-gene expression patterns with many other genes in different cell states. By discovering such patterns, we may be able to identify carcinogenesis related genes.
Yinglei Lai, Baolin Wu, Hongyu Zhao 0003
Bioinform.4
2004 Information assessment on predicting protein-protein interactions
abstract
BACKGROUND: Identifying protein-protein interactions is fundamental for understanding the molecular machinery of the cell. Proteome-wide studies of protein-protein interactions are of significant value, but the high-throughput experimental technologies suffer from high rates of both false positive and false negative predictions. In addition to high-throughput experimental data, many diverse types of genomic data can help predict protein-protein interactions, such as mRNA expression, localization, essentiality, and functional annotation. Evaluations of the information contributions from different evidences help to establish more parsimonious models with comparable or better prediction accuracy, and to obtain biological insights of the relationships between protein-protein interactions and other genomic information. RESULTS: Our assessment is based on the genomic features used in a Bayesian network approach to predict protein-protein interactions genome-wide in yeast. In the special case, when one does not have any missing information about any of the features, our analysis shows that there is a larger information contribution from the functional-classification than from expression correlations or essentiality. We also show that in this case alternative models, such as logistic regression and random forest, may be more effective than Bayesian networks for predicting interactions. CONCLUSIONS: In the restricted problem posed by the complete-information subset, we identified that the MIPS and Gene Ontology (GO) functional similarity datasets as the dominating information contributors for predicting the protein-protein interactions under the framework proposed by Jansen et al. Random forests based on the MIPS and GO information alone can give highly accurate classifications. In this particular subset of complete information, adding other genomic data does little for improving predictions. We also found that the data discretizations used in the Bayesian methods decreased classification performance.
Baolin Wu, Ronald Jansen, Mark Gerstein, Hongyu Zhao 0003
BMC Bioinform.5
2004 A computational approach for ordering signal transduction pathway components from genomics and proteomics Data
abstract
BACKGROUND: Signal transduction is one of the most important biological processes by which cells convert an external signal into a response. Novel computational approaches to mapping proteins onto signaling pathways are needed to fully take advantage of the rapid accumulation of genomic and proteomics information. However, despite their importance, research on signaling pathways reconstruction utilizing large-scale genomics and proteomics information has been limited. RESULTS: We have developed an approach for predicting the order of signaling pathway components, assuming all the components on the pathways are known. Our method is built on a score function that integrates protein-protein interaction data and microarray gene expression data. Compared to the individual datasets, either protein interactions or gene transcript abundance measurements, the integrated approach leads to better identification of the order of the pathway components. CONCLUSIONS: As demonstrated in our study on the yeast MAPK signaling pathways, the integration analysis of high-throughput genomics and proteomics data can be a powerful means to infer the order of pathway components, enabling the transformation from molecular data into knowledge of cellular mechanisms.
Hongyu Zhao 0003
BMC Bioinform.2
2004 Handling multiple testing while interpreting microarrays with the Gene Ontology Database
abstract
BACKGROUND: The development of software tools that analyze microarray data in the context of genetic knowledgebases is being pursued by multiple research groups using different methods. A common problem for many of these tools is how to correct for multiple statistical testing since simple corrections are overly conservative and more sophisticated corrections are currently impractical. A careful study of the nature of the distribution one would expect by chance, such as by a simulation study, may be able to guide the development of an appropriate correction that is not overly time consuming computationally. RESULTS: We present the results from a preliminary study of the distribution one would expect for analyzing sets of genes extracted from Drosophila, S. cerevisiae, Wormbase, and Gramene databases using the Gene Ontology Database. CONCLUSIONS: We found that the estimated distribution is not regular and is not predictable outside of a particular set of genes. Permutation-based simulations may be necessary to determine the confidence in results of such analyses.
Michael V. Osier, Hongyu Zhao 0003, Kei-Hoi Cheung
BMC Bioinform.2
2003 Comparison of statistical methods for classification of ovarian cancer using mass spectrometry data
abstract
MOTIVATION: Novel methods, both molecular and statistical, are urgently needed to take advantage of recent advances in biotechnology and the human genome project for disease diagnosis and prognosis. Mass spectrometry (MS) holds great promise for biomarker identification and genome-wide protein profiling. It has been demonstrated in the literature that biomarkers can be identified to distinguish normal individuals from cancer patients using MS data. Such progress is especially exciting for the detection of early-stage ovarian cancer patients. Although various statistical methods have been utilized to identify biomarkers from MS data, there has been no systematic comparison among these approaches in their relative ability to analyze MS data. RESULTS: We compare the performance of several classes of statistical methods for the classification of cancer based on MS spectra. These methods include: linear discriminant analysis, quadratic discriminant analysis, k-nearest neighbor classifier, bagging and boosting classification trees, support vector machine, and random forest (RF). The methods are applied to ovarian cancer and control serum samples from the National Ovarian Cancer Early Detection Program clinic at Northwestern University Hospital. We found that RF outperforms other methods in the analysis of MS data.
Baolin Wu, Tom Abbott, David Fishman, Walter McMurray, Gil Mor, Kathryn L. Stone, David Ward, Kenneth R. Williams, Hongyu Zhao 0003
Bioinform.9
2003 PathMAPA: a tool for displaying gene expression and performing statistical tests on metabolic pathways at multiple levels for Arabidopsis
abstract
BACKGROUND: To date, many genomic and pathway-related tools and databases have been developed to analyze microarray data. In published web-based applications to date, however, complex pathways have been displayed with static image files that may not be up-to-date or are time-consuming to rebuild. In addition, gene expression analyses focus on individual probes and genes with little or no consideration of pathways. These approaches reveal little information about pathways that are key to a full understanding of the building blocks of biological systems. Therefore, there is a need to provide useful tools that can generate pathways without manually building images and allow gene expression data to be integrated and analyzed at pathway levels for such experimental organisms as Arabidopsis. RESULTS: We have developed PathMAPA, a web-based application written in Java that can be easily accessed over the Internet. An Oracle database is used to store, query, and manipulate the large amounts of data that are involved. PathMAPA allows its users to (i) upload and populate microarray data into a database; (ii) integrate gene expression with enzymes of the pathways; (iii) generate pathway diagrams without building image files manually; (iv) visualize gene expressions for each pathway at enzyme, locus, and probe levels; and (v) perform statistical tests at pathway, enzyme and gene levels. PathMAPA can be used to examine Arabidopsis thaliana gene expression patterns associated with metabolic pathways. CONCLUSION: PathMAPA provides two unique features for the gene expression analysis of Arabidopsis thaliana: (i) automatic generation of pathways associated with gene expression and (ii) statistical tests at pathway level. The first feature allows for the periodical updating of genomic data for pathways, while the second feature can provide insight into how treatments affect relevant pathways for the selected experiment(s).
Deyun Pan, Kei-Hoi Cheung, Ligeng Ma, Matthew Holford, Xingwang Deng, Hongyu Zhao 0003
BMC Bioinform.8
2002 YMD: a microarray database for large-scale gene expression analysis
Kei-Hoi Cheung, Kevin P. White, Janet Hager, Mark Gerstein, Valerie Reinke, Kenneth Nelson, Peter Masiar, Ranjana Srivastava, Yuli Li, Hongyu Zhao 0003, David B. Allison, Michael Snyder 0001, Perry L. Miller, Kenneth R. Williams
AMIA11