Zhaohui S. Qin

dblp:52/1149 · DBLP profile ↗
← Back
29ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-1583-146XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 29 · 4 first-author · 6 since 2021
YearPublicationVenuePosition
2026 A novel machine learning-based algorithm for eQTL identification reveals complex pleiotropic effects in the MHC region
abstract
Abstract Expression quantitative trait loci (eQTLs) are regulatory variants that affect the expression level of their target genes and have significant impact on disease biology. However, eQTL mapping has been done mostly in one tissue at a time, despite the known prevalence of correlations among tissues. Multivariate analyses incorporating multiple phenotypes are available, but they emphasize linear combinations of phenotypes. We present MTClass, a machine learning framework that attempts to classify an individual’s genotype based on a vector of multi-phenotype expression levels of a given gene. We conduct simulation studies and multiple case studies using real and imputed data, and we demonstrate that MTClass detects more functionally relevant variants and genes compared to existing single-tissue approaches as well as multi-phenotype association tests. Our results suggest that the importance of expression regulation at the MHC region may have been underestimated, and they provide fresh biological insights into genetic variants that have pleiotropic effects, influencing gene expression in a complex manner. Key points MTClass is a machine learning-based approach that classifies genotypes based on multi-phenotype expression data, providing a novel method for identifying eQTLs. MTClass outperforms traditional linear methods like MultiPhen and MANOVA in detecting eQTLs with greater functional impact and in capturing complex genotype-phenotype relationships. MTClass identified immune-related variants in the HLA region, suggesting that existing approaches may have underestimated the complexity of these variants’ effects across tissues. MTClass is more flexible and reliable than linear multivariate methods, handling multicollinearity, zero-expressed features, and various input values with greater ease.
Ronnie Y. Li, Zhaohui S. Qin
Briefings Bioinform.3
2025 A review on knowledge graphs for healthcare: Resources, applications, and promises
Hejie Cui, Jiaying Lu 0001, Ran Xu 0002, Shiyu Wang 0002, Wenjing Ma, Yue Yu 0001, Shaojun Yu, Xuan Kan, Chen Ling 0003, Liang Zhao 0002, Zhaohui S. Qin, Joyce C. Ho, Tianfan Fu, Jing Ma 0005, Mengdi Huai, Carl Yang 0001
J. Biomed. Informatics11
2024 A novel classification framework for genome-wide association study of whole brain MRI images using deep learning
abstract
Genome-wide association studies (GWASs) have been widely applied in the neuroimaging field to discover genetic variants associated with brain-related traits. So far, almost all GWASs conducted in neuroimaging genetics are performed on univariate quantitative features summarized from brain images. On the other hand, powerful deep learning technologies have dramatically improved our ability to classify images. In this study, we proposed and implemented a novel machine learning strategy for systematically identifying genetic variants that lead to detectable nuances on Magnetic Resonance Images (MRI). For a specific single nucleotide polymorphism (SNP), if MRI images labeled by genotypes of this SNP can be reliably distinguished using machine learning, we then hypothesized that this SNP is likely to be associated with brain anatomy or function which is manifested in MRI brain images. We applied this strategy to a catalog of MRI image and genotype data collected by the Alzheimer's Disease Neuroimaging Initiative (ADNI) consortium. From the results, we identified novel variants that show strong association to brain phenotypes.
Shaojun Yu, Yumeng Shao, Deqiang Qiu, Zhaohui S. Qin
PLoS Comput. Biol.5
2022 Disease category-specific annotation of variants using an ensemble learning framework
abstract
Understanding the impact of non-coding sequence variants on complex diseases is an essential problem. We present a novel ensemble learning framework-CASAVA, to predict genomic loci in terms of disease category-specific risk. Using disease-associated variants identified by GWAS as training data, and diverse sequencing-based genomics and epigenomics profiles as features, CASAVA provides risk prediction of 24 major categories of diseases throughout the human genome. Our studies showed that CASAVA scores at a genomic locus provide a reasonable prediction of the disease-specific and disease category-specific risk prediction for non-coding variants located within the locus. Taking MHC2TA and immune system diseases as an example, we demonstrate the potential of CASAVA in revealing variant-disease associations. A website (http://zhanglabtools.org/CASAVA) has been built to facilitate easily access to CASAVA scores.
Zhaohui S. Qin
Briefings Bioinform.5
2022 LRcell: detecting the source of differential expression at the sub-cell-type level from bulk RNA-seq data
abstract
Given most tissues are consist of abundant and diverse (sub-)cell types, an important yet unaddressed problem in bulk RNA-seq analysis is to identify at which (sub-)cell type(s) the differential expression occurs. Single-cell RNA-sequencing (scRNA-seq) technologies can answer the question, but they are often labor-intensive and cost-prohibitive. Here, we present LRcell, a computational method aiming to identify specific (sub-)cell type(s) that drives the changes observed in a bulk RNA-seq experiment. In addition, LRcell provides pre-embedded marker genes computed from putative scRNA-seq experiments as options to execute the analyses. We conduct a simulation study to demonstrate the effectiveness and reliability of LRcell. Using three different real datasets, we show that LRcell successfully identifies known cell types involved in psychiatric disorders. Applying LRcell to bulk RNA-seq results can produce a hypothesis on which (sub-)cell type(s) contributes to the differential expression. LRcell is complementary to cell type deconvolution methods.
Wenjing Ma, Sumeet Sharma, Shannon L. Gourley, Zhaohui S. Qin
Briefings Bioinform.5
2022 Guest Editorial for Selected Papers From BIOKDD 2021
abstract
The papers in this special section were presented at the 20th International Workshop on Data Mining in Bioinformatics (BIOKDD 2021) that was held virtually on August 15, 2021. The conference featured the special theme of “Artificial Intelligence in Medicine” which particularly welcomed paper submissions and invited talks related to the use of machine learning and data mining techniques for the analysis of large amounts of heterogeneous complex biological and medical data, with a particular focus on deep learning methods that see fast advancement and wider adoption in Bioinformatics.
Da Yan 0001, Zhaohui S. Qin, Debswapna Bhattacharya, Jake Yue Chen
IEEE ACM Trans. Comput. Biol. Bioinform.2
2020 A comprehensive review of computational prediction of genome-wide features
abstract
There are significant correlations among different types of genetic, genomic and epigenomic features within the genome. These correlations make the in silico feature prediction possible through statistical or machine learning models. With the accumulation of a vast amount of high-throughput data, feature prediction has gained significant interest lately, and a plethora of papers have been published in the past few years. Here we provide a comprehensive review on these published works, categorized by the prediction targets, including protein binding site, enhancer, DNA methylation, chromatin structure and gene expression. We also provide discussions on some important points and possible future directions.
Tianlei Xu, Xiaoqi Zheng, Zhaohui S. Qin, Hao Wu 0003
Briefings Bioinform.5
2020 Regulatory annotation of genomic intervals based on tissue-specific expression QTLs
abstract
MOTIVATION: Annotating a given genomic locus or a set of genomic loci is an important yet challenging task. This is especially true for the non-coding part of the genome which is enormous yet poorly understood. Since gene set enrichment analyses have demonstrated to be effective approach to annotate a set of genes, the same idea can be extended to explore the enrichment of functional elements or features in a set of genomic intervals to reveal potential functional connections. RESULTS: In this study, we describe a novel computational strategy named loci2path that takes advantage of the newly emerged, genome-wide and tissue-specific expression quantitative trait loci (eQTL) information to help annotate a set of genomic intervals in terms of transcription regulation. By checking the presence or the absence of millions of eQTLs in a set of input genomic intervals, combined with grouping eQTLs by the pathways or gene sets that their target genes belong to, loci2path build a bridge connecting genomic intervals to functional pathways and pre-defined biological-meaningful gene sets, revealing potential for regulatory connection. Our method enjoys two key advantages over existing methods: first, we no longer rely on proximity to link a locus to a gene which has shown to be unreliable; second, eQTL allows us to provide the regulatory annotation under the context of specific tissue types. To demonstrate its utilities, we apply loci2path on sets of genomic intervals harboring disease-associated variants as query. Using 1 702 612 eQTLs discovered by the Genotype-Tissue Expression (GTEx) project across 44 tissues and 6320 pathways or gene sets cataloged in MSigDB as annotation resource, our method successfully identifies highly relevant biological pathways and revealed disease mechanisms for psoriasis and other immune-related diseases. Tissue specificity analysis of associated eQTLs provide additional evidence of the distinct roles of different tissues played in the disease mechanisms. AVAILABILITY AND IMPLEMENTATION: loci2path is published as an open source Bioconductor package, and it is available at http://bioconductor.org/packages/release/bioc/html/loci2path.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tianlei Xu, Zhaohui S. Qin
Bioinform.3
2020 Application of topic models to a compendium of ChIP-Seq datasets uncovers recurrent transcriptional regulatory modules
abstract
MOTIVATION: The availability of thousands of genome-wide coupling chromatin immunoprecipitation (ChIP)-Seq datasets across hundreds of transcription factors (TFs) and cell lines provides an unprecedented opportunity to jointly analyze large-scale TF-binding in vivo, making possible the discovery of the potential interaction and cooperation among different TFs. The interacted and cooperated TFs can potentially form a transcriptional regulatory module (TRM) (e.g. co-binding TFs), which helps decipher the combinatorial regulatory mechanisms. RESULTS: We develop a computational method tfLDA to apply state-of-the-art topic models to multiple ChIP-Seq datasets to decipher the combinatorial binding events of multiple TFs. tfLDA is able to learn high-order combinatorial binding patterns of TFs from multiple ChIP-Seq profiles, interpret and visualize the combinatorial patterns. We apply the tfLDA to two cell lines with a rich collection of TFs and identify combinatorial binding patterns that show well-known TRMs and related TF co-binding events. AVAILABILITY AND IMPLEMENTATION: A software R package tfLDA is freely available at https://github.com/lichen-lab/tfLDA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Aiqun Ma, Zhaohui S. Qin, Li Chen 0029
Bioinform.3
2020 Proceedings of the 2019 MidSouth Computational Biology and Bioinformatics Society (MCBIOS) Conference
abstract
The 16th Annual MidSouth Computational Biology and Bioinformatics Society (MCBIOS XVI) conference was held in Hilton Birmingham at University of Alabama Birmingham (UAB) conference center on March 28–30, 2019. The theme of the conference was Informatics for Precision Medicine. The co-chairs and conference hosts were Drs. Jake Y. Chen and Matthew Might from the University of Alabama Birmingham.
Jonathan D. Wren, Yongsheng Bai, Zhaohui S. Qin, Da Yan 0001, Ramin Homayouni
BMC Bioinform.3
2020 Inferring Spatial Organization of Individual Topologically Associated Domains via Piecewise Helical Model
abstract
The recently developed Hi-C technology enables a genome-wide view of chromosome spatial organizations, and has shed deep insights into genome structure and genome function. However, multiple sources of uncertainties make downstream data analysis and interpretation challenging. Specifically, statistical models for inferring three-dimensional (3D) chromosomal structure from Hi-C data are far from their maturity. Most existing methods are highly over-parameterized, lacking clear interpretations, and sensitive to outliers. In this study, we propose a parsimonious, easy to interpret, and robust piecewise helical model for the inference of 3D chromosomal structure of individual topologically associated domain from Hi-C data. When applied to a real Hi-C dataset, the piecewise helical model not only achieves much better model fitting than existing models, but also reveals that geometric properties of chromatin spatial organization are closely related to genome function.
Ming Hu 0001, Zhaohui S. Qin, Jun S. Liu
IEEE ACM Trans. Comput. Biol. Bioinform.4
2019 RT States: systematic annotation of the human genome using cell type-specific replication timing programs
abstract
MOTIVATION: The replication timing (RT) program has been linked to many key biological processes including cell fate commitment, 3D chromatin organization and transcription regulation. Significant technology progress now allows to characterize the RT program in the entire human genome in a high-throughput and high-resolution fashion. These experiments suggest that RT changes dynamically during development in coordination with gene activity. Since RT is such a fundamental biological process, we believe that an effective quantitative profile of the local RT program from a diverse set of cell types in various developmental stages and lineages can provide crucial biological insights for a genomic locus. RESULTS: In this study, we explored recurrent and spatially coherent combinatorial profiles from 42 RT programs collected from multiple lineages at diverse differentiation states. We found that a Hidden Markov Model with 15 hidden states provide a good model to describe these genome-wide RT profiling data. Each of the hidden state represents a unique combination of RT profiles across different cell types which we refer to as 'RT states'. To understand the biological properties of these RT states, we inspected their relationship with chromatin states, gene expression, functional annotation and 3D chromosomal organization. We found that the newly defined RT states possess interesting genome-wide functional properties that add complementary information to the existing annotation of the human genome. AVAILABILITY AND IMPLEMENTATION: R scripts for inferring HMM models and Perl scripts for further analysis are available https://github.com/PouletAxel/script_HMM_Replication_timing. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Axel Poulet, Tristan Dubos, Juan Carlos Rivera-Mulia, David M. Gilbert, Zhaohui S. Qin
Bioinform.6
2016 traseR: an R package for performing trait-associated SNP enrichment analysis in genomic intervals
abstract
UNLABELLED: Genome-wide association studies (GWASs) have successfully identified many sequence variants that are significantly associated with common diseases and traits. Tens of thousands of such trait-associated SNPs have already been cataloged, which we believe form a great resource for genomic research. Recent studies have demonstrated that the collection of trait-associated SNPs can be exploited to indicate whether a given genomic interval or intervals are likely to be functionally connected with certain phenotypes or diseases. Despite this importance, currently, there is no ready-to-use computational tool able to connect genomic intervals to phenotypes. Here, we present traseR, an easy-to-use R Bioconductor package that performs enrichment analyses of trait-associated SNPs in arbitrary genomic intervals with flexible options, including testing method, type of background and inclusion of SNPs in LD. AVAILABILITY AND IMPLEMENTATION: The traseR R package preloaded with up-to-date collection of trait-associated SNPs are freely available in Bioconductor CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Li Chen 0029, Zhaohui S. Qin
Bioinform.2
2016 Bayesian inference with historical data-based informative priors improves detection of differentially expressed genes
abstract
MOTIVATION: Modern high-throughput biotechnologies such as microarray are capable of producing a massive amount of information for each sample. However, in a typical high-throughput experiment, only limited number of samples were assayed, thus the classical 'large p, small n' problem. On the other hand, rapid propagation of these high-throughput technologies has resulted in a substantial collection of data, often carried out on the same platform and using the same protocol. It is highly desirable to utilize the existing data when performing analysis and inference on a new dataset. RESULTS: Utilizing existing data can be carried out in a straightforward fashion under the Bayesian framework in which the repository of historical data can be exploited to build informative priors and used in new data analysis. In this work, using microarray data, we investigate the feasibility and effectiveness of deriving informative priors from historical data and using them in the problem of detecting differentially expressed genes. Through simulation and real data analysis, we show that the proposed strategy significantly outperforms existing methods including the popular and state-of-the-art Bayesian hierarchical model-based approaches. Our work illustrates the feasibility and benefits of exploiting the increasingly available genomics big data in statistical inference and presents a promising practical strategy for dealing with the 'large p, small n' problem. AVAILABILITY AND IMPLEMENTATION: Our method is implemented in R package IPBT, which is freely available from https://github.com/benliemory/IPBT CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhaonan Sun, Zhaohui S. Qin
Bioinform.5
2016 A hidden Markov random field-based Bayesian method for the detection of long-range chromosomal interactions in Hi-C data
abstract
MOTIVATION: Advances in chromosome conformation capture and next-generation sequencing technologies are enabling genome-wide investigation of dynamic chromatin interactions. For example, Hi-C experiments generate genome-wide contact frequencies between pairs of loci by sequencing DNA segments ligated from loci in close spatial proximity. One essential task in such studies is peak calling, that is, detecting non-random interactions between loci from the two-dimensional contact frequency matrix. Successful fulfillment of this task has many important implications including identifying long-range interactions that assist interpreting a sizable fraction of the results from genome-wide association studies. The task - distinguishing biologically meaningful chromatin interactions from massive numbers of random interactions - poses great challenges both statistically and computationally. Model-based methods to address this challenge are still lacking. In particular, no statistical model exists that takes the underlying dependency structure into consideration. RESULTS: In this paper, we propose a hidden Markov random field (HMRF) based Bayesian method to rigorously model interaction probabilities in the two-dimensional space based on the contact frequency matrix. By borrowing information from neighboring loci pairs, our method demonstrates superior reproducibility and statistical power in both simulation studies and real data analysis. AVAILABILITY AND IMPLEMENTATION: The Source codes can be downloaded at: http://www.unc.edu/∼yunmli/HMRFBayesHiC CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zheng Xu 0010, Fulai Jin, Mengjie Chen, Terrence S. Furey, Patrick F. Sullivan, Zhaohui S. Qin, Ming Hu 0001
Bioinform.7
2015 A novel statistical method for quantitative comparison of multiple ChIP-seq datasets
abstract
MOTIVATION: ChIP-seq is a powerful technology to measure the protein binding or histone modification strength in the whole genome scale. Although there are a number of methods available for single ChIP-seq data analysis (e.g. 'peak detection'), rigorous statistical method for quantitative comparison of multiple ChIP-seq datasets with the considerations of data from control experiment, signal to noise ratios, biological variations and multiple-factor experimental designs is under-developed. RESULTS: In this work, we develop a statistical method to perform quantitative comparison of multiple ChIP-seq datasets and detect genomic regions showing differential protein binding or histone modification. We first detect peaks from all datasets and then union them to form a single set of candidate regions. The read counts from IP experiment at the candidate regions are assumed to follow Poisson distribution. The underlying Poisson rates are modeled as an experiment-specific function of artifacts and biological signals. We then obtain the estimated biological signals and compare them through the hypothesis testing procedure in a linear model framework. Simulations and real data analyses demonstrate that the proposed method provides more accurate and robust results compared with existing ones. AVAILABILITY AND IMPLEMENTATION: An R software package ChIPComp is freely available at http://web1.sph.emory.edu/users/hwu30/software/ChIPComp.html.
Li Chen 0029, Chi Wang 0003, Zhaohui S. Qin, Hao Wu 0003
Bioinform.3
2015 One Size Doesn't Fit All - RefEditor: Building Personalized Diploid Reference Genome to Improve Read Mapping and Genotype Calling in Next Generation Sequencing Studies
abstract
With rapid decline of the sequencing cost, researchers today rush to embrace whole genome sequencing (WGS), or whole exome sequencing (WES) approach as the next powerful tool for relating genetic variants to human diseases and phenotypes. A fundamental step in analyzing WGS and WES data is mapping short sequencing reads back to the reference genome. This is an important issue because incorrectly mapped reads affect the downstream variant discovery, genotype calling and association analysis. Although many read mapping algorithms have been developed, the majority of them uses the universal reference genome and do not take sequence variants into consideration. Given that genetic variants are ubiquitous, it is highly desirable if they can be factored into the read mapping procedure. In this work, we developed a novel strategy that utilizes genotypes obtained a priori to customize the universal haploid reference genome into a personalized diploid reference genome. The new strategy is implemented in a program named RefEditor. When applying RefEditor to real data, we achieved encouraging improvements in read mapping, variant discovery and genotype calling. Compared to standard approaches, RefEditor can significantly increase genotype calling consistency (from 43% to 61% at 4X coverage; from 82% to 92% at 20X coverage) and reduce Mendelian inconsistency across various sequencing depths. Because many WGS and WES studies are conducted on cohorts that have been genotyped using array-based genotyping platforms previously or concurrently, we believe the proposed strategy will be of high value in practice, which can also be applied to the scenario where multiple NGS experiments are conducted on the same cohort. The RefEditor sources are available at https://github.com/superyuan/refeditor.
H. Richard Johnston, Yi-Juan Hu, Zhaohui S. Qin
PLoS Comput. Biol.6
2013 Sparsely correlated hidden Markov models with application to genome-wide location studies
abstract
MOTIVATION: Multiply correlated datasets have become increasingly common in genome-wide location analysis of regulatory proteins and epigenetic modifications. Their correlation can be directly incorporated into a statistical model to capture underlying biological interactions, but such modeling quickly becomes computationally intractable. RESULTS: We present sparsely correlated hidden Markov models (scHMM), a novel method for performing simultaneous hidden Markov model (HMM) inference for multiple genomic datasets. In scHMM, a single HMM is assumed for each series, but the transition probability in each series depends on not only its own hidden states but also the hidden states of other related series. For each series, scHMM uses penalized regression to select a subset of the other data series and estimate their effects on the odds of each transition in the given series. Following this, hidden states are inferred using a standard forward-backward algorithm, with the transition probabilities adjusted by the model at each position, which helps retain the order of computation close to fitting independent HMMs (iHMM). Hence, scHMM is a collection of inter-dependent non-homogeneous HMMs, capable of giving a close approximation to a fully multivariate HMM fit. A simulation study shows that scHMM achieves comparable sensitivity to the multivariate HMM fit at a much lower computational cost. The method was demonstrated in the joint analysis of 39 histone modifications, CTCF and RNA polymerase II in human CD4+ T cells. scHMM reported fewer high-confidence regions than iHMM in this dataset, but scHMM could recover previously characterized histone modifications in relevant genomic regions better than iHMM. In addition, the resulting combinatorial patterns from scHMM could be better mapped to the 51 states reported by the multivariate HMM method of Ernst and Kellis. AVAILABILITY: The scHMM package can be freely downloaded from http://sourceforge.net/p/schmm/ and is recommended for use in a linux environment.
Hyungwon Choi, Damian Fermin, Alexey I. Nesvizhskii, Debashis Ghosh, Zhaohui S. Qin
Bioinform.5
2013 Bayesian Inference of Spatial Organizations of Chromosomes
abstract
Knowledge of spatial chromosomal organizations is critical for the study of transcriptional regulation and other nuclear processes in the cell. Recently, chromosome conformation capture (3C) based technologies, such as Hi-C and TCC, have been developed to provide a genome-wide, three-dimensional (3D) view of chromatin organization. Appropriate methods for analyzing these data and fully characterizing the 3D chromosomal structure and its structural variations are still under development. Here we describe a novel Bayesian probabilistic approach, denoted as "Bayesian 3D constructor for Hi-C data" (BACH), to infer the consensus 3D chromosomal structure. In addition, we describe a variant algorithm BACH-MIX to study the structural variations of chromatin in a cell population. Applying BACH and BACH-MIX to a high resolution Hi-C dataset generated from mouse embryonic stem cells, we found that most local genomic regions exhibit homogeneous 3D chromosomal structures. We further constructed a model for the spatial arrangement of chromatin, which reveals structural properties associated with euchromatic and heterochromatic regions in the genome. We observed strong associations between structural properties and several genomic and epigenetic features of the chromosome. Using BACH-MIX, we further found that the structural variations of chromatin are correlated with these genomic and epigenetic features. Our results demonstrate that BACH and BACH-MIX have the potential to provide new insights into the chromosomal architecture of mammalian cells.
Ming Hu 0001, Zhaohui S. Qin, Jesse R. Dixon, Siddarth Selvaraj, Jennifer Fang, Jun S. Liu
PLoS Comput. Biol.3
2012 HiCNorm: removing biases in Hi-C data via Poisson regression
abstract
SUMMARY: We propose a parametric model, HiCNorm, to remove systematic biases in the raw Hi-C contact maps, resulting in a simple, fast, yet accurate normalization procedure. Compared with the existing Hi-C normalization method developed by Yaffe and Tanay, HiCNorm has fewer parameters, runs >1000 times faster and achieves higher reproducibility. AVAILABILITY: Freely available on the web at: http://www.people.fas.harvard.edu/∼junliu/HiCNorm/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ming Hu 0001, Siddarth Selvaraj, Zhaohui S. Qin, Jun S. Liu
Bioinform.4
2012 Using Poisson mixed-effects model to quantify transcript-level gene expression in RNA-Seq
abstract
MOTIVATION: RNA sequencing (RNA-Seq) is a powerful new technology for mapping and quantifying transcriptomes using ultra high-throughput next-generation sequencing technologies. Using deep sequencing, gene expression levels of all transcripts including novel ones can be quantified digitally. Although extremely promising, the massive amounts of data generated by RNA-Seq, substantial biases and uncertainty in short read alignment pose challenges for data analysis. In particular, large base-specific variation and between-base dependence make simple approaches, such as those that use averaging to normalize RNA-Seq data and quantify gene expressions, ineffective. RESULTS: In this study, we propose a Poisson mixed-effects (POME) model to characterize base-level read coverage within each transcript. The underlying expression level is included as a key parameter in this model. Since the proposed model is capable of incorporating base-specific variation as well as between-base dependence that affect read coverage profile throughout the transcript, it can lead to improved quantification of the true underlying expression level. AVAILABILITY AND IMPLEMENTATION: POME can be freely downloaded at http://www.stat.purdue.edu/~yuzhu/pome.html. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ming Hu 0001, Jeremy M. G. Taylor, Jun S. Liu, Zhaohui S. Qin
Bioinform.5
2010 HPeak: an HMM-based algorithm for defining read-enriched regions in ChIP-Seq data
abstract
BACKGROUND: Protein-DNA interaction constitutes a basic mechanism for the genetic regulation of target gene expression. Deciphering this mechanism has been a daunting task due to the difficulty in characterizing protein-bound DNA on a large scale. A powerful technique has recently emerged that couples chromatin immunoprecipitation (ChIP) with next-generation sequencing, (ChIP-Seq). This technique provides a direct survey of the cistrom of transcription factors and other chromatin-associated proteins. In order to realize the full potential of this technique, increasingly sophisticated statistical algorithms have been developed to analyze the massive amount of data generated by this method. RESULTS: Here we introduce HPeak, a Hidden Markov model (HMM)-based Peak-finding algorithm for analyzing ChIP-Seq data to identify protein-interacting genomic regions. In contrast to the majority of available ChIP-Seq analysis software packages, HPeak is a model-based approach allowing for rigorous statistical inference. This approach enables HPeak to accurately infer genomic regions enriched with sequence reads by assuming realistic probability distributions, in conjunction with a novel weighting scheme on the sequencing read coverage. CONCLUSIONS: Using biologically relevant data collections, we found that HPeak showed a higher prevalence of the expected transcription factor binding motifs in ChIP-enriched sequences relative to the control sequences when compared to other currently available ChIP-Seq analysis approaches. Additionally, in comparison to the ChIP-chip assay, ChIP-Seq provides higher resolution along with improved sensitivity and specificity of binding site detection. Additional file and the HPeak program are freely available at http://www.sph.umich.edu/csg/qin/HPeak.
Zhaohui S. Qin, Jianjun Yu, Jincheng Shen, Christopher A. Maher, Ming Hu 0001, Shanker Kalyana-Sundaram, Jindan Yu, Arul M. Chinnaiyan
BMC Bioinform.1
2009 Hierarchical hidden Markov model with application to joint analysis of ChIP-chip and ChIP-seq data
abstract
MOTIVATION: Chromatin immunoprecipitation (ChIP) experiments followed by array hybridization, or ChIP-chip, is a powerful approach for identifying transcription factor binding sites (TFBS) and has been widely used. Recently, massively parallel sequencing coupled with ChIP experiments (ChIP-seq) has been increasingly used as an alternative to ChIP-chip, offering cost-effective genome-wide coverage and resolution up to a single base pair. For many well-studied TFs, both ChIP-seq and ChIP-chip experiments have been applied and their data are publicly available. Previous analyses have revealed substantial technology-specific binding signals despite strong correlation between the two sets of results. Therefore, it is of interest to see whether the two data sources can be combined to enhance the detection of TFBS. RESULTS: In this work, hierarchical hidden Markov model (HHMM) is proposed for combining data from ChIP-seq and ChIP-chip. In HHMM, inference results from individual HMMs in ChIP-seq and ChIP-chip experiments are summarized by a higher level HMM. Simulation studies show the advantage of HHMM when data from both technologies co-exist. Analysis of two well-studied TFs, NRSF and CCCTC-binding factor (CTCF), also suggests that HHMM yields improved TFBS identification in comparison to analyses using individual data sources or a simple merger of the two. AVAILABILITY: Source code for the software ChIPmeta is freely available for download at http://www.umich.edu/~hwchoi/HHMMsoftware.zip, implemented in C and supported on linux.
Hyungwon Choi, Alexey I. Nesvizhskii, Debashis Ghosh, Zhaohui S. Qin
Bioinform.4
2008 Data Analysis and Graphics Using R: An Example-Based Approach, Second Edition.: John Maindonald and John Braun
abstract
A volume in the ‘Cambridge Series in Statistical and Probabilistic Mathematics’, ‘Data Analysis and Graphics Using R’ is presented as a gentle tour guide for new R users, aiming to help them navigate through many powerful tools that the open source R system offers. As the authors point out in the Preface, the book is ‘aimed at scientists who wish to do statistical analysis on their own’. Every effort has been made to ensure the book is useful for practical data analysis. This book is particularly useful for researchers in the life science realm who have limited exposure to statistical methodology or theory. The R system, an open source, free software package for data analysis and graphics, has gained substantial popularity among statisticians over the years. Unlike commercial software packages, R is specifically designed such that it is easy for regular users to create contributed packages and share them with other users. The openness generated surprisingly rich resources in all areas of statistics and beyond. The R system has evolved into an all-purpose software platform for scientific computing and graphics. Researchers in genomics and bioinformatics fields routinely face challenges in analyzing large datasets generated from high-throughput assays such as DNA microarray, genotyping and sequencing technologies. The Bioconductor package, built on top of R, offers a great variety of tools for analyzing -omics data. However, current books on Bioconductor require knowledge of the R system. This book can serve as an introductory guide or reference for these users, especially researchers who are only interested in using R as a data analysis tool. There are 14 chapters in this book. The first chapter provides a brief introduction to the R system. Topics include defining a variable, reading in data from an external file. This chapter also introduces basic R programming, such as writing functions, loops. The chapter ends with a brief introduction to R graphics tools. Chapter 2 illustrates basic functions in R for data visualization and summary, and how to use them to address basic data analysis questions. Chapter 3 introduces basic concepts in statistical models including probability distributions and random samples. Chapter 4 provides a rather intuitive yet accurate introduction on statistical inference including point estimation, confidence intervals construction and hypothesis testing. There is also a starred section that briefly survey basic theories behind maximum likelihood estimate and Bayesian inference. The next four chapters discuss basic regression models including simple linear regression, multiple linear regression, ANOVA, generalized linear models and basic survival models. Practical issues such as model fitting assessment, model assumption diagnostics and model comparison are discussed in detail. Although the authors started from the most basic models, they also briefly covered advanced topics such as polynomial regression and smoothing splines techniques in nonlinear and nonparametric models. Chapter 9 covers basic time series models. Chapter 10 offers a detailed introduction of random effect and multi-level models, experimental design and repeated measure methodology. This chapter may be of particular interest to biomedical researchers who conduct clinical studies since the models discussed in this chapter are commonly used in clinical data analysis. Chapter 11 provides a nice introduction to tree-based supervised learning techniques. Powerful tools such as regression tree and classification tree are explained in detail. Chapter 12 and 13 cover multivariate data analysis. Important techniques such as principle component analysis, multi-dimensional scaling and discriminant analysis are described. The authors also survey graphical representation tools that are useful for exploring multi-dimensional data. The final chapter discusses practical issues using the R software and provides more details about its powerful graphical display environment. Instead of the regular references, the authors provide separate reference lists for methodology, datasets, R packages and websites. And in addition to regular term and author index, the authors also provide an index of R symbols and functions. These resources are a great help for R users who can use the book as a reference guide for finding information they need fast. The authors should be complimented for their attention-to-details writing style. Each chapter is ended with a Recap section to summarize the material covered; a Further Reading section, which listed reference for special, in-depth and advanced topics; an Exercise section for readers to practice using the R tools discussed. Solutions to selected exercises, together with R scripts, codes, figures and additional notes can be found at a website maintained by the authors. As the R system is rapidly evolving (a new version every half year), it is important to keep the content of the book up-to-date. It seems that the authors are aware of this issue. They have made updates from the previous edition and are planning for a new edition. At the mean time, I think it would be great for the readers if the authors can provide continued and detailed updates on their website. Another place the book can be improved is the organization of the many datasets used. There are so many different datasets discussed in this book, while a great strength, I found it difficult to keep track of them. It would be nice if an index of all dataset as well as some background information of them can be provided. Such that the users can quickly identify an example dataset of their interest or a dataset that resembles the one on their hand. The material presented in this volume made it an ideal textbook for an applied statistics course designed for students who need training in data analysis. Given the comprehensiveness, this book can also serve as a reference for general R users. Although most of the examples in the book are not directly relevant for scientists in genomics and bioinformatics fields, I do think they will benefit from the examples and in-depth coverage of essential data analysis techniques and methods.
Zhaohui S. Qin
Briefings Bioinform.1
2007 CRCView: a web server for analyzing and visualizing microarray gene expression data using model-based clustering
abstract
UNLABELLED: CRCView is a user-friendly point-and-click web server for analyzing and visualizing microarray gene expression data using a Dirichlet process mixture model-based clustering algorithm. CRCView is designed to clustering genes based on their expression profiles. It allows flexible input data format, rich graphical illustration as well as integrated GO term based annotation/interpretation of clustering results. AVAILABILITY: http://helab.bioinformatics.med.umich.edu/crcview/.
Zuoshuang Xiang, Zhaohui S. Qin, Yongqun He
Bioinform.2
2006 Clustering microarray gene expression data using weighted Chinese restaurant process
abstract
MOTIVATION: Clustering microarray gene expression data is a powerful tool for elucidating co-regulatory relationships among genes. Many different clustering techniques have been successfully applied and the results are promising. However, substantial fluctuation contained in microarray data, lack of knowledge on the number of clusters and complex regulatory mechanisms underlying biological systems make the clustering problems tremendously challenging. RESULTS: We devised an improved model-based Bayesian approach to cluster microarray gene expression data. Cluster assignment is carried out by an iterative weighted Chinese restaurant seating scheme such that the optimal number of clusters can be determined simultaneously with cluster assignment. The predictive updating technique was applied to improve the efficiency of the Gibbs sampler. An additional step is added during reassignment to allow genes that display complex correlation relationships such as time-shifted and/or inverted to be clustered together. Analysis done on a real dataset showed that as much as 30% of significant genes clustered in the same group display complex relationships with the consensus pattern of the cluster. Other notable features including automatic handling of missing data, quantitative measures of cluster strength and assignment confidence. Synthetic and real microarray gene expression datasets were analyzed to demonstrate its performance. AVAILABILITY: A computer program named Chinese restaurant cluster (CRC) has been developed based on this algorithm. The program can be downloaded at http://www.sph.umich.edu/csg/qin/CRC/.
Zhaohui S. Qin
Bioinform.1
2006 An efficient comprehensive search algorithm for tagSNP selection using linkage disequilibrium criteria
abstract
MOTIVATION: Selecting SNP markers for genome-wide association studies is an important and challenging task. The goal is to minimize the number of markers selected for genotyping in a particular platform and therefore reduce genotyping cost while simultaneously maximizing the information content provided by selected markers. RESULTS: We devised an improved algorithm for tagSNP selection using the pairwise r(2) criterion. We first break down large marker sets into disjoint pieces, where more exhaustive searches can replace the greedy algorithm for tagSNP selection. These exhaustive searches lead to smaller tagSNP sets being generated. In addition, our method evaluates multiple solutions that are equivalent according to the linkage disequilibrium criteria to accommodate additional constraints. Its performance was assessed using HapMap data. AVAILABILITY: A computer program named FESTA has been developed based on this algorithm. The program is freely available and can be downloaded at http://www.sph.umich.edu/csg/qin/FESTA/
Zhaohui S. Qin, Shyam Gopalakrishnan, Gonçalo R. Abecasis
Bioinform.1
2005 HapBlock: haplotype block partitioning and tag SNP selection software using a set of dynamic programming algorithms
abstract
UNLABELLED: Recent studies have revealed that linkage disequilibrium (LD) patterns vary across the human genome with some regions of high LD interspersed with regions of low LD. Such LD patterns make it possible to select a set of single nucleotide polymorphism (SNPs; tag SNPs) for genome-wide association studies. We have developed a suite of computer programs to analyze the block-like LD patterns and to select the corresponding tag SNPs. Compared to other programs for haplotype block partitioning and tag SNP selection, our program has several notable features. First, the dynamic programming algorithms implemented are guaranteed to find the block partition with minimum number of tag SNPs for the given criteria of blocks and tag SNPs. Second, both haplotype data and genotype data from unrelated individuals and/or from general pedigrees can be analyzed. Third, several existing measures/criteria for haplotype block partitioning and tag SNP selection have been implemented in the program. Finally, the programs provide flexibility to include specific SNPs (e.g. non-synonymous SNPs) as tag SNPs. AVAILABILITY: The HapBlock program and its supplemental documents can be downloaded from the website http://www.cmb.usc.edu/~msms/HapBlock.
Zhaohui S. Qin, Ting Chen 0006, Jun S. Liu, Michael S. Waterman, Fengzhu Sun
Bioinform.2
2005 Structural comparison of metabolic networks in selected single cell organisms
abstract
BACKGROUND: There has been tremendous interest in the study of biological network structure. An array of measurements has been conceived to assess the topological properties of these networks. In this study, we compared the metabolic network structures of eleven single cell organisms representing the three domains of life using these measurements, hoping to find out whether the intrinsic network design principle(s), reflected by these measurements, are different among species in the three domains of life. RESULTS: Three groups of topological properties were used in this study: network indices, degree distribution measures and motif profile measure. All of which are higher-level topological properties except for the marginal degree distribution. Metabolic networks in Archaeal species are found to be different from those in S. cerevisiae and the six Bacterial species in almost all measured higher-level topological properties. Our findings also indicate that the metabolic network in Archaeal species is similar to the exponential random network. CONCLUSION: If these metabolic network properties of the organisms studied can be extended to other species in their respective domains (which is likely), then the design principle(s) of Archaea are fundamentally different from those of Bacteria and Eukaryote. Furthermore, the functional mechanisms of Archaeal metabolic networks revealed in this study differentiate significantly from those of Bacterial and Eukaryotic organisms, which warrant further investigation.
Dongxiao Zhu, Zhaohui S. Qin
BMC Bioinform.2