Hao Wu 0003

dblp:72/4250-3 · DBLP profile ↗
← Back
27ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0003-1269-7354ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 27 · 2 first-author · 12 since 2021
YearPublicationVenuePosition
2026 CEMUSA: a graph-based integrative metric for evaluating clusters in spatial transcriptomics
abstract
MOTIVATION: Spatial clustering is a critical analytical task in spatial transcriptomics (ST) that aids in uncovering the spatial molecular mechanisms underlying biological phenotypes. Along with the numerous spatial clustering methods, there comes the imperative need for an effective metric to evaluate their performance. An ideal metric should consider three factors: label agreement, spatial organization, and error severity. However, existing evaluation metrics focus solely on either label agreement or spatial organization, leading to biased and misleading evaluations. RESULTS: To fill this gap, we propose CEMUSA, a novel graph-based metric that integrates these factors into a unified evaluation framework. Extensive testing on both simulated and real datasets demonstrate CEMUSA's superiority over conventional metrics in differentiating clustering results with subtle differences in topology and error severity, while maintaining computational efficiency. AVAILABILITY AND IMPLEMENTATION: The source code and data are freely available at https://github.com/YihDu/CEMUSA. CEMUSA is implemented as an R package at https://yihdu.github.io/CEMUSA.
Jiaying Hu, Yihang Du, Suyang Hou, Yueyang Ding, Hao Wu 0003
Bioinform.6
2025 scDETECT: a novel statistical model accounting for cell type correlation in single-cell RNA-seq differential expression analysis
abstract
Differential expression (DE) is one of the most important analyses in single-cell RNA-seq (scRNA-seq). Due to similarity of cell types, the DE states often have strong correlation among different cell types. Existing methods perform DE analysis for each cell type separately and ignore such correlation, leading to low accuracy, and statistical power. We develop single cell Differential Expression TEst with Cell Type correlation (scDETECT), a novel statistical method, for scRNA-seq DE analysis accounting for the cell type correlations. scDETECT implements a Bayesian hierarchical model to incorporate the cell type correlations into the modeling of the gene expression, and then the DE genes are called based on the derived posterior probabilities. Simulation and real data studies show that scDETECT significantly improves the accuracy and statistical power compared with existing methods.
Yuhan Xu, Hao Wu 0003
Briefings Bioinform.3
2024 Target-oriented Reference Construction for supervised cell type identification in scRNA-seq
abstract
Cell type identification is a crucial step in single-cell RNA-seq (scRNA-seq) data analysis. The supervised cell type identification method is a preferred solution due to its accuracy and efficiency. The performance of such methods highly depends on the quality of the reference data. Although there are many supervised cell type identification tools, no method currently exists for constructing reference data. Here, we develop Target-Oriented Reference Construction (TORC), a widely applicable strategy for constructing references based on target datasets in scRNA-seq supervised cell type identification. TORC focuses on alleviating differences in cell type composition between the reference and target sets. Extensive benchmarks on simulated and real data analyses demonstrate consistent improvements in cell type identification with TORC.
Wenjing Ma, Zhijin Wu, Hao Wu 0003
BIBM4
2024 SCIntRuler: guiding the integration of multiple single-cell RNA-seq datasets with a novel statistical metric
abstract
MOTIVATION: The growing number of single-cell RNA-seq (scRNA-seq) studies highlights the potential benefits of integrating multiple datasets, such as augmenting sample sizes and enhancing analytical robustness. Inherent diversity and batch discrepancies within samples or across studies continue to pose significant challenges for computational analyses. Questions persist in practice, lacking definitive answers: Should we use a specific integration method or opt for simply merging the datasets during joint analysis? Among all the existing data integration methods, which one is more suitable in specific scenarios? RESULT: To fill the gap, we introduce SCIntRuler, a novel statistical metric for guiding the integration of multiple scRNA-seq datasets. SCIntRuler helps researchers make informed decisions regarding the necessity of data integration and the selection of an appropriate integration method. Our simulations and real data applications demonstrate that SCIntRuler streamlines decision-making processes and facilitates the analysis of diverse scRNA-seq datasets under varying contexts, thereby alleviating the complexities associated with the integration of heterogeneous scRNA-seq datasets. AVAILABILITY AND IMPLEMENTATION: The implementation of our method is available on CRAN as an open-source R package with a user-friendly manual available: https://cloud.r-project.org/web/packages/SCIntRuler/index.html.
Yue Lyu, Steven H. Lin, Hao Wu 0003, Ziyi Li 0001
Bioinform.3
2023 A cofunctional grouping-based approach for non-redundant feature gene selection in unannotated single-cell RNA-seq analysis
abstract
Feature gene selection has significant impact on the performance of cell clustering in single-cell RNA sequencing (scRNA-seq) analysis. A well-rounded feature selection (FS) method should consider relevance, redundancy and complementarity of the features. Yet most existing FS methods focus on gene relevance to the cell types but neglect redundancy and complementarity, which undermines the cell clustering performance. We develop a novel computational method GeneClust to select feature genes for scRNA-seq cell clustering. GeneClust groups genes based on their expression profiles, then selects genes with the aim of maximizing relevance, minimizing redundancy and preserving complementarity. It can work as a plug-in tool for FS with any existing cell clustering method. Extensive benchmark results demonstrate that GeneClust significantly improve the clustering performance. Moreover, GeneClust can group cofunctional genes in biological process and pathway into clusters, thus providing a means of investigating gene interactions and identifying potential genes relevant to biological characteristics of the dataset. GeneClust is freely available at https://github.com/ToryDeng/scGeneClust.
Siyu Chen 0033, Yuanbin Xu, Da Feng, Hao Wu 0003
Briefings Bioinform.6
2022 A comprehensive comparison of supervised and unsupervised methods for cell type identification in single-cell RNA-seq
abstract
The cell type identification is among the most important tasks in single-cell RNA-sequencing (scRNA-seq) analysis. Many in silico methods have been developed and can be roughly categorized as either supervised or unsupervised. In this study, we investigated the performances of 8 supervised and 10 unsupervised cell type identification methods using 14 public scRNA-seq datasets of different tissues, sequencing protocols and species. We investigated the impacts of a number of factors, including total amount of cells, number of cell types, sequencing depth, batch effects, reference bias, cell population imbalance, unknown/novel cell type, and computational efficiency and scalability. Instead of merely comparing individual methods, we focused on factors' impacts on the general category of supervised and unsupervised methods. We found that in most scenarios, the supervised methods outperformed the unsupervised methods, except for the identification of unknown cell types. This is particularly true when the supervised methods use a reference dataset with high informational sufficiency, low complexity and high similarity to the query dataset. However, such outperformance could be undermined by some undesired dataset properties investigated in this study, which lead to uninformative and biased reference datasets. In these scenarios, unsupervised methods could be comparable to supervised methods. Our study not only explained the cell typing methods' behaviors under different experimental settings but also provided a general guideline for the choice of method according to the scientific goal and dataset properties. Finally, our evaluation workflow is implemented as a modularized R pipeline that allows future evaluation of new methods. Availability: All the source codes are available at https://github.com/xsun28/scRNAIdent.
Xiaochu Lin, Ziyi Li 0001, Hao Wu 0003
Briefings Bioinform.4
2022 Differential RNA methylation analysis for MeRIP-seq data under general experimental design
abstract
MOTIVATION: RNA epigenetics is an emerging field to study the post-transcriptional gene regulation. The dynamics of RNA epigenetic modification have been reported to associate with many human diseases. Recently developed high-throughput technology named Methylated RNA Immunoprecipitation Sequencing (MeRIP-seq) enables the transcriptome-wide profiling of N6-methyladenosine (m6A) modification and comparison of RNA epigenetic modifications. There are a few computational methods for the comparison of mRNA modifications under different conditions but they all suffer from serious limitations. RESULTS: In this work, we develop a novel statistical method to detect differentially methylated mRNA regions from MeRIP-seq data. We model the sequence count data by a hierarchical negative binomial model that accounts for various sources of variations and derive parameter estimation and statistical testing procedures for flexible statistical inferences under general experimental designs. Extensive benchmark evaluations in simulation and real data analyses demonstrate that our method is more accurate, robust and flexible compared to existing methods. AVAILABILITY AND IMPLEMENTATION: Our method TRESS is implemented as an R/Bioconductor package and is available at https://bioconductor.org/packages/devel/TRESS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhenxing Guo 0001, Andrew M. Shafik, Hao Wu 0003
Bioinform.4
2022 EDClust: an EM-MM hybrid method for cell clustering in multiple-subject single-cell RNA sequencing
abstract
MOTIVATION: Single-cell RNA sequencing (scRNA-seq) has revolutionized biological research by enabling the measurement of transcriptomic profiles at the single-cell level. With the increasing application of scRNA-seq in larger-scale studies, the problem of appropriately clustering cells emerges when the scRNA-seq data are from multiple subjects. One challenge is the subject-specific variation; systematic heterogeneity from multiple subjects may have a significant impact on clustering accuracy. Existing methods seeking to address such effects suffer from several limitations. RESULTS: We develop a novel statistical method, EDClust, for multi-subject scRNA-seq cell clustering. EDClust models the sequence read counts by a mixture of Dirichlet-multinomial distributions and explicitly accounts for cell-type heterogeneity, subject heterogeneity and clustering uncertainty. An EM-MM hybrid algorithm is derived for maximizing the data likelihood and clustering the cells. We perform a series of simulation studies to evaluate the proposed method and demonstrate the outstanding performance of EDClust. Comprehensive benchmarking on four real scRNA-seq datasets with various tissue types and species demonstrates the substantial accuracy improvement of EDClust compared to existing methods. AVAILABILITY AND IMPLEMENTATION: The R package is freely available at https://github.com/weix21/EDClust. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ziyi Li 0001, Hongkai Ji, Hao Wu 0003
Bioinform.4
2021 A novel virtual drug screening pipeline with deep-leaning as core component identifies inhibitor of pancreatic alpha-amylase
abstract
Virtual drug screening that provides possible drug candidates facilitates early-stage drug discovery. It works by large scale predicting native-like protein-ligand complexes (PLC) from an abundance of docking decoys. Many affinity predicting models currently in use fail to provide reliable prediction because of a lack of non-binding data during model training, lost critical physical-chemical features, and difficulties in learning abstract information with limited neural layers. In this paper, we developed a deep learning model, DeepBindBC for classifying putative ligands as binding or non-binding. Our model incorporates information of non-binding interactions, making it more suitable for real applications. ResNet model architecture and more detailed atom type representation guarantee implicit features can be learned more accurately. DeepBindBC identified a novel human pancreatic $\alpha$-amylase binder validated by a fluorescence spectral experiment (Ka $=1.0\times 10^{5}\mathrm{M}$). Furthermore, we proposed a virtual screening pipeline by incorporating multiple complementary methods, such as DFCNN, Autodock vina docking, DeepBindBC, and pocket molecular dynamics simulation. Three potential inhibitors of pancreatic $\alpha$-amylase were identified by the proposed pipeline, and interestingly most of them contain glycan groups. Additionally, an online webserver based on the model is available at http://cbblab.siat.ac.cn/DeepBindBC/index.php for the convenience of the users.
Konda Mani Saravanan, Linbu Liao, Hao Wu 0003, Haishan Zhang, Yi Pan 0001, Xuli Wu, Yanjie Wei
BIBM5
2021 Accurate feature selection improves single-cell RNA-seq cell clustering
abstract
Cell clustering is one of the most important and commonly performed tasks in single-cell RNA sequencing (scRNA-seq) data analysis. An important step in cell clustering is to select a subset of genes (referred to as 'features'), whose expression patterns will then be used for downstream clustering. A good set of features should include the ones that distinguish different cell types, and the quality of such set could have a significant impact on the clustering accuracy. All existing scRNA-seq clustering tools include a feature selection step relying on some simple unsupervised feature selection methods, mostly based on the statistical moments of gene-wise expression distributions. In this work, we carefully evaluate the impact of feature selection on cell clustering accuracy. In addition, we develop a feature selection algorithm named FEAture SelecTion (FEAST), which provides more representative features. We apply the method on 12 public scRNA-seq datasets and demonstrate that using features selected by FEAST with existing clustering tools significantly improve the clustering accuracy.
Kenong Su, Tianwei Yu, Hao Wu 0003
Briefings Bioinform.3
2021 Detecting m6A methylation regions from Methylated RNA Immunoprecipitation Sequencing
abstract
MOTIVATION: The post-transcriptional epigenetic modification on mRNA is an emerging field to study the gene regulatory mechanism and their association with diseases. Recently developed high-throughput sequencing technology named Methylated RNA Immunoprecipitation Sequencing (MeRIP-seq) enables one to profile mRNA epigenetic modification transcriptome wide. A few computational methods are available to identify transcriptome-wide mRNA modification, but they are either limited by over-simplified model ignoring the biological variance across replicates or suffer from low accuracy and efficiency. RESULTS: In this work, we develop a novel statistical method, based on an empirical Bayesian hierarchical model, to identify mRNA epigenetic modification regions from MeRIP-seq data. Our method accounts for various sources of variations in the data through rigorous modeling and applies shrinkage estimation by borrowing information from transcriptome-wide data to stabilize the parameter estimation. Simulation and real data analyses demonstrate that our method is more accurate, robust and efficient than the existing peak calling methods. AVAILABILITY AND IMPLEMENTATION: Our method TRES is implemented as an R package and is freely available on Github at https://github.com/ZhenxingGuo0015/TRES. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhenxing Guo 0001, Andrew M. Shafik, Zhijin Wu, Hao Wu 0003
Bioinform.5
2021 Complete deconvolution of DNA methylation signals from complex tissues: a geometric approach
abstract
MOTIVATION: It is a common practice in epigenetics research to profile DNA methylation on tissue samples, which is usually a mixture of different cell types. To properly account for the mixture, estimating cell compositions has been recognized as an important first step. Many methods were developed for quantifying cell compositions from DNA methylation data, but they mostly have limited applications due to lack of reference or prior information. RESULTS: We develop Tsisal, a novel complete deconvolution method which accurately estimate cell compositions from DNA methylation data without any prior knowledge of cell types or their proportions. Tsisal is a full pipeline to estimate number of cell types, cell compositions and identify cell-type-specific CpG sites. It can also assign cell type labels when (full or part of) reference panel is available. Extensive simulation studies and analyses of seven real datasets demonstrate the favorable performance of our proposed method compared with existing deconvolution methods serving similar purpose. AVAILABILITY AND IMPLEMENTATION: The proposed method Tsisal is implemented as part of the R/Bioconductor package TOAST at https://bioconductor.org/packages/TOAST. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hao Wu 0003, Ziyi Li 0001
Bioinform.2
2020 A comprehensive review of computational prediction of genome-wide features
abstract
There are significant correlations among different types of genetic, genomic and epigenomic features within the genome. These correlations make the in silico feature prediction possible through statistical or machine learning models. With the accumulation of a vast amount of high-throughput data, feature prediction has gained significant interest lately, and a plethora of papers have been published in the past few years. Here we provide a comprehensive review on these published works, categorized by the prediction targets, including protein binding site, enhancer, DNA methylation, chromatin structure and gene expression. We also provide discussions on some important points and possible future directions.
Tianlei Xu, Xiaoqi Zheng, Zhaohui S. Qin, Hao Wu 0003
Briefings Bioinform.6
2020 Robust partial reference-free cell composition estimation from tissue expression
abstract
MOTIVATION: In the analysis of high-throughput omics data from tissue samples, estimating and accounting for cell composition have been recognized as important steps. High cost, intensive labor requirements and technical limitations hinder the cell composition quantification using cell-sorting or single-cell technologies. Computational methods for cell composition estimation are available, but they are either limited by the availability of a reference panel or suffer from low accuracy. RESULTS: We introduce TOols for the Analysis of heterogeneouS Tissues TOAST/-P and TOAST/+P, two partial reference-free algorithms for estimating cell composition of heterogeneous tissues based on their gene expression profiles. TOAST/-P and TOAST/+P incorporate additional biological information, including cell-type-specific markers and prior knowledge of compositions, in the estimation procedure. Extensive simulation studies and real data analyses demonstrate that the proposed methods provide more accurate and robust cell composition estimation than existing methods. AVAILABILITY AND IMPLEMENTATION: The proposed methods TOAST/-P and TOAST/+P are implemented as part of the R/Bioconductor package TOAST at https://bioconductor.org/packages/TOAST. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ziyi Li 0001, Zhenxing Guo 0001, Hao Wu 0003
Bioinform.5
2020 Simulation, power evaluation and sample size recommendation for single-cell RNA-seq
abstract
MOTIVATION: Determining the sample size for adequate power to detect statistical significance is a crucial step at the design stage for high-throughput experiments. Even though a number of methods and tools are available for sample size calculation for microarray and RNA-seq in the context of differential expression (DE), this topic in the field of single-cell RNA sequencing is understudied. Moreover, the unique data characteristics present in scRNA-seq such as sparsity and heterogeneity increase the challenge. RESULTS: We propose POWSC, a simulation-based method, to provide power evaluation and sample size recommendation for single-cell RNA-sequencing DE analysis. POWSC consists of a data simulator that creates realistic expression data, and a power assessor that provides a comprehensive evaluation and visualization of the power and sample size relationship. The data simulator in POWSC outperforms two other state-of-art simulators in capturing key characteristics of real datasets. The power assessor in POWSC provides a variety of power evaluations including stratified and marginal power analyses for DEs characterized by two forms (phase transition or magnitude tuning), under different comparison scenarios. In addition, POWSC offers information for optimizing the tradeoffs between sample size and sequencing depth with the same total reads. AVAILABILITY AND IMPLEMENTATION: POWSC is an open-source R package available online at https://github.com/suke18/POWSC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Kenong Su, Zhijin Wu, Hao Wu 0003
Bioinform.3
2019 Disease prediction by cell-free DNA methylation
abstract
Disease diagnosis using cell-free DNA (cfDNA) has been an active research field recently. Most existing approaches perform diagnosis based on the detection of sequence variants on cfDNA; thus, their applications are limited to diseases associated with high mutation rate such as cancer. Recent developments start to exploit the epigenetic information on cfDNA, which could have substantially wider applications. In this work, we provide thorough reviews and discussions on the statistical method developments and data analysis strategies for using cfDNA epigenetic profiles, in particular DNA methylation, to construct disease diagnostic models. We focus on two important aspects: marker selection and prediction model construction, under different scenarios. We perform simulations and real data analysis to compare different approaches, and provide recommendations for data analysis.
Hao Feng 0005, Hao Wu 0003
Briefings Bioinform.3
2019 Dissecting differential signals in high-throughput data from complex tissues
abstract
MOTIVATION: Samples from clinical practices are often mixtures of different cell types. The high-throughput data obtained from these samples are thus mixed signals. The cell mixture brings complications to data analysis, and will lead to biased results if not properly accounted for. RESULTS: We develop a method to model the high-throughput data from mixed, heterogeneous samples, and to detect differential signals. Our method allows flexible statistical inference for detecting a variety of cell-type specific changes. Extensive simulation studies and analyses of two real datasets demonstrate the favorable performance of our proposed method compared with existing ones serving similar purpose. AVAILABILITY AND IMPLEMENTATION: The proposed method is implemented as an R package and is freely available on GitHub (https://github.com/ziyili20/TOAST). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ziyi Li 0001, Zhijin Wu, Hao Wu 0003
Bioinform.4
2018 Two-phase differential expression analysis for single cell RNA-seq
abstract
Motivation: Single-cell RNA-sequencing (scRNA-seq) has brought the study of the transcriptome to higher resolution and makes it possible for scientists to provide answers with more clarity to the question of 'differential expression'. However, most computational methods still stick with the old mentality of viewing differential expression as a simple 'up or down' phenomenon. We advocate that we should fully embrace the features of single cell data, which allows us to observe binary (from Off to On) as well as continuous (the amount of expression) regulations. Results: We develop a method, termed SC2P, that first identifies the phase of expression a gene is in, by taking into account of both cell- and gene-specific contexts, in a model-based and data-driven fashion. We then identify two forms of transcription regulation: phase transition, and magnitude tuning. We demonstrate that compared with existing methods, SC2P provides substantial improvement in sensitivity without sacrificing the control of false discovery, as well as better robustness. Furthermore, the analysis provides better interpretation of the nature of regulation types in different genes. Availability and implementation: SC2P is implemented as an open source R package publicly available at https://github.com/haowulab/SC2P. Supplementary information: Supplementary data are available at Bioinformatics online.
Zhijin Wu, Michael L. Stitzel, Hao Wu 0003
Bioinform.4
2017 Accounting for tumor purity improves cancer subtype classification from DNA methylation data
abstract
MOTIVATION: Tumor sample classification has long been an important task in cancer research. Classifying tumors into different subtypes greatly benefits therapeutic development and facilitates application of precision medicine on patients. In practice, solid tumor tissue samples obtained from clinical settings are always mixtures of cancer and normal cells. Thus, the data obtained from these samples are mixed signals. The 'tumor purity', or the percentage of cancer cells in cancer tissue sample, will bias the clustering results if not properly accounted for. RESULTS: In this article, we developed a model-based clustering method and an R function which uses DNA methylation microarray data to infer tumor subtypes with the consideration of tumor purity. Simulation studies and the analyses of The Cancer Genome Atlas data demonstrate improved results compared with existing methods. AVAILABILITY AND IMPLEMENTATION: InfiniumClust is part of R package InfiniumPurify , which is freely available from CRAN ( https://cran.r-project.org/web/packages/InfiniumPurify/index.html ). CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hao Feng 0005, Hao Wu 0003, Xiaoqi Zheng
Bioinform.3
2016 Differential methylation analysis for BS-seq data under general experimental design
abstract
MOTIVATION: DNA methylation is an epigenetic modification with important roles in many biological processes and diseases. Bisulfite sequencing (BS-seq) has emerged recently as the technology of choice to profile DNA methylation because of its accuracy, genome coverage and higher resolution. Current statistical methods to identify differential methylation mainly focus on comparing two treatment groups. With an increasing number of experiments performed under a general and multiple-factor design, particularly in reduced representation bisulfite sequencing, there is a need to develop more flexible, powerful and computationally efficient methods. RESULTS: We present a novel statistical model to detect differentially methylated loci from BS-seq data under general experimental design, based on a beta-binomial regression model with 'arcsine' link function. Parameter estimation is based on transformed data with generalized least square approach without relying on iterative algorithm. Simulation and real data analyses demonstrate that our method is accurate, powerful, robust and computationally efficient. AVAILABILITY AND IMPLEMENTATION: It is available as Bioconductor package DSS. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yongseok Park, Hao Wu 0003
Bioinform.2
2016 Measuring the spatial correlations of protein binding sites
abstract
MOTIVATION: Understanding the interactions of different DNA binding proteins is a crucial first step toward deciphering gene regulatory mechanism. With advances of high-throughput sequencing technology such as ChIP-seq, the genome-wide binding sites of many proteins have been profiled under different biological contexts. It is of great interest to quantify the spatial correlations of the binding sites, such as their overlaps, to provide information for the interactions of proteins. Analyses of the overlapping patterns of binding sites have been widely performed, mostly based on ad hoc methods. Due to the heterogeneity and the tremendous size of the genome, such methods often lead to biased even erroneous results. RESULTS: In this work, we discover a Simpson's paradox phenomenon in assessing the genome-wide spatial correlation of protein binding sites. Leveraging information from publicly available data, we propose a testing procedure for evaluating the significance of overlapping from a pair of proteins, which accounts for background artifacts and genome heterogeneity. Real data analyses demonstrate that the proposed method provide more biologically meaningful results. AVAILABILITY AND IMPLEMENTATION: An R package is available at http://www.sta.cuhk.edu.hk/YWei/ChIPCor.html CONTACTS: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hao Wu 0003
Bioinform.2
2015 A novel statistical method for quantitative comparison of multiple ChIP-seq datasets
abstract
MOTIVATION: ChIP-seq is a powerful technology to measure the protein binding or histone modification strength in the whole genome scale. Although there are a number of methods available for single ChIP-seq data analysis (e.g. 'peak detection'), rigorous statistical method for quantitative comparison of multiple ChIP-seq datasets with the considerations of data from control experiment, signal to noise ratios, biological variations and multiple-factor experimental designs is under-developed. RESULTS: In this work, we develop a statistical method to perform quantitative comparison of multiple ChIP-seq datasets and detect genomic regions showing differential protein binding or histone modification. We first detect peaks from all datasets and then union them to form a single set of candidate regions. The read counts from IP experiment at the candidate regions are assumed to follow Poisson distribution. The underlying Poisson rates are modeled as an experiment-specific function of artifacts and biological signals. We then obtain the estimated biological signals and compare them through the hypothesis testing procedure in a linear model framework. Simulations and real data analyses demonstrate that the proposed method provides more accurate and robust results compared with existing ones. AVAILABILITY AND IMPLEMENTATION: An R software package ChIPComp is freely available at http://web1.sph.emory.edu/users/hwu30/software/ChIPComp.html.
Li Chen 0029, Chi Wang 0003, Zhaohui S. Qin, Hao Wu 0003
Bioinform.4
2015 PROPER: comprehensive power evaluation for differential expression using RNA-seq
abstract
MOTIVATION: RNA-seq has become a routine technique in differential expression (DE) identification. Scientists face a number of experimental design decisions, including the sample size. The power for detecting differential expression is affected by several factors, including the fraction of DE genes, distribution of the magnitude of DE, distribution of gene expression level, sequencing coverage and the choice of type I error control. The complexity and flexibility of RNA-seq experiments, the high-throughput nature of transcriptome-wide expression measurements and the unique characteristics of RNA-seq data make the power assessment particularly challenging. RESULTS: We propose prospective power assessment instead of a direct sample size calculation by making assumptions on all of these factors. Our power assessment tool includes two components: (i) a semi-parametric simulation that generates data based on actual RNA-seq experiments with flexible choices on baseline expressions, biological variations and patterns of DE; and (ii) a power assessment component that provides a comprehensive view of power. We introduce the concepts of stratified power and false discovery cost, and demonstrate the usefulness of our method in experimental design (such as sample size and sequencing depth), as well as analysis plan (gene filtering). AVAILABILITY: The proposed method is implemented in a freely available R software package PROPER. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hao Wu 0003, Chi Wang 0003, Zhijin Wu
Bioinform.1
2015 Predicting tumor purity from methylation microarray data
abstract
MOTIVATION: In cancer genomics research, one important problem is that the solid tissue sample obtained from clinical settings is always a mixture of cancer and normal cells. The sample mixture brings complication in data analysis and results in biased findings if not correctly accounted for. Estimating tumor purity is of great interest, and a number of methods have been developed using gene expression, copy number variation or point mutation data. RESULTS: We discover that in cancer samples, the distributions of data from Illumina Infinium 450 k methylation microarray are highly correlated with tumor purities. We develop a simple but effective method to estimate purities from the microarray data. Analyses of the Cancer Genome Atlas lung cancer data demonstrate favorable performance of the proposed method. AVAILABILITY AND IMPLEMENTATION: The method is implemented in InfiniumPurify, which is freely available at https://bitbucket.org/zhengxiaoqi/infiniumpurify. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Naiqian Zhang, Hua-Jun Wu, Hao Wu 0003, Xiaoqi Zheng
Bioinform.5
2010 JAMIE: joint analysis of multiple ChIP-chip experiments
abstract
MOTIVATION: Chromatin immunoprecipitation followed by genome tiling array hybridization (ChIP-chip) is a powerful approach to identify transcription factor binding sites (TFBSs) in target genomes. When multiple related ChIP-chip datasets are available, analyzing them jointly allows one to borrow information across datasets to improve peak detection. This is particularly useful for analyzing noisy datasets. RESULTS: We propose a hierarchical mixture model and develop an R package JAMIE to perform the joint analysis. The genome is assumed to consist of background and potential binding regions (PBRs). PBRs have context-dependent probabilities to become bona fide binding sites in individual datasets. This model captures the correlation among datasets, which provides basis for sharing information across experiments. Real data tests illustrate the advantage of JAMIE over a strategy that analyzes individual datasets separately. AVAILABILITY: JAMIE is freely available from http://www.biostat.jhsph.edu/~hji/jamie
Hao Wu 0003, Hongkai Ji
Bioinform.1
2007 R/qtlbim: QTL with Bayesian Interval Mapping in experimental crosses
abstract
UNLABELLED: R/qtlbim is an extensible, interactive environment for the Bayesian Interval Mapping of QTL, built on top of R/qtl (Broman et al., 2003), providing Bayesian analysis of multiple interacting quantitative trait loci (QTL) models for continuous, binary and ordinal traits in experimental crosses. It includes several efficient Markov chain Monte Carlo (MCMC) algorithms for evaluating the posterior of genetic architectures, i.e. the number and locations of QTL, their main and epistatic effects and gene-environment interactions. R/qtlbim provides extensive informative graphical and numerical summaries, and model selection and convergence diagnostics of the MCMC output, illustrated through the vignette, example and demo capabilities of R (R Development Core Team 2006). AVAILABILITY: The package is freely available from cran.r-project.org.
Brian S. Yandell, Tapan Mehta, Samprit Banerjee, Daniel Shriner, Ramprasad Venkataraman, Jee Young Moon, W. Whipple Neely, Hao Wu 0003, Randy von Smith, Nengjun Yi
Bioinform.8
2003 R/qtl: QTL Mapping in Experimental Crosses
abstract
Abstract Summary: R/qtl is an extensible, interactive environment for mapping quantitative trait loci (QTLs) in experimental populations derived from inbred lines. It is implemented as an add-on package for the freely-available statistical software, R, and includes functions for estimating genetic maps, identifying genotyping errors, and performing single-QTL and two-dimensional, two-QTL genome scans by multiple methods, with the possible inclusion of covariates. Availability: The package is freely available at http://www.biostat.jhsph.edu/~kbroman/qtl. Contact: [email protected] * To whom correspondence should be addressed. † Present address: Department of Epidemiology and Biostatistics, University of California, San Francisco, CA 94143, USA.
Karl W. Broman, Hao Wu 0003, Saunak Sen, Gary A. Churchill
Bioinform.2