Zhijin Wu

dblp:13/7027 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
4since 2021 · last 2024
0000-0002-9596-9134ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 4 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2024 Target-oriented Reference Construction for supervised cell type identification in scRNA-seq
abstract
Cell type identification is a crucial step in single-cell RNA-seq (scRNA-seq) data analysis. The supervised cell type identification method is a preferred solution due to its accuracy and efficiency. The performance of such methods highly depends on the quality of the reference data. Although there are many supervised cell type identification tools, no method currently exists for constructing reference data. Here, we develop Target-Oriented Reference Construction (TORC), a widely applicable strategy for constructing references based on target datasets in scRNA-seq supervised cell type identification. TORC focuses on alleviating differences in cell type composition between the reference and target sets. Extensive benchmarks on simulated and real data analyses demonstrate consistent improvements in cell type identification with TORC.
Wenjing Ma, Zhijin Wu, Hao Wu 0003
BIBM3
2024 m6A peak calling accounting for varying sequencing bias across regions and samples
abstract
N6-Methyladenosine (m6A) is the most abundant type of mRNA methylation and is most widely measured by methylated RNA immunoprecipitation sequencing (MeRIP-seq). In MeRIP-seq, an immunoprecipitation (IP) sample and a pairing control (input) sample are sequenced for each biological sample. Methylated regions are identified as peaks showing increased counts in the IP sample versus the input. We report that technical bias in sequencing can vary substantially in the IP and input samples depending on the local sequence context. Current sequencing depth-based normalization does not appropriately account for the varying technical bias along the transcriptome and leads to inaccurate identification of m6A regions. We describe a method to estimate a local size factor that reflects the RNA sequence context and show that peak calling using these region-specific size factors identifies more accurate peak regions.
Lanyu Zhang, Zhenxing Guo 0001, Zhaohui Qin, Zhijin Wu
BIBM4
2023 Accurate Detection of MicroRNAs from NanoString nCounter with a Latent Mixture Model
abstract
MicroRNAs (miRNA) are promising biomarker candidates for diagnosing neurodegenerative diseases due to their presence in easy-to-obtain biofluids. The NanoString nCounter is a popular platform measuring miRNA for it avoids amplification bias. Existing methods for nCounter data processing and analysis rely heavily on the handful of control probes and housekeeping genes for background estimation and/or normalization. Motivated by the observations from hundreds of samples compiled from multiple studies, we propose a multi-study joint processing method, multi-study miRNA detection (MMD). MMD is based on a latent mixture model that accounts for both probe-specific and sample-specific effects. The probe effects are estimated jointly from samples across studies. Sample-specific background and normalization factors are estimated from all probes instead of relying on a few controls. We demonstrate that MMD outperforms the built-in method from Nanostring in signal detection and has greater power in identifying differentially present miRNAs which are largely overlooked by alternative methods, in both simulation and real data comparison.
Zhijin Wu
BIBM2
2021 Detecting m6A methylation regions from Methylated RNA Immunoprecipitation Sequencing
abstract
MOTIVATION: The post-transcriptional epigenetic modification on mRNA is an emerging field to study the gene regulatory mechanism and their association with diseases. Recently developed high-throughput sequencing technology named Methylated RNA Immunoprecipitation Sequencing (MeRIP-seq) enables one to profile mRNA epigenetic modification transcriptome wide. A few computational methods are available to identify transcriptome-wide mRNA modification, but they are either limited by over-simplified model ignoring the biological variance across replicates or suffer from low accuracy and efficiency. RESULTS: In this work, we develop a novel statistical method, based on an empirical Bayesian hierarchical model, to identify mRNA epigenetic modification regions from MeRIP-seq data. Our method accounts for various sources of variations in the data through rigorous modeling and applies shrinkage estimation by borrowing information from transcriptome-wide data to stabilize the parameter estimation. Simulation and real data analyses demonstrate that our method is more accurate, robust and efficient than the existing peak calling methods. AVAILABILITY AND IMPLEMENTATION: Our method TRES is implemented as an R package and is freely available on Github at https://github.com/ZhenxingGuo0015/TRES. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhenxing Guo 0001, Andrew M. Shafik, Zhijin Wu, Hao Wu 0003
Bioinform.4
2020 Simulation, power evaluation and sample size recommendation for single-cell RNA-seq
abstract
MOTIVATION: Determining the sample size for adequate power to detect statistical significance is a crucial step at the design stage for high-throughput experiments. Even though a number of methods and tools are available for sample size calculation for microarray and RNA-seq in the context of differential expression (DE), this topic in the field of single-cell RNA sequencing is understudied. Moreover, the unique data characteristics present in scRNA-seq such as sparsity and heterogeneity increase the challenge. RESULTS: We propose POWSC, a simulation-based method, to provide power evaluation and sample size recommendation for single-cell RNA-sequencing DE analysis. POWSC consists of a data simulator that creates realistic expression data, and a power assessor that provides a comprehensive evaluation and visualization of the power and sample size relationship. The data simulator in POWSC outperforms two other state-of-art simulators in capturing key characteristics of real datasets. The power assessor in POWSC provides a variety of power evaluations including stratified and marginal power analyses for DEs characterized by two forms (phase transition or magnitude tuning), under different comparison scenarios. In addition, POWSC offers information for optimizing the tradeoffs between sample size and sequencing depth with the same total reads. AVAILABILITY AND IMPLEMENTATION: POWSC is an open-source R package available online at https://github.com/suke18/POWSC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Kenong Su, Zhijin Wu, Hao Wu 0003
Bioinform.2
2019 Dissecting differential signals in high-throughput data from complex tissues
abstract
MOTIVATION: Samples from clinical practices are often mixtures of different cell types. The high-throughput data obtained from these samples are thus mixed signals. The cell mixture brings complications to data analysis, and will lead to biased results if not properly accounted for. RESULTS: We develop a method to model the high-throughput data from mixed, heterogeneous samples, and to detect differential signals. Our method allows flexible statistical inference for detecting a variety of cell-type specific changes. Extensive simulation studies and analyses of two real datasets demonstrate the favorable performance of our proposed method compared with existing ones serving similar purpose. AVAILABILITY AND IMPLEMENTATION: The proposed method is implemented as an R package and is freely available on GitHub (https://github.com/ziyili20/TOAST). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ziyi Li 0001, Zhijin Wu, Hao Wu 0003
Bioinform.2
2018 Two-phase differential expression analysis for single cell RNA-seq
abstract
Motivation: Single-cell RNA-sequencing (scRNA-seq) has brought the study of the transcriptome to higher resolution and makes it possible for scientists to provide answers with more clarity to the question of 'differential expression'. However, most computational methods still stick with the old mentality of viewing differential expression as a simple 'up or down' phenomenon. We advocate that we should fully embrace the features of single cell data, which allows us to observe binary (from Off to On) as well as continuous (the amount of expression) regulations. Results: We develop a method, termed SC2P, that first identifies the phase of expression a gene is in, by taking into account of both cell- and gene-specific contexts, in a model-based and data-driven fashion. We then identify two forms of transcription regulation: phase transition, and magnitude tuning. We demonstrate that compared with existing methods, SC2P provides substantial improvement in sensitivity without sacrificing the control of false discovery, as well as better robustness. Furthermore, the analysis provides better interpretation of the nature of regulation types in different genes. Availability and implementation: SC2P is implemented as an open source R package publicly available at https://github.com/haowulab/SC2P. Supplementary information: Supplementary data are available at Bioinformatics online.
Zhijin Wu, Michael L. Stitzel, Hao Wu 0003
Bioinform.1
2018 The International Conference on Intelligent Biology and Medicine (ICIBM) 2018: bioinformatics towards translational applications
Xiaoming Liu 0021, Lei Xie 0006, Zhijin Wu, Kai Wang 0063, Zhongming Zhao, Jianhua Ruan, Degui Zhi
BMC Bioinform.3
2015 PROPER: comprehensive power evaluation for differential expression using RNA-seq
abstract
MOTIVATION: RNA-seq has become a routine technique in differential expression (DE) identification. Scientists face a number of experimental design decisions, including the sample size. The power for detecting differential expression is affected by several factors, including the fraction of DE genes, distribution of the magnitude of DE, distribution of gene expression level, sequencing coverage and the choice of type I error control. The complexity and flexibility of RNA-seq experiments, the high-throughput nature of transcriptome-wide expression measurements and the unique characteristics of RNA-seq data make the power assessment particularly challenging. RESULTS: We propose prospective power assessment instead of a direct sample size calculation by making assumptions on all of these factors. Our power assessment tool includes two components: (i) a semi-parametric simulation that generates data based on actual RNA-seq experiments with flexible choices on baseline expressions, biological variations and patterns of DE; and (ii) a power assessment component that provides a comprehensive view of power. We introduce the concepts of stratified power and false discovery cost, and demonstrate the usefulness of our method in experimental design (such as sample size and sequencing depth), as well as analysis plan (gene filtering). AVAILABILITY: The proposed method is implemented in a freely available R software package PROPER. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hao Wu 0003, Chi Wang 0003, Zhijin Wu
Bioinform.3
2010 Empirical Bayes Analysis of Sequencing-based Transcriptional Profiling without Replicates
abstract
BACKGROUND: Recent technological advancements have made high throughput sequencing an increasingly popular approach for transcriptome analysis. Advantages of sequencing-based transcriptional profiling over microarrays have been reported, including lower technical variability. However, advances in technology do not remove biological variation between replicates and this variation is often neglected in many analyses. RESULTS: We propose an empirical Bayes method, titled Analysis of Sequence Counts (ASC), to detect differential expression based on sequencing technology. ASC borrows information across sequences to establish prior distribution of sample variation, so that biological variation can be accounted for even when replicates are not available. Compared to current approaches that simply tests for equality of proportions in two samples, ASC is less biased towards highly expressed sequences and can identify more genes with a greater log fold change at lower overall abundance. CONCLUSIONS: ASC unifies the biological and statistical significance of differential expression by estimating the posterior mean of log fold change and estimating false discovery rates based on the posterior mean. The implementation in R is available at http://www.stat.brown.edu/Zwu/research.aspx.
Zhijin Wu, Bethany D. Jenkins, Tatiana A. Rynearson, Sonya T. Dyhrman, Mak A. Saito, Melissa Mercier, LeAnn P. Whitney
BMC Bioinform.1
2009 A novel highly accurate log skew normal approximation method to lognormal sum distributions
abstract
Sums of lognormal random variables occur in many important problems in wireless communications. However, the lognormal sum distribution is known to have no close-form and is difficult to compute numerically. Several approximation methods have already been proposed to approximate the lognormal sum distribution. However, these approximation methods all have their drawbacks: some widely used approximation methods are not very accurate at the lower region, some other approximation methods require the CDF curve from Monte Carlo simulation first. In this paper, we propose a novel approximation method, namely the Log Skew Normal (LSN) approximation, to model and approximate the sum of M lognormal distributed random variables. The proposed LSN approximation method has very high accuracy in most of the region, especially in the lower region. Furthermore, this approximation method does not require the CDF curve from Monte Carlo simulation first. The closed-form probability density function (PDF) of the resulting LSN random variable is presented and its parameters are derived from those of the M individual lognormal random variables by using an moment matching technique. Simulation results on the cumulative distribution function (CDF) of sum of M lognormal random variables in different conditions are used as reference curves to compare various approximation techniques. LSN approximation is found to provide better accuracy over a wide CDF range over other approximation methods.
Zhijin Wu, Xue Li 0002, Robert Husnay, Vasu Chakravarthy, Bin Wang 0002, Zhiqiang Wu 0001
WCNC1
2006 Comparison of Affymetrix GeneChip expression measures
abstract
MOTIVATION: In the Affymetrix GeneChip system, preprocessing occurs before one obtains expression level measurements. Because the number of competing preprocessing methods was large and growing we developed a benchmark to help users identify the best method for their application. A webtool was made available for developers to benchmark their procedures. At the time of writing over 50 methods had been submitted. RESULTS: We benchmarked 31 probe set algorithms using a U95A dataset of spike in controls. Using this dataset, we found that background correction, one of the main steps in preprocessing, has the largest effect on performance. In particular, background correction appears to improve accuracy but, in general, worsen precision. The benchmark results put this balance in perspective. Furthermore, we have improved some of the original benchmark metrics to provide more detailed information regarding precision and accuracy. A handful of methods stand out as providing the best balance using spike-in data with the older U95A array, although different experiments on more current arrays may benchmark differently. AVAILABILITY: The affycomp package, now version 1.5.2, continues to be available as part of the Bioconductor project (http://www.bioconductor.org). The webtool continues to be available at http://affycomp.biostat.jhsph.edu CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Rafael A. Irizarry, Zhijin Wu, Harris A. Jaffee
Bioinform.2
2004 Stochastic models inspired by hybridization theory for short oligonucleotide arrays
abstract
High density oligonucleotide expression arrays are a widely used tool for the measurement of gene expression on a large scale. Affymetrix GeneChip arrays appear to dominate this market. These arrays use short oligonucleotides to probe for genes in an RNA sample. Due to optical noise, non-specific hybridization, probe-specific effects, and measurement error, ad-hoc measures of expression, that summarize probe intensities, can lead to imprecise and inaccurate results. Various researchers have demonstrated that expression measures based on simple statistical models can provide great improvements over the ad-hoc procedure offered by Affymetrix. Recently, physical models based on molecular hybridization theory, have been proposed as useful tools for prediction of, for example, non-specific hybridization. These physical models show great potential in terms of improving existing expression measures. In this paper we suggest that the system producing the measured intensities is too complex to be fully described with these relatively simple physical models and we propose empirically motivated stochastic models that compliment the above mentioned molecular hybridization theory to provide a comprehensive description of the data. We discuss how the proposed model can be used to obtain improved measures of expression useful for the data analysts.
Zhijin Wu, Rafael A. Irizarry
RECOMB1
2004 A benchmark for Affymetrix GeneChip expression measures
abstract
Abstract Motivation: The defining feature of oligonucleotide expression arrays is the use of several probes to assay each targeted transcript. This is a bonanza for the statistical geneticist, who can create probeset summaries with specific characteristics. There are now several methods available for summarizing probe level data from the popular Affymetrix GeneChips, but it is difficult to identify the best method for a given inquiry. Results: We have developed a graphical tool to evaluate summaries of Affymetrix probe level data. Plots and summary statistics offer a picture of how an expression measure performs in several important areas. This picture facilitates the comparison of competing expression measures and the selection of methods suitable for a specific investigation. The key is a benchmark data set consisting of a dilution study and a spike-in study. Because the truth is known for these data, we can identify statistical features of the data for which the expected outcome is known in advance. Those features highlighted in our suite of graphs are justified by questions of biological interest and motivated by the presence of appropriate data. Availability: In conjunction with the release of a graphics toolbox as part of the Bioconductor project (http://www.bioconductor.org), a webtool is available at http://affycomp.biostat.jhsph.edu. Supplemental material is available at http://www.biostat.jhsph.edu/~ririzarr/papers/suppaffycomp.pdf
Leslie Cope, Rafael A. Irizarry, Harris A. Jaffee, Zhijin Wu, Terence P. Speed
Bioinform.4