Zhenxing Guo 0001

dblp:174/6582-1 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
5since 2021 · last 2024
0000-0003-1681-1337ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2024 m6A peak calling accounting for varying sequencing bias across regions and samples
abstract
N6-Methyladenosine (m6A) is the most abundant type of mRNA methylation and is most widely measured by methylated RNA immunoprecipitation sequencing (MeRIP-seq). In MeRIP-seq, an immunoprecipitation (IP) sample and a pairing control (input) sample are sequenced for each biological sample. Methylated regions are identified as peaks showing increased counts in the IP sample versus the input. We report that technical bias in sequencing can vary substantially in the IP and input samples depending on the local sequence context. Current sequencing depth-based normalization does not appropriately account for the varying technical bias along the transcriptome and leads to inaccurate identification of m6A regions. We describe a method to estimate a local size factor that reflects the RNA sequence context and show that peak calling using these region-specific size factors identifies more accurate peak regions.
Lanyu Zhang, Zhenxing Guo 0001, Zhaohui Qin, Zhijin Wu
BIBM2
2024 magpie: A power evaluation method for differential RNA methylation analysis in N6-methyladenosine sequencing
abstract
Recently, novel biotechnologies to quantify RNA modifications became an increasingly popular choice for researchers who study epitranscriptome. When studying RNA methylations such as N6-methyladenosine (m6A), researchers need to make several decisions in its experimental design, especially the sample size and a proper statistical power. Due to the complexity and high-throughput nature of m6A sequencing measurements, methods for power calculation and study design are still currently unavailable. In this work, we propose a statistical power assessment tool, magpie, for power calculation and experimental design for epitranscriptome studies using m6A sequencing data. Our simulation-based power assessment tool will borrow information from real pilot data, and inspect various influential factors including sample size, sequencing depth, effect size, and basal expression ranges. We integrate two modules in magpie: (i) a flexible and realistic simulator module to synthesize m6A sequencing data based on real data; and (ii) a power assessment module to examine a set of comprehensive evaluation metrics.
Zhenxing Guo 0001, Daoyu Duan, Wen Tang 0003, Julia Zhu, William S. Bush, Fulai Jin, Hao Feng 0005
PLoS Comput. Biol.1
2023 Evaluation of epitranscriptome-wide N6-methyladenosine differential analysis methods
abstract
RNA methylation has emerged recently as an active research domain to study post-transcriptional alteration in gene expression regulation. Various types of RNA methylation, including N6-methyladenosine (m6A), are involved in human disease development. As a newly developed sequencing biotechnology to quantify the m6A level on a transcriptome-wide scale, MeRIP-seq expands RNA epigenetics study in both basic and clinical applications, with an upward trend. One of the fundamental questions in RNA methylation data analysis is to identify the Differentially Methylated Regions (DMRs), by contrasting cases and controls. Multiple statistical approaches have been recently developed for DMR detection, but there is a lack of a comprehensive evaluation for these analytical methods. Here, we thoroughly assess all eight existing methods for DMR calling, using both synthetic and real data. Our simulation adopts a Gamma-Poisson model and logit linear framework, and accommodates various sample sizes and DMR proportions for benchmarking. For all methods, low sensitivities are observed among regions with low input levels, but they can be drastically boosted by an increase in sample size. TRESS and exomePeak2 perform the best using metrics of detection precision, FDR, type I error control and runtime, though hampered by low sensitivity. DRME and exomePeak obtain high sensitivities, at the expense of inflated FDR and type I error. Analyses on three real datasets suggest differential preference on identified DMR length and uniquely discovered regions, between these methods.
Daoyu Duan, Wen Tang 0003, Runshu Wang, Zhenxing Guo 0001, Hao Feng 0005
Briefings Bioinform.4
2022 Differential RNA methylation analysis for MeRIP-seq data under general experimental design
abstract
MOTIVATION: RNA epigenetics is an emerging field to study the post-transcriptional gene regulation. The dynamics of RNA epigenetic modification have been reported to associate with many human diseases. Recently developed high-throughput technology named Methylated RNA Immunoprecipitation Sequencing (MeRIP-seq) enables the transcriptome-wide profiling of N6-methyladenosine (m6A) modification and comparison of RNA epigenetic modifications. There are a few computational methods for the comparison of mRNA modifications under different conditions but they all suffer from serious limitations. RESULTS: In this work, we develop a novel statistical method to detect differentially methylated mRNA regions from MeRIP-seq data. We model the sequence count data by a hierarchical negative binomial model that accounts for various sources of variations and derive parameter estimation and statistical testing procedures for flexible statistical inferences under general experimental designs. Extensive benchmark evaluations in simulation and real data analyses demonstrate that our method is more accurate, robust and flexible compared to existing methods. AVAILABILITY AND IMPLEMENTATION: Our method TRESS is implemented as an R/Bioconductor package and is available at https://bioconductor.org/packages/devel/TRESS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhenxing Guo 0001, Andrew M. Shafik, Hao Wu 0003
Bioinform.1
2021 Detecting m6A methylation regions from Methylated RNA Immunoprecipitation Sequencing
abstract
MOTIVATION: The post-transcriptional epigenetic modification on mRNA is an emerging field to study the gene regulatory mechanism and their association with diseases. Recently developed high-throughput sequencing technology named Methylated RNA Immunoprecipitation Sequencing (MeRIP-seq) enables one to profile mRNA epigenetic modification transcriptome wide. A few computational methods are available to identify transcriptome-wide mRNA modification, but they are either limited by over-simplified model ignoring the biological variance across replicates or suffer from low accuracy and efficiency. RESULTS: In this work, we develop a novel statistical method, based on an empirical Bayesian hierarchical model, to identify mRNA epigenetic modification regions from MeRIP-seq data. Our method accounts for various sources of variations in the data through rigorous modeling and applies shrinkage estimation by borrowing information from transcriptome-wide data to stabilize the parameter estimation. Simulation and real data analyses demonstrate that our method is more accurate, robust and efficient than the existing peak calling methods. AVAILABILITY AND IMPLEMENTATION: Our method TRES is implemented as an R package and is freely available on Github at https://github.com/ZhenxingGuo0015/TRES. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhenxing Guo 0001, Andrew M. Shafik, Zhijin Wu, Hao Wu 0003
Bioinform.1
2020 Robust partial reference-free cell composition estimation from tissue expression
abstract
MOTIVATION: In the analysis of high-throughput omics data from tissue samples, estimating and accounting for cell composition have been recognized as important steps. High cost, intensive labor requirements and technical limitations hinder the cell composition quantification using cell-sorting or single-cell technologies. Computational methods for cell composition estimation are available, but they are either limited by the availability of a reference panel or suffer from low accuracy. RESULTS: We introduce TOols for the Analysis of heterogeneouS Tissues TOAST/-P and TOAST/+P, two partial reference-free algorithms for estimating cell composition of heterogeneous tissues based on their gene expression profiles. TOAST/-P and TOAST/+P incorporate additional biological information, including cell-type-specific markers and prior knowledge of compositions, in the estimation procedure. Extensive simulation studies and real data analyses demonstrate that the proposed methods provide more accurate and robust cell composition estimation than existing methods. AVAILABILITY AND IMPLEMENTATION: The proposed methods TOAST/-P and TOAST/+P are implemented as part of the R/Bioconductor package TOAST at https://bioconductor.org/packages/TOAST. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ziyi Li 0001, Zhenxing Guo 0001, Hao Wu 0003
Bioinform.2