Hao Feng 0005

dblp:46/4184-5 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
5since 2021 · last 2024
0000-0003-2243-9949ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2024 cypress: an R/Bioconductor package for cell-type-specific differential expression analysis power assessment
abstract
SUMMARY: Recent methodology advances in computational signal deconvolution have enabled bulk transcriptome data analysis at a finer cell-type level. Through deconvolution, identifying cell-type-specific differentially expressed (csDE) genes is drawing increasing attention in clinical applications. However, researchers still face a number of difficulties in adopting csDE genes detection methods in practice, especially in their experimental design. Here we present cypress, the first experimental design and statistical power analysis tool in csDE genes identification. This tool can reliably model purified cell-type-specific (CTS) profiles, cell-type compositions, biological and technical variations, offering a high-fidelity simulator for bulk RNA-seq convolution and deconvolution. cypress conducts simulation and evaluates the impact of multiple influencing factors, by various statistical metrics, to help researchers optimize experimental design and conduct power analysis. AVAILABILITY AND IMPLEMENTATION: cypress is an open-source R/Bioconductor package at https://bioconductor.org/packages/cypress/.
Shilin Yu, Guanqun Meng, Wen Tang 0003, Wenjing Ma, Xiongwei Zhu, Hao Feng 0005
Bioinform.8
2024 magpie: A power evaluation method for differential RNA methylation analysis in N6-methyladenosine sequencing
abstract
Recently, novel biotechnologies to quantify RNA modifications became an increasingly popular choice for researchers who study epitranscriptome. When studying RNA methylations such as N6-methyladenosine (m6A), researchers need to make several decisions in its experimental design, especially the sample size and a proper statistical power. Due to the complexity and high-throughput nature of m6A sequencing measurements, methods for power calculation and study design are still currently unavailable. In this work, we propose a statistical power assessment tool, magpie, for power calculation and experimental design for epitranscriptome studies using m6A sequencing data. Our simulation-based power assessment tool will borrow information from real pilot data, and inspect various influential factors including sample size, sequencing depth, effect size, and basal expression ranges. We integrate two modules in magpie: (i) a flexible and realistic simulator module to synthesize m6A sequencing data based on real data; and (ii) a power assessment module to examine a set of comprehensive evaluation metrics.
Zhenxing Guo 0001, Daoyu Duan, Wen Tang 0003, Julia Zhu, William S. Bush, Fulai Jin, Hao Feng 0005
PLoS Comput. Biol.9
2023 Evaluation of epitranscriptome-wide N6-methyladenosine differential analysis methods
abstract
RNA methylation has emerged recently as an active research domain to study post-transcriptional alteration in gene expression regulation. Various types of RNA methylation, including N6-methyladenosine (m6A), are involved in human disease development. As a newly developed sequencing biotechnology to quantify the m6A level on a transcriptome-wide scale, MeRIP-seq expands RNA epigenetics study in both basic and clinical applications, with an upward trend. One of the fundamental questions in RNA methylation data analysis is to identify the Differentially Methylated Regions (DMRs), by contrasting cases and controls. Multiple statistical approaches have been recently developed for DMR detection, but there is a lack of a comprehensive evaluation for these analytical methods. Here, we thoroughly assess all eight existing methods for DMR calling, using both synthetic and real data. Our simulation adopts a Gamma-Poisson model and logit linear framework, and accommodates various sample sizes and DMR proportions for benchmarking. For all methods, low sensitivities are observed among regions with low input levels, but they can be drastically boosted by an increase in sample size. TRESS and exomePeak2 perform the best using metrics of detection precision, FDR, type I error control and runtime, though hampered by low sensitivity. DRME and exomePeak obtain high sensitivities, at the expense of inflated FDR and type I error. Analyses on three real datasets suggest differential preference on identified DMR length and uniquely discovered regions, between these methods.
Daoyu Duan, Wen Tang 0003, Runshu Wang, Zhenxing Guo 0001, Hao Feng 0005
Briefings Bioinform.5
2023 A comprehensive assessment of cell type-specific differential expression methods in bulk data
abstract
Accounting for cell type compositions has been very successful at analyzing high-throughput data from heterogeneous tissues. Differential gene expression analysis at cell type level is becoming increasingly popular, yielding biomarker discovery in a finer granularity within a particular cell type. Although several computational methods have been developed to identify cell type-specific differentially expressed genes (csDEG) from RNA-seq data, a systematic evaluation is yet to be performed. Here, we thoroughly benchmark six recently published methods: CellDMC, CARseq, TOAST, LRCDE, CeDAR and TCA, together with two classical methods, csSAM and DESeq2, for a comprehensive comparison. We aim to systematically evaluate the performance of popular csDEG detection methods and provide guidance to researchers. In simulation studies, we benchmark available methods under various scenarios of baseline expression levels, sample sizes, cell type compositions, expression level alterations, technical noises and biological dispersions. Real data analyses of three large datasets on inflammatory bowel disease, lung cancer and autism provide evaluation in both the gene level and the pathway level. We find that csDEG calling is strongly affected by effect size, baseline expression level and cell type compositions. Results imply that csDEG discovery is a challenging task itself, with room to improvements on handling low signal-to-noise ratio and low expression genes.
Guanqun Meng, Wen Tang 0003, Emina Huang, Ziyi Li 0001, Hao Feng 0005
Briefings Bioinform.5
2022 NeuCA web server: a neural network-based cell annotation tool with web-app and GUI
abstract
SUMMARY: Correctly annotating individual cell's type is an important initial step in single-cell RNA sequencing (scRNA-seq) data analysis. Here, we present NeuCA web server, a neural network-based scRNA-seq cell annotation tool with web-app portal and graphical user interface, for automatically assigning cell labels. NeuCA algorithm is accurate and exhaustive, maximizing the usage of measured cells for downstream analysis. NeuCA web server provides over 20 ready-to-use pre-trained classifiers for commonly used tissue types. As the first web-app tool with neural-network infrastructure implemented, NeuCA web will facilitate the research community in analyzing and annotating scRNA-seq data. AVAILABILITY AND IMPLEMENTATION: NeuCA web server is implemented with R Shiny application online at https://statbioinfo.shinyapps.io/NeuCA/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Daoyu Duan, Sijia He, Emina Huang, Ziyi Li 0001, Hao Feng 0005
Bioinform.5
2019 Disease prediction by cell-free DNA methylation
abstract
Disease diagnosis using cell-free DNA (cfDNA) has been an active research field recently. Most existing approaches perform diagnosis based on the detection of sequence variants on cfDNA; thus, their applications are limited to diseases associated with high mutation rate such as cancer. Recent developments start to exploit the epigenetic information on cfDNA, which could have substantially wider applications. In this work, we provide thorough reviews and discussions on the statistical method developments and data analysis strategies for using cfDNA epigenetic profiles, in particular DNA methylation, to construct disease diagnostic models. We focus on two important aspects: marker selection and prediction model construction, under different scenarios. We perform simulations and real data analysis to compare different approaches, and provide recommendations for data analysis.
Hao Feng 0005, Hao Wu 0003
Briefings Bioinform.1
2017 Accounting for tumor purity improves cancer subtype classification from DNA methylation data
abstract
MOTIVATION: Tumor sample classification has long been an important task in cancer research. Classifying tumors into different subtypes greatly benefits therapeutic development and facilitates application of precision medicine on patients. In practice, solid tumor tissue samples obtained from clinical settings are always mixtures of cancer and normal cells. Thus, the data obtained from these samples are mixed signals. The 'tumor purity', or the percentage of cancer cells in cancer tissue sample, will bias the clustering results if not properly accounted for. RESULTS: In this article, we developed a model-based clustering method and an R function which uses DNA methylation microarray data to infer tumor subtypes with the consideration of tumor purity. Simulation studies and the analyses of The Cancer Genome Atlas data demonstrate improved results compared with existing methods. AVAILABILITY AND IMPLEMENTATION: InfiniumClust is part of R package InfiniumPurify , which is freely available from CRAN ( https://cran.r-project.org/web/packages/InfiniumPurify/index.html ). CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hao Feng 0005, Hao Wu 0003, Xiaoqi Zheng
Bioinform.2