EDBT 2026 Demo / reviewers in the wild / expert
Tao Wang 0067
dblp:12/5838-67
· DBLP profile ↗
11ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0002-1218-4017ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CRAmed: a conditional randomization test for high-dimensional mediation analysis in sparse microbiome dataabstractMOTIVATION: Numerous microbiome studies have revealed significant associations between the microbiome and human health and disease. These findings have motivated researchers to explore the causal role of the microbiome in human complex traits and diseases. However, the complexities of microbiome data pose challenges for statistical analysis and interpretation of causal effects. RESULTS: We introduced a novel statistical framework, CRAmed, for inferring the mediating role of the microbiome between treatment and outcome. CRAmed improved the interpretability of the mediation analysis by decomposing the natural indirect effect into two parts, corresponding to the presence-absence and abundance of a microbe, respectively. Comprehensive simulations demonstrated the superior performance of CRAmed in Recall, precision, and F1 score, with a notable level of robustness, compared to existing mediation analysis methods. Furthermore, two real data applications illustrated the effectiveness and interpretability of CRAmed. Our research revealed that CRAmed holds promise for uncovering the mediating role of the microbiome and understanding of the factors influencing host health. AVAILABILITY AND IMPLEMENTATION: The R package CRAmed implementing the proposed methods is available online at https://github.com/liudoubletian/CRAmed. Xiangnan Xu, Tao Wang 0067, Peirong Xu |
Bioinform. | 3 |
| 2025 | InterVelo: a mutually enhancing model for estimating pseudotime and RNA velocity in multi-omic single-cell dataabstractMOTIVATION: RNA velocity has become a powerful tool for uncovering transcriptional dynamics in snapshot single-cell data. However, current RNA velocity approaches often assume constant transcriptional rates and treat genes independently with gene-specific times, which may introduce biases and deviate from biological realities. Here, we present InterVelo, a novel deep learning framework that simultaneously learns cellular pseudotime and RNA velocity. RESULTS: InterVelo leverages an unsupervised cellular time to guide RNA velocity estimation, while the estimated RNA velocity in turn refines the direction of pseudotime. By benchmarking InterVelo against existing methods on both simulated and real datasets, we demonstrate its superior performance in recovering pseudotime and RNA velocity. InterVelo yields more precise velocity estimations in terms of both direction and magnitude, with outstanding robustness across diverse scenarios. Furthermore, it successfully identifies driver genes and enables reliable gene activity enrichment analysis. The flexible architecture of InterVelo also allows for the integration of multi-omic data, enhancing its applicability to complex biological systems. AVAILABILITY AND IMPLEMENTATION: InterVelo is implemented using Python, and the code is available on GitHub https://github.com/yurouwang-rosie/InterVelo and has been archived with a DOI https://doi.org/10.5281/zenodo.16158798 for reproducibility. Yurou Wang, Zhixiang Lin, Tao Wang 0067 |
Bioinform. | 3 |
| 2024 | mbDriver: identifying driver microbes in microbial communities based on time-series microbiome dataabstractAlterations in human microbial communities are intricately linked to the onset and progression of diseases. Identifying the key microbes driving these community changes is crucial, as they may serve as valuable biomarkers for disease prevention, diagnosis, and treatment. However, there remains a need for further research to develop effective methods for addressing this critical task. This is primarily because defining the driver microbe requires consideration not only of each microbe's individual contributions but also their interactions. This paper introduces a novel framework, called mbDriver, for identifying driver microbes based on microbiome abundance data collected at discrete time points. mbDriver comprises three main components: (i) data preprocessing of time-series abundance data using smoothing splines based on the negative binomial distribution, (ii) parameter estimation for the generalized Lotka-Volterra (gLV) model using regularized least squares, and (iii) quantification of each microbe's contribution to the community's steady state by manipulating the causal graph implied by gLV equations. The performance of nonparametric spline-based denoising and regularized least squares estimation is comprehensively evaluated on simulated datasets, demonstrating superiority over existing methods. Furthermore, the practical applicability and effectiveness of mbDriver are showcased using a dietary fiber intervention dataset and an ulcerative colitis dataset. Notably, driver microbes identified in the dietary fiber intervention dataset exhibit significant effects on the abundances of short-chain fatty acids, while those identified in the ulcerative colitis dataset show a significant correlation with metabolism-related pathways. Xiaoxiu Tan, Chenhong Zhang, Tao Wang 0067 |
Briefings Bioinform. | 4 |
| 2024 | mbDecoda: a debiased approach to compositional data analysis for microbiome surveysabstractPotentially pathogenic or probiotic microbes can be identified by comparing their abundance levels between healthy and diseased populations, or more broadly, by linking microbiome composition with clinical phenotypes or environmental factors. However, in microbiome studies, feature tables provide relative rather than absolute abundance of each feature in each sample, as the microbial loads of the samples and the ratios of sequencing depth to microbial load are both unknown and subject to considerable variation. Moreover, microbiome abundance data are count-valued, often over-dispersed and contain a substantial proportion of zeros. To carry out differential abundance analysis while addressing these challenges, we introduce mbDecoda, a model-based approach for debiased analysis of sparse compositions of microbiomes. mbDecoda employs a zero-inflated negative binomial model, linking mean abundance to the variable of interest through a log link function, and it accommodates the adjustment for confounding factors. To efficiently obtain maximum likelihood estimates of model parameters, an Expectation Maximization algorithm is developed. A minimum coverage interval approach is then proposed to rectify compositional bias, enabling accurate and reliable absolute abundance analysis. Through extensive simulation studies and analysis of real-world microbiome datasets, we demonstrate that mbDecoda compares favorably with state-of-the-art methods in terms of effectiveness, robustness and reproducibility. Yuxuan Zong, Tao Wang 0067 |
Briefings Bioinform. | 3 |
| 2022 | MZINBVA: variational approximation for multilevel zero-inflated negative-binomial models for association analysis in microbiome surveysabstractAs our understanding of the microbiome has expanded, so has the recognition of its critical role in human health and disease, thereby emphasizing the importance of testing whether microbes are associated with environmental factors or clinical outcomes. However, many of the fundamental challenges that concern microbiome surveys arise from statistical and experimental design issues, such as the sparse and overdispersed nature of microbiome count data and the complex correlation structure among samples. For example, in the human microbiome project (HMP) dataset, the repeated observations across time points (level 1) are nested within body sites (level 2), which are further nested within subjects (level 3). Therefore, there is a great need for the development of specialized and sophisticated statistical tests. In this paper, we propose multilevel zero-inflated negative-binomial models for association analysis in microbiome surveys. We develop a variational approximation method for maximum likelihood estimation and inference. It uses optimization, rather than sampling, to approximate the log-likelihood and compute parameter estimates, provides a robust estimate of the covariance of parameter estimates and constructs a Wald-type test statistic for association testing. We evaluate and demonstrate the performance of our method using extensive simulation studies and an application to the HMP dataset. We have developed an R package MZINBVA to implement the proposed method, which is available from the GitHub repository https://github.com/liudoubletian/MZINBVA. Peirong Xu, Yueyao Du, Hui Lu 0004, Hongyu Zhao 0003, Tao Wang 0067 |
Briefings Bioinform. | 6 |
| 2022 | fastANCOM: a fast method for analysis of compositions of microbiomesabstractSUMMARY: Analysis of compositions of microbiomes (ANCOM) compares the absolute abundances of microbes between two or more ecosystems using relative abundances in specimens derived from these ecosystems. Despite its impressive performance, there are two drawbacks to ANCOM. First, with K microbes it requires fitting K(K-1)/2 models for log-ratios of counts, and so can be computationally intensive. Second, it does not output P-values for microbes detected as differentially abundant. We propose a fast implementation of ANCOM, fastANCOM, that fits only K models for log-transformed counts. fastANCOM provides P-values to declare statistical significance and outputs log fold changes of abundance between groups. We demonstrate that fastANCOM compares favorably with existing differential abundance testing methods in terms of running time, false discovery rate and power. AVAILABILITY AND IMPLEMENTATION: fastANCOM is available at https://github.com/ZRChao/fastANCOM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hongyu Zhao 0003, Tao Wang 0067 |
Bioinform. | 4 |
| 2022 | phyloMDA: an R package for phylogeny-aware microbiome data analysisabstractBACKGROUND: Modern sequencing technologies have generated low-cost microbiome survey datasets, across sample sites, conditions, and treatments, on an unprecedented scale and throughput. These datasets often come with a phylogenetic tree that provides a unique opportunity to examine how shared evolutionary history affects the different patterns in host-associated microbial communities. RESULTS: In this paper, we describe an R package, phyloMDA, for phylogeny-aware microbiome data analysis. It includes the Dirichlet-tree multinomial model for multivariate abundance data, tree-guided empirical Bayes estimation of microbial compositions, and tree-based multiscale regression methods with relative abundances as predictors. CONCLUSION: phyloMDA is a versatile and user-friendly tool to analyze microbiome data while incorporating the phylogenetic information and addressing some of the challenges posed by the data. Hongyu Zhao 0003, Tao Wang 0067 |
BMC Bioinform. | 5 |
| 2021 | Transformation and differential abundance analysis of microbiome data incorporating phylogenyabstractMOTIVATION: Microbiome data have proven extremely useful for understanding microbial communities and their impacts in health and disease. Although microbiome analysis methods and standards are evolving rapidly, obtaining meaningful and interpretable results from microbiome studies still requires careful statistical treatment. In particular, many existing and emerging methods for differential abundance (DA) analysis fail to account for the fact that microbiome data are high-dimensional and sparse, compositional, negatively and positively correlated and phylogenetically structured. To better describe microbiome data and improve the power of DA testing, there is still a great need for the continued development of appropriate statistical methodology. RESULTS: In this article, we propose a model-based approach for microbiome data transformation, and a phylogenetically informed procedure for DA testing based on the transformed data. First, we extend the Dirichlet-tree multinomial (DTM) to zero-inflated DTM for multivariate modeling of microbial counts, addressing data sparsity and correlation and phylogeny among bacterial taxa. Then, within this framework and using a Bayesian formulation, we introduce posterior mean transformation to convert raw counts into non-zero relative abundances that sum to one, accounting for the compositionality nature of microbiome data. Second, using the transformed data, we propose adaptive analysis of composition of microbiomes (adaANCOM) for DA testing by constructing log-ratios adaptively on the tree for each taxon, greatly reducing the computational complexity of ANCOM in high dimensions. Finally, we present extensive simulation studies, an analysis of HMP data across 18 body sites and 2 visits, and an application to a gut microbiome and malnutrition study, to investigate the performance of posterior mean transformation and adaANCOM. Comparisons with ANCOM and other DA testing procedures show that adaANCOM controls the false discovery rate well, allows for easy interpretation of the results, and is computationally efficient for high-dimensional problems. AVAILABILITY AND IMPLEMENTATION: The developed R package is available at https://github.com/ZRChao/adaANCOM. For replicability purposes, scripts for our simulations and data analysis are available at https://github.com/ZRChao/Papers_supplementary. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hongyu Zhao 0003, Tao Wang 0067 |
Bioinform. | 3 |
| 2020 | LPM: a latent probit model to characterize the relationship among complex traits using summary statistics from multiple GWASs and functional annotationsabstractMOTIVATION: Much effort has been made toward understanding the genetic architecture of complex traits and diseases. In the past decade, fruitful GWAS findings have highlighted the important role of regulatory variants and pervasive pleiotropy. Because of the accumulation of GWAS data on a wide range of phenotypes and high-quality functional annotations in different cell types, it is timely to develop a statistical framework to explore the genetic architecture of human complex traits by integrating rich data resources. RESULTS: In this study, we propose a unified statistical approach, aiming to characterize relationship among complex traits, and prioritize risk variants by leveraging regulatory information collected in functional annotations. Specifically, we consider a latent probit model (LPM) to integrate summary-level GWAS data and functional annotations. The developed computational framework not only makes LPM scalable to hundreds of annotations and phenotypes but also ensures its statistically guaranteed accuracy. Through comprehensive simulation studies, we evaluated LPM's performance and compared it with related methods. Then, we applied it to analyze 44 GWASs with 9 genic category annotations and 127 cell-type specific functional annotations. The results demonstrate the benefits of LPM and gain insights of genetic architecture of complex traits. AVAILABILITY AND IMPLEMENTATION: The LPM package, all simulation codes and real datasets in this study are available at https://github.com/mingjingsi/LPM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jingsi Ming, Tao Wang 0067, Can Yang 0002 |
Bioinform. | 2 |
| 2020 | Predicting viral exposure response from modeling the changes of co-expression networks using time series gene expression dataabstractBACKGROUND: Deciphering the relationship between clinical responses and gene expression profiles may shed light on the mechanisms underlying diseases. Most existing literature has focused on exploring such relationship from cross-sectional gene expression data. It is likely that the dynamic nature of time-series gene expression data is more informative in predicting clinical response and revealing the physiological process of disease development. However, it remains challenging to extract useful dynamic information from time-series gene expression data. RESULTS: We propose a statistical framework built on considering co-expression network changes across time from time series gene expression data. It first detects change point for co-expression networks and then employs a Bayesian multiple kernel learning method to predict exposure response. There are two main novelties in our method: the use of change point detection to characterize the co-expression network dynamics, and the use of kernel function to measure the similarity between subjects. Our algorithm allows exposure response prediction using dynamic network information across a collection of informative gene sets. Through parameter estimations, our model has clear biological interpretations. The performance of our method on the simulated data under different scenarios demonstrates that the proposed algorithm has better explanatory power and classification accuracy than commonly used machine learning algorithms. The application of our method to time series gene expression profiles measured in peripheral blood from a group of subjects with respiratory viral exposure shows that our method can predict exposure response at early stage (within 24 h) and the informative gene sets are enriched for pathways related to respiratory and influenza virus infection. CONCLUSIONS: The biological hypothesis in this paper is that the dynamic changes of the biological system are related to the clinical response. Our results suggest that when the relationship between the clinical response and a single gene or a gene set is not significant, we may benefit from studying the relationships among genes in gene sets that may lead to novel biological insights. Fangli Dong, Tao Wang 0067, Hui Lu 0004, Hongyu Zhao 0003 |
BMC Bioinform. | 3 |
| 2020 | An empirical Bayes approach to normalization and differential abundance testing for microbiome dataabstractBACKGROUND: Advances in DNA sequencing have offered researchers an unprecedented opportunity to better study the variety of species living in and on the human body. However, the analysis of microbiome data is complicated by several challenges. First, the sequencing depth may vary by orders of magnitude across samples. Second, species are rare and the data often contain many zeros. Third, the specimen is a fraction of the microbial ecosystem, and so the data are compositional carrying only relative information. Other characteristics of microbiome data include pronounced over-dispersion in taxon abundances, and the existence of a phylogenetic tree that relates all bacterial species. To address some of these challenges, microbiome analysis workflows often normalize the read counts prior to downstream analysis. However, there are limitations in the current literature on the normalization of microbiome data. RESULTS: Under the multinomial distribution for the read counts and a prior for the unknown proportions, we propose an empirical Bayes approach to microbiome data normalization. Using a tree-based extension of the Dirichlet prior, we further extend our method by incorporating the phylogenetic tree into the normalization process. We study the impact of normalization on differential abundance analysis. In the presence of tree structure, we propose a phylogeny-aware detection procedure. CONCLUSIONS: Extensive simulations and gut microbiome data applications are conducted to demonstrate the superior performance of our empirical Bayes method over other normalization methods, and over commonly-used methods for differential abundance testing. Original R scripts are available at GitHub (https://github.com/liudoubletian/eBay). Hongyu Zhao 0003, Tao Wang 0067 |
BMC Bioinform. | 3 |