EDBT 2026 Demo / reviewers in the wild / expert
Tianzhou Ma
dblp:202/9212
· DBLP profile ↗
9ranked-venue papers
1as first author
7since 2021 · last 2024
0000-0003-3605-0811ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | dCCA: detecting differential covariation patterns between two types of high-throughput omics dataabstractMOTIVATION: The advent of multimodal omics data has provided an unprecedented opportunity to systematically investigate underlying biological mechanisms from distinct yet complementary angles. However, the joint analysis of multi-omics data remains challenging because it requires modeling interactions between multiple sets of high-throughput variables. Furthermore, these interaction patterns may vary across different clinical groups, reflecting disease-related biological processes. RESULTS: We propose a novel approach called Differential Canonical Correlation Analysis (dCCA) to capture differential covariation patterns between two multivariate vectors across clinical groups. Unlike classical Canonical Correlation Analysis, which maximizes the correlation between two multivariate vectors, dCCA aims to maximally recover differentially expressed multivariate-to-multivariate covariation patterns between groups. We have developed computational algorithms and a toolkit to sparsely select paired subsets of variables from two sets of multivariate variables while maximizing the differential covariation. Extensive simulation analyses demonstrate the superior performance of dCCA in selecting variables of interest and recovering differential correlations. We applied dCCA to the Pan-Kidney cohort from the Cancer Genome Atlas Program database and identified differentially expressed covariations between noncoding RNAs and gene expressions. AVAILABILITY AND IMPLEMENTATION: The R package that implements dCCA is available at https://github.com/hwiyoungstat/dCCA. Hwiyoung Lee, Tianzhou Ma, Hongjie Ke, Zhenyao Ye |
Briefings Bioinform. | 2 |
| 2024 | TIPS: a novel pathway-guided joint model for transcriptome-wide association studiesabstractIn the past two decades, genome-wide association studies (GWAS) have pinpointed numerous SNPs linked to human diseases and traits, yet many of these SNPs are in non-coding regions and hard to interpret. Transcriptome-wide association studies (TWAS) integrate GWAS and expression reference panels to identify the associations at gene level with tissue specificity, potentially improving the interpretability. However, the list of individual genes identified from univariate TWAS contains little unifying biological theme, leaving the underlying mechanisms largely elusive. In this paper, we propose a novel multivariate TWAS method that Incorporates Pathway or gene Set information, namely TIPS, to identify genes and pathways most associated with complex polygenic traits. We jointly modeled the imputation and association steps in TWAS, incorporated a sparse group lasso penalty in the model to induce selection at both gene and pathway levels and developed an expectation-maximization algorithm to estimate the parameters for the penalized likelihood. We applied our method to three different complex traits: systolic and diastolic blood pressure, as well as a brain aging biomarker white matter brain age gap in UK Biobank and identified critical biologically relevant pathways and genes associated with these traits. These pathways cannot be detected by traditional univariate TWAS + pathway enrichment analysis approach, showing the power of our model. We also conducted comprehensive simulations with varying heritability levels and genetic architectures and showed our method outperformed other established TWAS methods in feature selection, statistical power, and prediction. The R package that implements TIPS is available at https://github.com/nwang123/TIPS. Zhenyao Ye, Tianzhou Ma |
Briefings Bioinform. | 3 |
| 2024 | Systems approach for congruence and selection of cancer models towards precision medicineabstractCancer models are instrumental as a substitute for human studies and to expedite basic, translational, and clinical cancer research. For a given cancer type, a wide selection of models, such as cell lines, patient-derived xenografts, organoids and genetically modified murine models, are often available to researchers. However, how to quantify their congruence to human tumors and to select the most appropriate cancer model is a largely unsolved issue. Here, we present Congruence Analysis and Selection of CAncer Models (CASCAM), a statistical and machine learning framework for authenticating and selecting the most representative cancer models in a pathway-specific manner using transcriptomic data. CASCAM provides harmonization between human tumor and cancer model omics data, systematic congruence quantification, and pathway-based topological visualization to determine the most appropriate cancer model selection. The systems approach is presented using invasive lobular breast carcinoma (ILC) subtype and suggesting CAMA1 followed by UACC3133 as the most representative cell lines for ILC research. Two additional case studies for triple negative breast cancer (TNBC) and patient-derived xenograft/organoid (PDX/PDO) are further investigated. CASCAM is generalizable to any cancer subtype and will authenticate cancer models for faithful non-human preclinical research towards precision medicine. Osama Shah, Yu-Chiao Chiu, Tianzhou Ma, Jennifer M. Atkinson, Steffi Oesterreich, Adrian V. Lee, George C. Tseng |
PLoS Comput. Biol. | 4 |
| 2023 | A Novel Neighborhood Rough Set-Based Feature Selection Method and Its Application to Biomarker Identification of SchizophreniaabstractFeature selection can disclose biomarkers of mental disorders that have unclear biological mechanisms. Although neighborhood rough set (NRS) has been applied to discover important sparse features, it has hardly ever been utilized in neuroimaging-based biomarker identification, probably due to the inadequate feature evaluation metric and incomplete information provided under a single-granularity. Here, we propose a new NRS-based feature selection method and successfully identify brain functional connectivity biomarkers of schizophrenia (SZ) using functional magnetic resonance imaging (fMRI) data. Specifically, we develop a new weighted metric based on NRS combined with information entropy to evaluate the capacity of features in distinguishing different groups. Inspired by multi-granularity information maximization theory, we further take advantage of the complementary information from different neighborhood sizes via a multi-granularity fusion to obtain the most discriminative and stable features. For validation, we compare our method with six popular feature selection methods using three public omics datasets as well as resting-state fMRI data of 393 SZ patients and 429 healthy controls. Results show that our method obtained higher classification accuracies on both omics data (100.0%, 88.6%, and 72.2% for three omics datasets, respectively) and fMRI data (93.9% for main dataset, and 76.3% and 83.8% for two independent datasets, respectively). Moreover, our findings reveal biologically meaningful substrates of SZ, notably involving the connectivity between the thalamus and superior temporal gyrus as well as between the postcentral gyrus and calcarine gyrus. Taken together, we propose a new NRS-based feature selection method that shows the potential of exploring effective and sparse neuroimaging-based biomarkers of mental disorders. Peter V. Kochunov, Theo G. M. van Erp, Tianzhou Ma, Vince D. Calhoun, Yuhui Du |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | A statistical framework for recovering pseudo-dynamic networks from static dataabstractMOTIVATION: The collection of temporal or perturbed data is often a prerequisite for reconstructing dynamic networks in most cases. However, these types of data are seldom available for genomic studies in medicine, thus significantly limiting the use of dynamic networks to characterize the biological principles underlying human health and diseases. RESULTS: We proposed a statistical framework to recover disease risk-associated pseudo-dynamic networks (DRDNet) from steady-state data. We incorporated a varying coefficient model with multiple ordinary differential equations to learn a series of networks. We analyzed the publicly available Genotype-Tissue Expression data to construct networks associated with hypertension risk, and biological findings showed that key genes constituting these networks had pivotal and biologically relevant roles associated with the vascular system. We also provided the selection consistency of the proposed learning procedure and evaluated its utility through extensive simulations. AVAILABILITY AND IMPLEMENTATION: DRDNet is implemented in the R language, and the source codes are available at https://github.com/chencxxy28/DRDnet/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chixiang Chen, Biyi Shen, Tianzhou Ma, Rongling Wu |
Bioinform. | 3 |
| 2022 | High-dimension to high-dimension screening for detecting genome-wide epigenetic and noncoding RNA regulators of gene expressionabstractMOTIVATION: The advancement of high-throughput technology characterizes a wide variety of epigenetic modifications and noncoding RNAs across the genome involved in disease pathogenesis via regulating gene expression. The high dimensionality of both epigenetic/noncoding RNA and gene expression data make it challenging to identify the important regulators of genes. Conducting univariate test for each possible regulator-gene pair is subject to serious multiple comparison burden, and direct application of regularization methods to select regulator-gene pairs is computationally infeasible. Applying fast screening to reduce dimension first before regularization is more efficient and stable than applying regularization methods alone. RESULTS: We propose a novel screening method based on robust partial correlation to detect epigenetic and noncoding RNA regulators of gene expression over the whole genome, a problem that includes both high-dimensional predictors and high-dimensional responses. Compared to existing screening methods, our method is conceptually innovative that it reduces the dimension of both predictor and response, and screens at both node (regulators or genes) and edge (regulator-gene pairs) levels. We develop data-driven procedures to determine the conditional sets and the optimal screening threshold, and implement a fast iterative algorithm. Simulations and applications to long noncoding RNA and microRNA regulation in Kidney cancer and DNA methylation regulation in Glioblastoma Multiforme illustrate the validity and advantage of our method. AVAILABILITY AND IMPLEMENTATION: The R package, related source codes and real datasets used in this article are provided at https://github.com/kehongjie/rPCor. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hongjie Ke, Zhao Ren, Jianfei Qi, George C. Tseng, Zhenyao Ye, Tianzhou Ma |
Bioinform. | 7 |
| 2021 | ICN: extracting interconnected communities in gene co-expression networksabstractMOTIVATION: The analysis of gene co-expression network (GCN) is critical in examining the gene-gene interactions and learning the underlying complex yet highly organized gene regulatory mechanisms. Numerous clustering methods have been developed to detect communities of co-expressed genes in the large network. The assumed independent community structure, however, can be oversimplified and may not adequately characterize the complex biological processes. RESULTS: We develop a new computational package to extract interconnected communities from gene co-expression network. We consider a pair of communities be interconnected if a subset of genes from one community is correlated with a subset of genes from another community. The interconnected community structure is more flexible and provides a better fit to the empirical co-expression matrix. To overcome the computational challenges, we develop efficient algorithms by leveraging advanced graph norm shrinkage approach. We validate and show the advantage of our method by extensive simulation studies. We then apply our interconnected community detection method to an RNA-seq data from The Cancer Genome Atlas (TCGA) Acute Myeloid Leukemia (AML) study and identify essential interacting biological pathways related to the immune evasion mechanism of tumor cells. AVAILABILITY: The software is available at Github: https://github.com/qwu1221/ICN and Figshare: https://figshare.com/articles/software/ICN-package/13229093. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tianzhou Ma, Qingzhi Liu, Donald K. Milton |
Bioinform. | 2 |
| 2019 | MetaOmics: analysis pipeline and browser-based software suite for transcriptomic meta-analysisabstractSUMMARY: The rapid advances of omics technologies have generated abundant genomic data in public repositories and effective analytical approaches are critical to fully decipher biological knowledge inside these data. Meta-analysis combines multiple studies of a related hypothesis to improve statistical power, accuracy and reproducibility beyond individual study analysis. To date, many transcriptomic meta-analysis methods have been developed, yet few thoughtful guidelines exist. Here, we introduce a comprehensive analytical pipeline and browser-based software suite, called MetaOmics, to meta-analyze multiple transcriptomic studies for various biological purposes, including quality control, differential expression analysis, pathway enrichment analysis, differential co-expression network analysis, prediction, clustering and dimension reduction. The pipeline includes many public as well as >10 in-house transcriptomic meta-analytic methods with data-driven and biological-aim-driven strategies, hands-on protocols, an intuitive user interface and step-by-step instructions. AVAILABILITY AND IMPLEMENTATION: MetaOmics is freely available at https://github.com/metaOmics/metaOmics. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tianzhou Ma, Zhiguang Huo, Anche Kuo, Chien-Wei Lin, Silvia Liu, Tanbin Rahman, Lun-Ching Chang, Yongseok Park, Chi Song, Steffi Oesterreich, Etienne Sibille, George C. Tseng |
Bioinform. | 1 |
| 2018 | Bayesian integrative model for multi-omics data with missingnessabstractMotivation: Integrative analysis of multi-omics data from different high-throughput experimental platforms provides valuable insight into regulatory mechanisms associated with complex diseases, and gains statistical power to detect markers that are otherwise overlooked by single-platform omics analysis. In practice, a significant portion of samples may not be measured completely due to insufficient tissues or restricted budget (e.g. gene expression profile are measured but not methylation). Current multi-omics integrative methods require complete data. A common practice is to ignore samples with any missing platform and perform complete case analysis, which leads to substantial loss of statistical power. Methods: In this article, inspired by the popular Integrative Bayesian Analysis of Genomics data (iBAG), we propose a full Bayesian model that allows incorporation of samples with missing omics data. Results: Simulation results show improvement of the new full Bayesian approach in terms of outcome prediction accuracy and feature selection performance when sample size is limited and proportion of missingness is large. When sample size is large or the proportion of missingness is low, incorporating samples with missingness may introduce extra inference uncertainty and generate worse prediction and feature selection performance. To determine whether and how to incorporate samples with missingness, we propose a self-learning cross-validation (CV) decision scheme. Simulations and a real application on child asthma dataset demonstrate superior performance of the CV decision scheme when various types of missing mechanisms are evaluated. Availability and implementation: Freely available on the GitHub at https://github.com/CHPGenetics/FBM. Supplementary information: Supplementary data are available at Bioinformatics online. Tianzhou Ma, Gong Tang, Qi Yan 0008, Ting Wang 0003, Juan C. Celedón, Wei Chen 0074, George C. Tseng |
Bioinform. | 2 |