Zhiguang Huo

dblp:119/6480 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0002-8032-4392ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 A Bayesian framework for genome-wide circadian rhythmicity biomarker detection
abstract
Circadian rhythms are endogenous $\sim $24-h cycles that significantly influence physiological and behavioral processes. These rhythms are governed by a transcriptional-translational feedback loop of core circadian genes and are essential for maintaining overall health. The study of circadian rhythms has expanded into various omics datasets, necessitating accurate analytical methodology for circadian biomarker detection. Here, we introduce a novel Bayesian framework for the detection of circadian rhythms in genome-wide transcriptomic applications that is capable of incorporating prior biological knowledge and adjusting for multiple testing issue via a false discovery rate (FDR) approach. Our framework leverages a Bayesian hierarchical model and employs a reverse jump Markov chain Monte Carlo technique for model selection. Through extensive simulations, our method, BayesCircRhy, demonstrated favorable FDR control over competing methods, robustness against heavier-tailed error distributions, and better performance compared with existing approaches. The method's efficacy was further validated in two RNA-sequencing data, including a human-restricted feeding data and a mouse aging data, where it successfully identified known and novel circadian genes.
Haocheng Ding, Lingsong Meng, Andrew J. Bryant, Chengguo Xing, Karyn A. Esser, Li Chen 0029, Yitong Feng, Zhiguang Huo
Briefings Bioinform.9
2025 A novel Bayesian hierarchical model for detecting differential circadian pattern in transcriptomic applications
abstract
Circadian rhythm plays a critical role in regulating various physiological processes, and disruptions in these rhythms have been linked to a wide range of diseases. Identifying molecular biomarkers showing differential circadian (DC) patterns between biological conditions or disease status is important for disease prevention, diagnosis, and treatment. However, circadian pattern is characterized by three key components: amplitude, phase, and MESOR, which poses a great challenge for DC analysis. Existing statistical methods focus on detecting differential shape (amplitude and phase) but often overlook MESOR difference. Additionally, these methods lack flexibility to incorporate external knowledge such as differential circadian information from similar clinical and biological context to improve the current DC analysis. To address these limitation, we introduce a novel Bayesian hierarchical model, BayesDCirc, designed for detecting differential circadian patterns in a two-group experimental design, which offer the advantage of testing MESOR difference and incorporating external knowledge. Benefiting from explicitly testing MESOR within the Bayesian modeling framework, BayesDCirc demonstrates superior FDR control over existing methods, with further performance improvement by leveraging external knowledge of DC genes. Applied to two real datasets, BayesDCirc successfully identify key circadian genes, particularly with external knowledge incorporated. The R package "BayesDCirc" for the method is publicly available on GitHub at https://github.com/lichen-lab/BayesDCirc.
Haocheng Ding, Zhiguang Huo, Li Chen 0029
Briefings Bioinform.3
2023 DiffCircaPipeline: a framework for multifaceted characterization of differential rhythmicity
abstract
SUMMARY: Circadian oscillations of gene expression regulate daily physiological processes, and their disruption is linked to many diseases. Circadian rhythms can be disrupted in a variety of ways, including differential phase, amplitude and rhythm fitness. Although many differential circadian biomarker detection methods have been proposed, a workflow for systematic detection of multifaceted differential circadian characteristics with accurate false positive control is not currently available. We propose a comprehensive and interactive pipeline to capture the multifaceted characteristics of differentially rhythmic biomarkers. Analysis outputs are accompanied by informative visualization and interactive exploration. The workflow is demonstrated in multiple case studies and is extensible to general omics applications. AVAILABILITY AND IMPLEMENTATION: R package, Shiny app and source code are available in GitHub (https://github.com/DiffCircaPipeline) and Zenodo (https://doi.org/10.5281/zenodo.7507989). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiangning Xue, Wei Zong, Zhiguang Huo, Kyle D. Ketchesin, Madeline R. Scott, Kaitlyn A Petersen, Ryan W. Logan, Marianne L. Seney, Colleen A. Mcclung, George C. Tseng
Bioinform.3
2021 Likelihood-based tests for detecting circadian rhythmicity and differential circadian patterns in transcriptomic applications
abstract
Circadian rhythmicity in transcriptomic profiles has been shown in many physiological processes, and the disruption of circadian patterns has been found to associate with several diseases. In this paper, we developed a series of likelihood-based methods to detect (i) circadian rhythmicity (denoted as LR_rhythmicity) and (ii) differential circadian patterns comparing two experimental conditions (denoted as LR_diff). In terms of circadian rhythmicity detection, we demonstrated that our proposed LR_rhythmicity could better control the type I error rate compared to existing methods under a wide variety of simulation settings. In terms of differential circadian patterns, we developed methods in detecting differential amplitude, differential phase, differential basal level and differential fit, which also successfully controlled the type I error rate. In addition, we demonstrated that the proposed LR_diff could achieve higher statistical power in detecting differential fit, compared to existing methods. The superior performance of LR_rhythmicity and LR_diff was demonstrated in four real data applications, including a brain aging data (gene expression microarray data of human postmortem brain), a time-restricted feeding data (RNA sequencing data of human skeletal muscles) and a scRNAseq data (single cell RNA sequencing data of mouse suprachiasmatic nucleus). An R package for our methods is publicly available on GitHub https://github.com/diffCircadian/diffCircadian.
Haocheng Ding, Lingsong Meng, Andrew C. Liu, Michelle L. Gumz, Andrew J. Bryant, Colleen A. Mcclung, George C. Tseng, Karyn A. Esser, Zhiguang Huo
Briefings Bioinform.9
2021 HCMMCNVs: hierarchical clustering mixture model of copy number variants detection using whole exome sequencing technology
abstract
SUMMARY: In this article, we introduce a hierarchical clustering and Gaussian mixture model with expectation-maximization (EM) algorithm for detecting copy number variants (CNVs) using whole exome sequencing (WES) data. The R shiny package 'HCMMCNVs' is also developed for processing user-provided bam files, running CNVs detection algorithm and conducting visualization. Through applying our approach to 325 cancer cell lines in 22 tumor types from Cancer Cell Line Encyclopedia (CCLE), we show that our algorithm is competitive with other existing methods and feasible in using multiple cancer cell lines for CNVs estimation. In addition, by applying our approach to WES data of 120 oral squamous cell carcinoma (OSCC) samples, our algorithm, using the tumor sample only, exhibits more power in detecting CNVs as compared with the methods using both tumors and matched normal counterparts. AVAILABILITY AND IMPLEMENTATION: HCMMCNVs R shiny software is freely available at github repository https://github.com/lunching/HCMM_CNVs.and Zenodo https://doi.org/10.5281/zenodo.4593371. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Chi Song, Shih-Chi Su, Zhiguang Huo, Suleyman Vural, James E. Galvin, Lun-Ching Chang
Bioinform.3
2020 P-value evaluation, variability index and biomarker categorization for adaptively weighted Fisher's meta-analysis method in omics applications
abstract
MOTIVATION: Meta-analysis methods have been widely used to combine results from multiple clinical or genomic studies to increase statistical powers and ensure robust and accurate conclusions. The adaptively weighted Fisher's method (AW-Fisher), initially developed for omics applications but applicable for general meta-analysis, is an effective approach to combine P-values from K independent studies and to provide better biological interpretability by characterizing which studies contribute to the meta-analysis. Currently, AW-Fisher suffers from the lack of fast P-value computation and variability estimate of AW weights. When the number of studies K is large, the 3K - 1 possible differential expression pattern categories generated by AW-Fisher can become intractable. In this paper, we develop an importance sampling scheme with spline interpolation to increase the accuracy and speed of the P-value calculation. We also apply bootstrapping to construct a variability index for the AW-Fisher weight estimator and a co-membership matrix to categorize (cluster) differentially expressed genes based on their meta-patterns for intuitive biological investigations. RESULTS: The superior performance of the proposed methods is shown in simulations as well as two real omics meta-analysis applications to demonstrate its insightful biological findings. AVAILABILITY AND IMPLEMENTATION: An R package AWFisher (calling C++) is available at Bioconductor and GitHub (https://github.com/Caleb-Huo/AWFisher), and all datasets and programing codes for this paper are available in the Supplementary Material. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhiguang Huo, Shaowu Tang, Yongseok Park, George C. Tseng
Bioinform.1
2019 MetaOmics: analysis pipeline and browser-based software suite for transcriptomic meta-analysis
abstract
SUMMARY: The rapid advances of omics technologies have generated abundant genomic data in public repositories and effective analytical approaches are critical to fully decipher biological knowledge inside these data. Meta-analysis combines multiple studies of a related hypothesis to improve statistical power, accuracy and reproducibility beyond individual study analysis. To date, many transcriptomic meta-analysis methods have been developed, yet few thoughtful guidelines exist. Here, we introduce a comprehensive analytical pipeline and browser-based software suite, called MetaOmics, to meta-analyze multiple transcriptomic studies for various biological purposes, including quality control, differential expression analysis, pathway enrichment analysis, differential co-expression network analysis, prediction, clustering and dimension reduction. The pipeline includes many public as well as >10 in-house transcriptomic meta-analytic methods with data-driven and biological-aim-driven strategies, hands-on protocols, an intuitive user interface and step-by-step instructions. AVAILABILITY AND IMPLEMENTATION: MetaOmics is freely available at https://github.com/metaOmics/metaOmics. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tianzhou Ma, Zhiguang Huo, Anche Kuo, Chien-Wei Lin, Silvia Liu, Tanbin Rahman, Lun-Ching Chang, Yongseok Park, Chi Song, Steffi Oesterreich, Etienne Sibille, George C. Tseng
Bioinform.2
2018 Image-derived generative modeling of pseudo-macromolecular structures - towards the statistical assessment of Electron CryoTomography template matching
Xiaodan Liang, Zhiguang Huo, Eric P. Xing, Min Xu 0009
BMVC4
2018 Meta-analytic principal component analysis in integrative omics application
abstract
Motivation: With the prevalent usage of microarray and massively parallel sequencing, numerous high-throughput omics datasets have become available in the public domain. Integrating abundant information among omics datasets is critical to elucidate biological mechanisms. Due to the high-dimensional nature of the data, methods such as principal component analysis (PCA) have been widely applied, aiming at effective dimension reduction and exploratory visualization. Results: In this article, we combine multiple omics datasets of identical or similar biological hypothesis and introduce two variations of meta-analytic framework of PCA, namely MetaPCA. Regularization is further incorporated to facilitate sparse feature selection in MetaPCA. We apply MetaPCA and sparse MetaPCA to simulations, three transcriptomic meta-analysis studies in yeast cell cycle, prostate cancer, mouse metabolism and a TCGA pan-cancer methylation study. The result shows improved accuracy, robustness and exploratory visualization of the proposed framework. Availability and implementation: An R package MetaPCA is available online. (http://tsenglab.biostat.pitt.edu/software.htm). Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Dongwan D. Kang, Zhiguang Huo, Yongseok Park, George C. Tseng
Bioinform.3
2017 MetaDCN: meta-analysis framework for differential co-expression network detection with an application in breast cancer
abstract
MOTIVATION: Gene co-expression network analysis from transcriptomic studies can elucidate gene-gene interactions and regulatory mechanisms. Differential co-expression analysis helps further detect alterations of regulatory activities in case/control comparison. Co-expression networks estimated from single transcriptomic study is often unstable and not generalizable due to cohort bias and limited sample size. With the rapid accumulation of publicly available transcriptomic studies, co-expression analysis combining multiple transcriptomic studies can provide more accurate and robust results. RESULTS: In this paper, we propose a meta-analytic framework for detecting differentially co-expressed networks (MetaDCN). Differentially co-expressed seed modules are first detected by optimizing an energy function via simulated annealing. Basic modules sharing common pathways are merged into pathway-centric supermodules and a Cytoscape plug-in (MetaDCNExplorer) is developed to visualize and explore the findings. We applied MetaDCN to two breast cancer applications: ER+/ER- comparison using five training and three testing studies, and ILC/IDC comparison with two training and two testing studies. We identified 20 and 4 supermodules for ER+/ER- and ILC/IDC comparisons, respectively. Ranking atop are 'immune response pathway' and 'complement cascades pathway' for ER comparison, and 'extracellular matrix pathway' for ILC/IDC comparison. Without the need for prior information, the results from MetaDCN confirm existing as well as discover novel disease mechanisms in a systems manner. AVAILABILITY AND IMPLEMENTATION: R package 'MetaDCN' and Cytoscape App 'MetaDCNExplorer' are available at http://tsenglab.biostat.pitt.edu/software.htm . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Cho-Yi Chen, Zhiguang Huo
Bioinform.5
2012 An R package suite for microarray meta-analysis in quality control, differentially expressed gene analysis and pathway enrichment detection
abstract
SUMMARY: With the rapid advances and prevalence of high-throughput genomic technologies, integrating information of multiple relevant genomic studies has brought new challenges. Microarray meta-analysis has become a frequently used tool in biomedical research. Little effort, however, has been made to develop a systematic pipeline and user-friendly software. In this article, we present MetaOmics, a suite of three R packages MetaQC, MetaDE and MetaPath, for quality control, differentially expressed gene identification and enriched pathway detection for microarray meta-analysis. MetaQC provides a quantitative and objective tool to assist study inclusion/exclusion criteria for meta-analysis. MetaDE and MetaPath were developed for candidate marker and pathway detection, which provide choices of marker detection, meta-analysis and pathway analysis methods. The system allows flexible input of experimental data, clinical outcome (case-control, multi-class, continuous or survival) and pathway databases. It allows missing values in experimental data and utilizes multi-core parallel computing for fast implementation. It generates informative summary output and visualization plots, operates on different operation systems and can be expanded to include new algorithms or combine different types of genomic data. This software suite provides a comprehensive tool to conveniently implement and compare various genomic meta-analysis pipelines. AVAILABILITY: http://www.biostat.pitt.edu/bioinfo/software.htm CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xingbin Wang, Dongwan D. Kang, Kui Shen, Chi Song, Shuya Lu, Lun-Ching Chang, Serena G. Liao, Zhiguang Huo, Shaowu Tang, Naftali Kaminski, Etienne Sibille, George C. Tseng
Bioinform.8