Zhen-Hao Guo

dblp:245/7622 · DBLP profile ↗
← Back
23ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 23 · 6 first-author · 12 since 2021
YearPublicationVenuePosition
2026 GraphLooper: predicting chromatin loops based on hierarchical multi-view graph pooling method
abstract
Chromatin loops serve as fundamental functional units of three-dimensional genome organization, playing pivotal roles in regulating gene expression and maintaining genomic spatial organization. Accurate identification of these fine-scale structures is crucial for advancing our understanding of cellular biological processes and the mechanisms underlying disease. However, due to the inherent complexity and dynamic of chromatin interactions, existing methods often fail to adequately characterize and capture multi-dimensional features. To address these limitations, we introduce GraphLooper, a novel framework using hierarchical multi-view graph pooling to enhance training and inference on large-scale data. GraphLooper transforms Hi-C data into a graph-structured representation, integrating multi-dimensional epigenomic features to construct a robust chromatin interaction model. Employing a hierarchical multi-view graph pooling mechanism, it effectively aggregates multi-scale features, enhancing representation learning. Evaluations across diverse cell lines demonstrate that GraphLooper outperforms state-of-the-art methods in prediction accuracy and generalization, particularly in capturing long-range chromatin interactions critical for precise spatial gene regulation.
Siguo Wang, Zhipeng Li 0002, Hailin Feng, Zhen-Hao Guo, Zuquan Hu, Qinhu Zhang, De-Shuang Huang
Briefings Bioinform.5
2025 NPENN: A Noise Perturbation Ensemble Neural Network for Microbiome Disease Phenotype Prediction
abstract
With advances in microbiomics, the crucial role of microbes in disease progression is increasingly recognized. However, predicting disease phenotypes using microbiome data remains challenging due to data complexity, heterogeneity, and limited model generalization. Current methods often depend on specific datasets and are vulnerable to adversarial attacks. To address these issues, this paper introduces a novel Noise Perturbation Ensemble Neural Network model (NPENN), which combines noise mechanisms with Gradient Boosting (GB) techniques for robust neural network ensemble learning. NPENN, validated on multiple microbiome datasets, shows superior accuracy and generalization compared to traditional methods, effectively handling data complexity and variability. This approach enhances model robustness and feature learning by integrating GB prior knowledge. Additionally, the study explores microbial community roles in various diseases, providing insights into disease mechanisms and potential biomarkers for personalized precision diagnosis and treatment strategies.
Yan Wu 0011, Qinhu Zhang, Siguo Wang, Zhen-Hao Guo
IEEE J. Biomed. Health Informatics5
2024 scCorrector: a robust method for integrating multi-study single-cell data
abstract
The advent of single-cell sequencing technologies has revolutionized cell biology studies. However, integrative analyses of diverse single-cell data face serious challenges, including technological noise, sample heterogeneity, and different modalities and species. To address these problems, we propose scCorrector, a variational autoencoder-based model that can integrate single-cell data from different studies and map them into a common space. Specifically, we designed a Study Specific Adaptive Normalization for each study in decoder to implement these features. scCorrector substantially achieves competitive and robust performance compared with state-of-the-art methods and brings novel insights under various circumstances (e.g. various batches, multi-omics, cross-species, and development stages). In addition, the integration of single-cell data and spatial data makes it possible to transfer information between different studies, which greatly expand the narrow range of genes covered by MERFISH technology. In summary, scCorrector can efficiently integrate multi-study single-cell datasets, thereby providing broad opportunities to tackle challenges emerging from noisy resources.
Zhen-Hao Guo, Siguo Wang, Qinhu Zhang, De-Shuang Huang
Briefings Bioinform.1
2023 scInterpreter: a knowledge-regularized generative model for interpretably integrating scRNA-seq data
abstract
BACKGROUND: The rapid emergence of single-cell RNA-seq (scRNA-seq) data presents remarkable opportunities for broad investigations through integration analyses. However, most integration models are black boxes that lack interpretability or are hard to train. RESULTS: To address the above issues, we propose scInterpreter, a deep learning-based interpretable model. scInterpreter substantially outperforms other state-of-the-art (SOTA) models in multiple benchmark datasets. In addition, scInterpreter is extensible and can integrate and annotate atlas scRNA-seq data. We evaluated the robustness of scInterpreter in a variety of situations. Through comparison experiments, we found that with a knowledge prior, the training process can be significantly accelerated. Finally, we conducted interpretability analysis for each dimension (pathway) of cell representation in the embedding space. CONCLUSIONS: The results showed that the cell representations obtained by scInterpreter are full of biological significance. Through weight sorting, we found several new genes related to pathways in PBMC dataset. In general, scInterpreter is an effective and interpretable integration tool. It is expected that scInterpreter will bring great convenience to the study of single-cell transcriptomics.
Zhen-Hao Guo, Siguo Wang, Qinhu Zhang, Jin-Ming Shi
BMC Bioinform.1
2023 In silico prediction methods of self-interacting proteins: an empirical and academic survey
Zhu-Hong You, Qinhu Zhang, Zhen-Hao Guo, Siguo Wang
Frontiers Comput. Sci.4
2023 Predicting the Sequence Specificities of DNA-Binding Proteins by DNA Fine-Tuned Language Model With Decaying Learning Rates
abstract
DNA-binding proteins (DBPs) play vital roles in the regulation of biological systems. Although there are already many deep learning methods for predicting the sequence specificities of DBPs, they face two challenges as follows. Classic deep learning methods for DBPs prediction usually fail to capture the dependencies between genomic sequences since their commonly used one-hot codes are mutually orthogonal. Besides, these methods usually perform poorly when samples are inadequate. To address these two challenges, we developed a novel language model for mining DBPs using human genomic data and ChIP-seq datasets with decaying learning rates, named DNA Fine-tuned Language Model (DFLM). It can capture the dependencies between genome sequences based on the context of human genomic data and then fine-tune the features of DBPs tasks using different ChIP-seq datasets. First, we compared DFLM with the existing widely used methods on 69 datasets and we achieved excellent performance. Moreover, we conducted comparative experiments on complex DBPs and small datasets. The results show that DFLM still achieved a significant improvement. Finally, through visualization analysis of one-hot encoding and DFLM, we found that one-hot encoding completely cut off the dependencies of DNA sequences themselves, while DFLM using language models can well represent the dependency of DNA sequences. Source code are available at: https://github.com/Deep-Bioinfo/DFLM.
Qinhu Zhang, Siguo Wang, Zhen-Hao Guo, De-Shuang Huang
IEEE ACM Trans. Comput. Biol. Bioinform.6
2022 DLoopCaller: A deep learning approach for predicting genome-wide chromatin loops by integrating accessible chromatin landscapes
abstract
In recent years, major advances have been made in various chromosome conformation capture technologies to further satisfy the needs of researchers for high-quality, high-resolution contact interactions. Discriminating the loops from genome-wide contact interactions is crucial for dissecting three-dimensional(3D) genome structure and function. Here, we present a deep learning method to predict genome-wide chromatin loops, called DLoopCaller, by combining accessible chromatin landscapes and raw Hi-C contact maps. Some available orthogonal data ChIA-PET/HiChIP and Capture Hi-C were used to generate positive samples with a wider contact matrix which provides the possibility to find more potential genome-wide chromatin loops. The experimental results demonstrate that DLoopCaller effectively improves the accuracy of predicting genome-wide chromatin loops compared to the state-of-the-art method Peakachu. Moreover, compared to two of most popular loop callers, such as HiCCUPS and Fit-Hi-C, DLoopCaller identifies some unique interactions. We conclude that a combination of chromatin landscapes on the one-dimensional genome contributes to understanding the 3D genome organization, and the identified chromatin loops reveal cell-type specificity and transcription factor motif co-enrichment across different cell lines and species.
Siguo Wang, Qinhu Zhang, Zhen-Hao Guo, Kyungsook Han, De-Shuang Huang
PLoS Comput. Biol.5
2022 Base-resolution prediction of transcription factor binding signals by a deep learning framework
abstract
Transcription factors (TFs) play an important role in regulating gene expression, thus the identification of the sites bound by them has become a fundamental step for molecular and cellular biology. In this paper, we developed a deep learning framework leveraging existing fully convolutional neural networks (FCN) to predict TF-DNA binding signals at the base-resolution level (named as FCNsignal). The proposed FCNsignal can simultaneously achieve the following tasks: (i) modeling the base-resolution signals of binding regions; (ii) discriminating binding or non-binding regions; (iii) locating TF-DNA binding regions; (iv) predicting binding motifs. Besides, FCNsignal can also be used to predict opening regions across the whole genome. The experimental results on 53 TF ChIP-seq datasets and 6 chromatin accessibility ATAC-seq datasets show that our proposed framework outperforms some existing state-of-the-art methods. In addition, we explored to use the trained FCNsignal to locate all potential TF-DNA binding regions on a whole chromosome and predict DNA sequences of arbitrary length, and the results show that our framework can find most of the known binding regions and accept sequences of arbitrary length. Furthermore, we demonstrated the potential ability of our framework in discovering causal disease-associated single-nucleotide polymorphisms (SNPs) through a series of experiments.
Qinhu Zhang, Siguo Wang, Zhen-Hao Guo, Qi Liu 0019, De-Shuang Huang
PLoS Comput. Biol.5
2021 Protein-Protein Interaction Prediction by Integrating Sequence Information and Heterogeneous Network Representation
Xiao-Rui Su 0001, Zhu-Hong You, Zhen-Hao Guo
ICIC (3)5
2021 MeSHHeading2vec: a new method for representing MeSH headings as vectors based on graph embedding algorithm
abstract
Effectively representing Medical Subject Headings (MeSH) headings (terms) such as disease and drug as discriminative vectors could greatly improve the performance of downstream computational prediction models. However, these terms are often abstract and difficult to quantify. In this paper, we converted the MeSH tree structure into a relationship network and applied several graph embedding algorithms on it to represent these terms. Specifically, the relationship network consisting of nodes (MeSH headings) and edges (relationships), which can be constructed by the tree num. Then, five graph embedding algorithms including DeepWalk, LINE, SDNE, LAP and HOPE were implemented on the relationship network to represent MeSH headings as vectors. In order to evaluate the performance of the proposed methods, we carried out the node classification and relationship prediction tasks. The results show that the MeSH headings characterized by graph embedding algorithms can not only be treated as an independent carrier for representation, but also can be utilized as additional information to enhance the representation ability of vectors. Thus, it can serve as an input and continue to play a significant role in any computational models related to disease, drug, microbe, etc. Besides, our method holds great hope to inspire relevant researchers to study the representation of terms in this network perspective.
Zhen-Hao Guo, Zhu-Hong You, De-Shuang Huang, Kai Zheng 0020
Briefings Bioinform.1
2021 A learning-based method to predict LncRNA-disease associations by combining CNN and ELM
abstract
BACKGROUND: lncRNAs play a critical role in numerous biological processes and life activities, especially diseases. Considering that traditional wet experiments for identifying uncovered lncRNA-disease associations is limited in terms of time consumption and labor cost. It is imperative to construct reliable and efficient computational models as addition for practice. Deep learning technologies have been proved to make impressive contributions in many areas, but the feasibility of it in bioinformatics has not been adequately verified. RESULTS: In this paper, a machine learning-based model called LDACE was proposed to predict potential lncRNA-disease associations by combining Extreme Learning Machine (ELM) and Convolutional Neural Network (CNN). Specifically, the representation vectors are constructed by integrating multiple types of biology information including functional similarity and semantic similarity. Then, CNN is applied to mine both local and global features. Finally, ELM is chosen to carry out the prediction task to detect the potential lncRNA-disease associations. The proposed method achieved remarkable Area Under Receiver Operating Characteristic Curve of 0.9086 in Leave-one-out cross-validation and 0.8994 in fivefold cross-validation, respectively. In addition, 2 kinds of case studies based on lung cancer and endometrial cancer indicate the robustness and efficiency of LDACE even in a real environment. CONCLUSIONS: Substantial results demonstrated that the proposed model is expected to be an auxiliary tool to guide and assist biomedical research, and the close integration of deep learning and biology big data will provide life sciences with novel insights.
Zhen-Hao Guo, Zhu-Hong You, Meineng Wang
BMC Bioinform.1
2021 Learning Representation of Molecules in Association Network for Predicting Intermolecular Associations
abstract
A key aim of post-genomic biomedical research is to systematically understand molecules and their interactions in human cells. Multiple biomolecules coordinate to sustain life activities, and interactions between various biomolecules are interconnected. However, existing studies usually only focusing on associations between two or very limited types of molecules. In this study, we propose a network representation learning based computational framework MAN-SDNE to predict any intermolecular associations. More specifically, we constructed a large-scale molecular association network of multiple biomolecules in human by integrating associations among long non-coding RNA, microRNA, protein, drug, and disease, containing 6,528 molecular nodes, 9 kind of,105,546 associations. And then, the feature of each node is represented by its network proximity and attribute features. Furthermore, these features are used to train Random Forest classifier to predict intermolecular associations. MAN-SDNE achieves a remarkable performance with an AUC of 0.9552 and an AUPR of 0.9338 under five-fold cross-validation. To indicate the ability to predict specific types of interactions, a case study for predicting lncRNA-protein interactions using MAN-SDNE is also executed. Experimental results demonstrate this work offers a systematic insight for understanding the synergistic associations between molecules and complex diseases and provides a network-based computational tool to systematically explore intermolecular interactions.
Zhu-Hong You, Zhen-Hao Guo, De-Shuang Huang, Keith C. C. Chan
IEEE ACM Trans. Comput. Biol. Bioinform.3
2020 Predicting Drug-Target Interactions by Node2vec Node Embedding in Molecular Associations Network
Zhu-Hong You, Zhen-Hao Guo, Gong-Xu Luo
ICIC (2)3
2020 Inferring Drug-miRNA Associations by Integrating Drug SMILES and MiRNA Sequence Information
Zhen-Hao Guo, Zhu-Hong You, Liping Li 0003
ICIC (2)1
2020 A Highly Efficient Biomolecular Network Representation Model for Predicting Drug-Disease Associations
Hanjing Jiang, Zhu-Hong You, Lun Hu, Zhen-Hao Guo, Leon Wong
ICIC (3)4
2020 A Unified Deep Biological Sequence Representation Learning with Pretrained Encoder-Decoder Model
Zhu-Hong You, Xiao-Rui Su 0001, De-Shuang Huang, Zhen-Hao Guo
ICIC (2)5
2020 A Novel Computational Method for Predicting LncRNA-Disease Associations from Heterogeneous Information Network with SDNE Embedding Model
Ping Zhang 0027, Bo-Wei Zhao, Leon Wong, Zhu-Hong You, Zhen-Hao Guo
ICIC (2)5
2020 RPI-SE: a stacking ensemble learning framework for ncRNA-protein interactions prediction using sequence information
abstract
BACKGROUND: The interactions between non-coding RNAs (ncRNA) and proteins play an essential role in many biological processes. Several high-throughput experimental methods have been applied to detect ncRNA-protein interactions. However, these methods are time-consuming and expensive. Accurate and efficient computational methods can assist and accelerate the study of ncRNA-protein interactions. RESULTS: In this work, we develop a stacking ensemble computational framework, RPI-SE, for effectively predicting ncRNA-protein interactions. More specifically, to fully exploit protein and RNA sequence feature, Position Weight Matrix combined with Legendre Moments is applied to obtain protein evolutionary information. Meanwhile, k-mer sparse matrix is employed to extract efficient feature of ncRNA sequences. Finally, an ensemble learning framework integrated different types of base classifier is developed to predict ncRNA-protein interactions using these discriminative features. The accuracy and robustness of RPI-SE was evaluated on three benchmark data sets under five-fold cross-validation and compared with other state-of-the-art methods. CONCLUSIONS: The results demonstrate that RPI-SE is competent for ncRNA-protein interactions prediction task with high accuracy and robustness. It's anticipated that this work can provide a computational prediction tool to advance ncRNA-protein interactions related biomedical research.
Zhu-Hong You, Meineng Wang, Zhen-Hao Guo, Ji-Ren Zhou
BMC Bioinform.4
2020 iCDA-CGR: Identification of circRNA-disease associations based on Chaos Game Representation
abstract
Found in recent research, tumor cell invasion, proliferation, or other biological processes are controlled by circular RNA. Understanding the association between circRNAs and diseases is an important way to explore the pathogenesis of complex diseases and promote disease-targeted therapy. Most methods, such as k-mer and PSSM, based on the analysis of high-throughput expression data have the tendency to think functionally similar nucleic acid lack direct linear homology regardless of positional information and only quantify nonlinear sequence relationships. However, in many complex diseases, the sequence nonlinear relationship between the pathogenic nucleic acid and ordinary nucleic acid is not much different. Therefore, the analysis of positional information expression can help to predict the complex associations between circRNA and disease. To fill up this gap, we propose a new method, named iCDA-CGR, to predict the circRNA-disease associations. In particular, we introduce circRNA sequence information and quantifies the sequence nonlinear relationship of circRNA by Chaos Game Representation (CGR) technology based on the biological sequence position information for the first time in the circRNA-disease prediction model. In the cross-validation experiment, our method achieved 0.8533 AUC, which was significantly higher than other existing methods. In the validation of independent data sets including circ2Disease, circRNADisease and CRDD, the prediction accuracy of iCDA-CGR reached 95.18%, 90.64% and 95.89%. Moreover, in the case studies, 19 of the top 30 circRNA-disease associations predicted by iCDA-CGR on circRDisease dataset were confirmed by newly published literature. These results demonstrated that iCDA-CGR has outstanding robustness and stability, and can provide highly credible candidates for biological experiments.
Kai Zheng 0020, Zhu-Hong You, Jianqiang Li 0001, Lei Wang 0121, Zhen-Hao Guo
PLoS Comput. Biol.5
2019 Combining LSTM Network Model and Wavelet Transform for Predicting Self-interacting Proteins
Zhu-Hong You, Liping Li 0003, Zhen-Hao Guo, Pengwei Hu 0001, Hanjing Jiang
ICIC (1)4
2019 Combining High Speed ELM with a CNN Feature Encoding to Predict LncRNA-Disease Associations
Zhen-Hao Guo, Zhu-Hong You, Liping Li 0003
ICIC (2)1
2019 Combining Evolutionary Information and Sparse Bayesian Probability Model to Accurately Predict Self-interacting Proteins
Zhu-Hong You, Zhen-Hao Guo, Kai Zheng 0020
ICIC (2)5
2019 In Silico Identification of Anticancer Peptides with Stacking Heterogeneous Ensemble Learning Model and Sequence Information
Zhu-Hong You, Zhen-Hao Guo
ICIC (2)5