VLDB 2026 Research / reviewers in the wild / expert
Qinhu Zhang
dblp:156/6796
· DBLP profile ↗
36ranked-venue papers
11as first author
30since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 32 · 10 first-author · 26 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Collaborative Graph Agents for LLM-Based Graph Reasoning
Jindi Wang, Meng Xing, Qinhu Zhang, Dijun Gong |
ICIC (14) | 5 |
| 2026 | GraphLooper: predicting chromatin loops based on hierarchical multi-view graph pooling methodabstractChromatin loops serve as fundamental functional units of three-dimensional genome organization, playing pivotal roles in regulating gene expression and maintaining genomic spatial organization. Accurate identification of these fine-scale structures is crucial for advancing our understanding of cellular biological processes and the mechanisms underlying disease. However, due to the inherent complexity and dynamic of chromatin interactions, existing methods often fail to adequately characterize and capture multi-dimensional features. To address these limitations, we introduce GraphLooper, a novel framework using hierarchical multi-view graph pooling to enhance training and inference on large-scale data. GraphLooper transforms Hi-C data into a graph-structured representation, integrating multi-dimensional epigenomic features to construct a robust chromatin interaction model. Employing a hierarchical multi-view graph pooling mechanism, it effectively aggregates multi-scale features, enhancing representation learning. Evaluations across diverse cell lines demonstrate that GraphLooper outperforms state-of-the-art methods in prediction accuracy and generalization, particularly in capturing long-range chromatin interactions critical for precise spatial gene regulation. Siguo Wang, Zhipeng Li 0002, Hailin Feng, Zhen-Hao Guo, Zuquan Hu, Qinhu Zhang, De-Shuang Huang |
Briefings Bioinform. | 9 |
| 2025 | RGCN-BA: relational graph convolutional network with batch awareness for single-cell RNA sequencing clusteringabstractSingle-cell RNA sequencing (scRNA-seq) technology has opened new frontiers in biomedical research, offering insights into cellular heterogeneity. Accurate cell clustering and batch effect correction are essential in single-cell RNA sequencing (scRNA-seq) data analysis, forming the foundation for downstream steps. However, most methods handle these tasks separately, limiting their applicability across diverse datasets. To address these challenges, we introduce Relational Graph Convolutional Network with Batch Awareness (RGCN-BA), a deep learning framework that integrates cell clustering and batch effect correction into a unified model. For multi-batch datasets, RGCN-BA leverages relational graph convolutional network to process batch information as distinct edge types, followed by a batch correction layer for global alignment. For single-batch data, it functions with a single edge type. Experiments on both multi-batch and single-batch datasets demonstrate that RGCN-BA outperforms both specialized clustering methods and batch effect correction methods. This versatility in handling both tasks positions RGCN-BA as a powerful tool for enhancing scRNA-seq data analysis. Pengrui Teng, Zheyu Wu, Yuna Zhang, Zhisen Shen, Qinhu Zhang, De-Shuang Huang |
Briefings Bioinform. | 6 |
| 2025 | scRGCL: a cell type annotation method for single-cell RNA-seq data using residual graph convolutional neural network with contrastive learningabstractCell type annotation is a critical step in analyzing single-cell RNA sequencing (scRNA-seq) data. A large number of deep learning (DL)-based methods have been proposed to annotate cell types of scRNA-seq data and have achieved impressive results. However, there are several limitations to these methods. First, they do not fully exploit cell-to-cell differential features. Second, they are developed based on shallow features and lack of flexibility in integrating high-order features in the data. Finally, the low-dimensional gene features may lead to overfitting in neural networks. To overcome those limitations, we propose a novel DL-based model, cell type annotation of single-cell RNA-seq data using residual graph convolutional neural network with contrastive learning (scRGCL), based on residual graph convolutional neural network and contrastive learning for cell type annotation of single-cell RNA-seq data. scRGCL mainly consists of a residual graph convolutional neural network, contrastive learning, and weight freezing. A residual graph convolutional neural network is utilized to extract complex high-order features from data. Contrastive learning can help the model learn meaningful cell-to-cell differential features. Weight freezing can avoid overfitting and help the model discover the impact of specific gene expression on cell type annotation. To verify the effectiveness of scRGCL, we compared its performance with six methods (three shallow learning algorithms and three state-of-the-art DL-based methods) on eight single-cell benchmark datasets from two species (seven in human and one in mouse). Experimental results not only show that scRGCL outperforms competing methods but also demonstrate the generalizability of scRGCL for cell type annotation. scRGCL is available at https://github.com/nathanyl/scRGCL. Lin Yuan 0001, Shengguo Sun, Qinhu Zhang, Lan Ye, Chun-Hou Zheng 0001, De-Shuang Huang |
Briefings Bioinform. | 4 |
| 2025 | A brief survey of deep learning-based models for CircRNA-protein binding sites predictionabstractCircRNAs are a particular single-stranded, circular structure and “non-coding” RNA molecules, with various biological functions . Existing studies have demonstrated the fundamental role of circRNAs in gene expression regulation and their significant involvement in the development of diverse complex diseases. Predicting the protein binding sites in circRNA can aid in comprehending the regulation mechanism involved in circRNA-protein binding during gene expression and facilitate the investigation of potential diagnosis and treatment strategies for complex diseases. This review begins by introducing the concept and functions of circRNAs, as well as their involvement in gene expression regulation . Then, some critical and publicly accessible databases about circRNA annotation, protein annotation, circRNA-protein binding were listed. Next, we present a brief introduction to the computational model for predicting circRNA-protein binding, followed by model performance comparison and suggestions for non-computer science experts on model selection. Finally, we examine the problems, limitations, and advantages of computational models and explore the further direction of circRNA-protein prediction, such as developing new and complex computational models, introducing complex biological sequence encoding schemes, and integrating additional biological data related to circRNA-protein binding. Zhen Shen 0003, Lin Yuan 0001, Wenzheng Bao, Siguo Wang, Qinhu Zhang, De-Shuang Huang |
Neurocomputing | 5 |
| 2025 | NPENN: A Noise Perturbation Ensemble Neural Network for Microbiome Disease Phenotype PredictionabstractWith advances in microbiomics, the crucial role of microbes in disease progression is increasingly recognized. However, predicting disease phenotypes using microbiome data remains challenging due to data complexity, heterogeneity, and limited model generalization. Current methods often depend on specific datasets and are vulnerable to adversarial attacks. To address these issues, this paper introduces a novel Noise Perturbation Ensemble Neural Network model (NPENN), which combines noise mechanisms with Gradient Boosting (GB) techniques for robust neural network ensemble learning. NPENN, validated on multiple microbiome datasets, shows superior accuracy and generalization compared to traditional methods, effectively handling data complexity and variability. This approach enhances model robustness and feature learning by integrating GB prior knowledge. Additionally, the study explores microbial community roles in various diseases, providing insights into disease mechanisms and potential biomarkers for personalized precision diagnosis and treatment strategies. Yan Wu 0011, Qinhu Zhang, Siguo Wang, Zhen-Hao Guo |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Dynamic Link Prediction for New Nodes in Temporal Graph NetworksabstractModeling temporal networks for dynamic link prediction of new nodes has many real-world applications, such as providing relevant item recommendations to new customers in recommender systems and suggesting appropriate posts to new users on social platforms. Unlike old nodes, new nodes have few historical links, which poses a challenge to the dynamic link prediction task. Most existing dynamic models treat all nodes equally and are not specialized for new nodes, resulting in suboptimal performance. In this paper, we consider the dynamic link prediction of new nodes as a few-shot problem and propose a novel model based on the meta-learning principle to effectively mitigate this problem. Specifically, we devise a meta-learning-specific temporal graph neural network module for node-level dynamic link prediction, featuring an encoder with node-wise span memory and a predictor. To overcome the few-shot problem, we design an adaptive meta-learner with span-wise adaptation and node-wise adaptation to extract two types of implicit information behind this problem. The acquired implicit information can serve as model initialization and facilitate rapid adaptation to new nodes through a fine-tuning process on just a few links. Experiments on three publicly available datasets demonstrate the superior performance of our model compared to existing state-of-the-art methods. Qinhu Zhang |
IJCNN | 3 |
| 2024 | scCorrector: a robust method for integrating multi-study single-cell dataabstractThe advent of single-cell sequencing technologies has revolutionized cell biology studies. However, integrative analyses of diverse single-cell data face serious challenges, including technological noise, sample heterogeneity, and different modalities and species. To address these problems, we propose scCorrector, a variational autoencoder-based model that can integrate single-cell data from different studies and map them into a common space. Specifically, we designed a Study Specific Adaptive Normalization for each study in decoder to implement these features. scCorrector substantially achieves competitive and robust performance compared with state-of-the-art methods and brings novel insights under various circumstances (e.g. various batches, multi-omics, cross-species, and development stages). In addition, the integration of single-cell data and spatial data makes it possible to transfer information between different studies, which greatly expand the narrow range of genes covered by MERFISH technology. In summary, scCorrector can efficiently integrate multi-study single-cell datasets, thereby providing broad opportunities to tackle challenges emerging from noisy resources. Zhen-Hao Guo, Siguo Wang, Qinhu Zhang, De-Shuang Huang |
Briefings Bioinform. | 4 |
| 2024 | scMGATGRN: a multiview graph attention network-based method for inferring gene regulatory networks from single-cell transcriptomic dataabstractThe gene regulatory network (GRN) plays a vital role in understanding the structure and dynamics of cellular systems, revealing complex regulatory relationships, and exploring disease mechanisms. Recently, deep learning (DL)-based methods have been proposed to infer GRNs from single-cell transcriptomic data and achieved impressive performance. However, these methods do not fully utilize graph topological information and high-order neighbor information from multiple receptive fields. To overcome those limitations, we propose a novel model based on multiview graph attention network, namely, scMGATGRN, to infer GRNs. scMGATGRN mainly consists of GAT, multiview, and view-level attention mechanism. GAT can extract essential features of the gene regulatory network. The multiview model can simultaneously utilize local feature information and high-order neighbor feature information of nodes in the gene regulatory network. The view-level attention mechanism dynamically adjusts the relative importance of node embedding representations and efficiently aggregates node embedding representations from two views. To verify the effectiveness of scMGATGRN, we compared its performance with 10 methods (five shallow learning algorithms and five state-of-the-art DL-based methods) on seven benchmark single-cell RNA sequencing (scRNA-seq) datasets from five cell lines (two in human and three in mouse) with four different kinds of ground-truth networks. The experimental results not only show that scMGATGRN outperforms competing methods but also demonstrate the potential of this model in inferring GRNs. The code and data of scMGATGRN are made freely available on GitHub (https://github.com/nathanyl/scMGATGRN). Lin Yuan 0001, Zhen Shen 0003, Qinhu Zhang, Chun-Hou Zheng 0001, De-Shuang Huang |
Briefings Bioinform. | 5 |
| 2024 | Identification of ferroptosis-related lncRNAs for predicting prognosis and immunotherapy response in non-small cell lung cancer
Lin Yuan 0001, Shengguo Sun, Qinhu Zhang, Hai-Tao Li, Zhen Shen 0003, Chunyu Hu 0001, Lan Ye, Chun-Hou Zheng 0001, De-Shuang Huang |
Future Gener. Comput. Syst. | 3 |
| 2024 | iCRBP-LKHA: Large convolutional kernel and hybrid channel-spatial attention for identifying circRNA-RBP interaction sitesabstractCircular RNAs (circRNAs) play vital roles in transcription and translation. Identification of circRNA-RBP (RNA-binding protein) interaction sites has become a fundamental step in molecular and cell biology. Deep learning (DL)-based methods have been proposed to predict circRNA-RBP interaction sites and achieved impressive identification performance. However, those methods cannot effectively capture long-distance dependencies, and cannot effectively utilize the interaction information of multiple features. To overcome those limitations, we propose a DL-based model iCRBP-LKHA using deep hybrid networks for identifying circRNA-RBP interaction sites. iCRBP-LKHA adopts five encoding schemes. Meanwhile, the neural network architecture, which consists of large kernel convolutional neural network (LKCNN), convolutional block attention module with one-dimensional convolution (CBAM-1D) and bidirectional gating recurrent unit (BiGRU), can explore local information, global context information and multiple features interaction information automatically. To verify the effectiveness of iCRBP-LKHA, we compared its performance with shallow learning algorithms on 37 circRNAs datasets and 37 circRNAs stringent datasets. And we compared its performance with state-of-the-art DL-based methods on 37 circRNAs datasets, 37 circRNAs stringent datasets and 31 linear RNAs datasets. The experimental results not only show that iCRBP-LKHA outperforms other competing methods, but also demonstrate the potential of this model in identifying other RNA-RBP interaction sites. Lin Yuan 0001, Jinling Lai, Qinhu Zhang, Zhen Shen 0003, Chun-Hou Zheng 0001, De-Shuang Huang |
PLoS Comput. Biol. | 5 |
| 2023 | Computational prediction and characterization of cell-type-specific and shared binding sitesabstractMOTIVATION: Cell-type-specific gene expression is maintained in large part by transcription factors (TFs) selectively binding to distinct sets of sites in different cell types. Recent research works have provided evidence that such cell-type-specific binding is determined by TF's intrinsic sequence preferences, cooperative interactions with co-factors, cell-type-specific chromatin landscapes and 3D chromatin interactions. However, computational prediction and characterization of cell-type-specific and shared binding sites is rarely studied. RESULTS: In this article, we propose two computational approaches for predicting and characterizing cell-type-specific and shared binding sites by integrating multiple types of features, in which one is based on XGBoost and another is based on convolutional neural network (CNN). To validate the performance of our proposed approaches, ChIP-seq datasets of 10 binding factors were collected from the GM12878 (lymphoblastoid) and K562 (erythroleukemic) human hematopoietic cell lines, each of which was further categorized into cell-type-specific (GM12878- and K562-specific) and shared binding sites. Then, multiple types of features for these binding sites were integrated to train the XGBoost- and CNN-based models. Experimental results show that our proposed approaches significantly outperform other competing methods on three classification tasks. Moreover, we identified independent feature contributions for cell-type-specific and shared sites through SHAP values and explored the ability of the CNN-based model to predict cell-type-specific and shared binding sites by excluding or including DNase signals. Furthermore, we investigated the generalization ability of our proposed approaches to different binding factors in the same cellular environment. AVAILABILITY AND IMPLEMENTATION: The source code is available at: https://github.com/turningpoint1988/CSSBS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qinhu Zhang, Pengrui Teng, Siguo Wang, Zhenghao Guo, Chang-an Yuan 0001, Qi Liu 0019, De-Shuang Huang |
Bioinform. | 1 |
| 2023 | scInterpreter: a knowledge-regularized generative model for interpretably integrating scRNA-seq dataabstractBACKGROUND: The rapid emergence of single-cell RNA-seq (scRNA-seq) data presents remarkable opportunities for broad investigations through integration analyses. However, most integration models are black boxes that lack interpretability or are hard to train. RESULTS: To address the above issues, we propose scInterpreter, a deep learning-based interpretable model. scInterpreter substantially outperforms other state-of-the-art (SOTA) models in multiple benchmark datasets. In addition, scInterpreter is extensible and can integrate and annotate atlas scRNA-seq data. We evaluated the robustness of scInterpreter in a variety of situations. Through comparison experiments, we found that with a knowledge prior, the training process can be significantly accelerated. Finally, we conducted interpretability analysis for each dimension (pathway) of cell representation in the embedding space. CONCLUSIONS: The results showed that the cell representations obtained by scInterpreter are full of biological significance. Through weight sorting, we found several new genes related to pathways in PBMC dataset. In general, scInterpreter is an effective and interpretable integration tool. It is expected that scInterpreter will bring great convenience to the study of single-cell transcriptomics. Zhen-Hao Guo, Siguo Wang, Qinhu Zhang, Jin-Ming Shi |
BMC Bioinform. | 4 |
| 2023 | In silico prediction methods of self-interacting proteins: an empirical and academic survey
Zhu-Hong You, Qinhu Zhang, Zhen-Hao Guo, Siguo Wang |
Frontiers Comput. Sci. | 3 |
| 2023 | iCircDA-NEAE: Accelerated attribute network embedding and dynamic convolutional autoencoder for circRNA-disease associations predictionabstractAccumulating evidence suggests that circRNAs play crucial roles in human diseases. CircRNA-disease association prediction is extremely helpful in understanding pathogenesis, diagnosis, and prevention, as well as identifying relevant biomarkers. During the past few years, a large number of deep learning (DL) based methods have been proposed for predicting circRNA-disease association and achieved impressive prediction performance. However, there are two main drawbacks to these methods. The first is these methods underutilize biometric information in the data. Second, the features extracted by these methods are not outstanding to represent association characteristics between circRNAs and diseases. In this study, we developed a novel deep learning model, named iCircDA-NEAE, to predict circRNA-disease associations. In particular, we use disease semantic similarity, Gaussian interaction profile kernel, circRNA expression profile similarity, and Jaccard similarity simultaneously for the first time, and extract hidden features based on accelerated attribute network embedding (AANE) and dynamic convolutional autoencoder (DCAE). Experimental results on the circR2Disease dataset show that iCircDA-NEAE outperforms other competing methods significantly. Besides, 16 of the top 20 circRNA-disease pairs with the highest prediction scores were validated by relevant literature. Furthermore, we observe that iCircDA-NEAE can effectively predict new potential circRNA-disease associations. Lin Yuan 0001, Jiawang Zhao 0002, Zhen Shen 0003, Qinhu Zhang, Chun-Hou Zheng 0001, De-Shuang Huang |
PLoS Comput. Biol. | 4 |
| 2023 | Predicting the Sequence Specificities of DNA-Binding Proteins by DNA Fine-Tuned Language Model With Decaying Learning RatesabstractDNA-binding proteins (DBPs) play vital roles in the regulation of biological systems. Although there are already many deep learning methods for predicting the sequence specificities of DBPs, they face two challenges as follows. Classic deep learning methods for DBPs prediction usually fail to capture the dependencies between genomic sequences since their commonly used one-hot codes are mutually orthogonal. Besides, these methods usually perform poorly when samples are inadequate. To address these two challenges, we developed a novel language model for mining DBPs using human genomic data and ChIP-seq datasets with decaying learning rates, named DNA Fine-tuned Language Model (DFLM). It can capture the dependencies between genome sequences based on the context of human genomic data and then fine-tune the features of DBPs tasks using different ChIP-seq datasets. First, we compared DFLM with the existing widely used methods on 69 datasets and we achieved excellent performance. Moreover, we conducted comparative experiments on complex DBPs and small datasets. The results show that DFLM still achieved a significant improvement. Finally, through visualization analysis of one-hot encoding and DFLM, we found that one-hot encoding completely cut off the dependencies of DNA sequences themselves, while DFLM using language models can well represent the dependency of DNA sequences. Source code are available at: https://github.com/Deep-Bioinfo/DFLM. Qinhu Zhang, Siguo Wang, Zhen-Hao Guo, De-Shuang Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | Using Fully Convolutional Network to Locate Transcription Factor Binding Sites Based on DNA Sequence and Conservation InformationabstractTranscription factors (TFs) play a part in gene expression. TFs can form complex gene expression regulation system by combining with DNA. Thereby, identifying the binding regions has become an indispensable step for understanding the regulatory mechanism of gene expression. Due to the great achievements of applying deep learning (DL) to computer vision and language processing in recent years, many scholars are inspired to use these methods to predict TF binding sites (TFBSs), achieving extraordinary results. However, these methods mainly focus on whether DNA sequences include TFBSs. In this paper, we propose a fully convolutional network (FCN) coupled with refinement residual block (RRB) and global average pooling layer (GAPL), namely FCNARRB. Our model could classify binding sequences at nucleotide level by outputting dense label for input data. Experimental results on human ChIP-seq datasets show that the RRB and GAPL structures are very useful for improving model performance. Adding GAPL improves the performance by 9.32% and 7.61% in terms of IoU (Intersection of Union) and PRAUC (Area Under Curve of Precision and Recall), and adding RRB improves the performance by 7.40% and 4.64%, respectively. In addition, we find that conservation information can help locate TFBSs. Qinhu Zhang, Youhong Xu, Siguo Wang, Yong Wu 0006, Yuan-Nong Ye, Chang-an Yuan 0001, Valeriya V. Gribova, Vladimir F. Filaretov, De-Shuang Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2023 | DeepTPpred: A Deep Learning Approach With Matrix Factorization for Predicting Therapeutic Peptides by Integrating Length InformationabstractThe abuse of traditional antibiotics has led to increased resistance of bacteria and viruses. Efficient therapeutic peptide prediction is critical for peptide drug discovery. However, most of the existing methods only make effective predictions for one class of therapeutic peptides. It is worth noting that currently no predictive method considers sequence length information as a distinct feature of therapeutic peptides. In this article, a novel deep learning approach with matrix factorization for predicting therapeutic peptides (DeepTPpred) by integrating length information are proposed. The matrix factorization layer can learn the potential features of the encoded sequence through the mechanism of first compression and then restoration. And the length features of the sequence of therapeutic peptides are embedded with encoded amino acid sequences. To automatically learn therapeutic peptide predictions, these latent features are input into the neural networks with self-attention mechanism. On eight therapeutic peptide datasets, DeepTPpred achieved excellent prediction results. Based on these datasets, we first integrated eight datasets to obtain a full therapeutic peptide integration dataset. Then, we obtained two functional integration datasets based on the functional similarity of the peptides. Finally, we also conduct experiments on the latest versions of the ACP and CPP datasets. Overall, the experimental results show that our work is effective for the identification of therapeutic peptides. Siguo Wang, Qinhu Zhang |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | DLoopCaller: A deep learning approach for predicting genome-wide chromatin loops by integrating accessible chromatin landscapesabstractIn recent years, major advances have been made in various chromosome conformation capture technologies to further satisfy the needs of researchers for high-quality, high-resolution contact interactions. Discriminating the loops from genome-wide contact interactions is crucial for dissecting three-dimensional(3D) genome structure and function. Here, we present a deep learning method to predict genome-wide chromatin loops, called DLoopCaller, by combining accessible chromatin landscapes and raw Hi-C contact maps. Some available orthogonal data ChIA-PET/HiChIP and Capture Hi-C were used to generate positive samples with a wider contact matrix which provides the possibility to find more potential genome-wide chromatin loops. The experimental results demonstrate that DLoopCaller effectively improves the accuracy of predicting genome-wide chromatin loops compared to the state-of-the-art method Peakachu. Moreover, compared to two of most popular loop callers, such as HiCCUPS and Fit-Hi-C, DLoopCaller identifies some unique interactions. We conclude that a combination of chromatin landscapes on the one-dimensional genome contributes to understanding the 3D genome organization, and the identified chromatin loops reveal cell-type specificity and transcription factor motif co-enrichment across different cell lines and species. Siguo Wang, Qinhu Zhang, Zhen-Hao Guo, Kyungsook Han, De-Shuang Huang |
PLoS Comput. Biol. | 2 |
| 2022 | Base-resolution prediction of transcription factor binding signals by a deep learning frameworkabstractTranscription factors (TFs) play an important role in regulating gene expression, thus the identification of the sites bound by them has become a fundamental step for molecular and cellular biology. In this paper, we developed a deep learning framework leveraging existing fully convolutional neural networks (FCN) to predict TF-DNA binding signals at the base-resolution level (named as FCNsignal). The proposed FCNsignal can simultaneously achieve the following tasks: (i) modeling the base-resolution signals of binding regions; (ii) discriminating binding or non-binding regions; (iii) locating TF-DNA binding regions; (iv) predicting binding motifs. Besides, FCNsignal can also be used to predict opening regions across the whole genome. The experimental results on 53 TF ChIP-seq datasets and 6 chromatin accessibility ATAC-seq datasets show that our proposed framework outperforms some existing state-of-the-art methods. In addition, we explored to use the trained FCNsignal to locate all potential TF-DNA binding regions on a whole chromosome and predict DNA sequences of arbitrary length, and the results show that our framework can find most of the known binding regions and accept sequences of arbitrary length. Furthermore, we demonstrated the potential ability of our framework in discovering causal disease-associated single-nucleotide polymorphisms (SNPs) through a series of experiments. Qinhu Zhang, Siguo Wang, Zhen-Hao Guo, Qi Liu 0019, De-Shuang Huang |
PLoS Comput. Biol. | 1 |
| 2022 | RMSCNN: A Random Multi-Scale Convolutional Neural Network for Marine Microbial Bacteriocins IdentificationabstractThe abuse of traditional antibiotics has led to an increase in the resistance of bacteria and viruses. Similar to the function of antibacterial peptides, bacteriocins are more common as a kind of peptides produced by bacteria that have bactericidal or bacterial effects. More importantly, the marine environment is one of the most abundant resources for extracting marine microbial bacteriocins (MMBs). Identifying bacteriocins from marine microorganisms is a common goal for the development of new drugs. Effective use of MMBs will greatly alleviate the current antibiotic abuse problem. In this work, deep learning is used to identify meaningful MMBs. We propose a random multi-scale convolutional neural network method. In the scale setting, we set a random model to update the scale value randomly. The scale selection method can reduce the contingency caused by artificial setting under certain conditions, thereby making the method more extensive. The results show that the classification performance of the proposed method is better than the state-of-the-art classification methods. In addition, some potential MMBs are predicted, and some different sequence analyses are performed on these candidates. It is worth mentioning that after sequence analysis, the HNH endonucleases of different marine bacteria are considered as potential bacteriocins. Qinhu Zhang, Valeriya V. Gribova, Vladimir F. Filaretov, De-Shuang Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | A Deep Learning Model for RNA-Protein Binding Preference Prediction Based on Hierarchical LSTM and Attention NetworkabstractAttention mechanism has the ability to find important information in the sequence. The regions of the RNA sequence that can bind to proteins are more important than those that cannot bind to proteins. Neither conventional methods nor deep learning-based methods, they are not good at learning this information. In this study, LSTM is used to extract the correlation features between different sites in RNA sequence. We also use attention mechanism to evaluate the importance of different sites in RNA sequence. We get the optimal combination of k-mer length, k-mer stride window, k-mer sentence length, k-mer sentence stride window, and optimization function through hyper-parm experiments. The results show that the performance of our method is better than other methods. We tested the effects of changes in k-mer vector length on model performance. We show model performance changes under various k-mer related parameter settings. Furthermore, we investigate the effect of attention mechanism and RNA structure data on model performance. Zhen Shen 0003, Qinhu Zhang, Kyungsook Han, De-Shuang Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | Predicting In-Vitro DNA-Protein Binding With a Spatially Aligned Fusion of Sequence and ShapeabstractDiscovery of transcription factor binding sites (TFBSs) is of primary importance for understanding the underlying binding mechanic and gene regulation process. Growing evidence indicates that apart from the primary DNA sequences, DNA shape landscape has a significant influence on transcription factor binding preference. To effectively model the co-influence of sequence and shape features, we emphasize the importance of position information of sequence motif and shape pattern. In this paper, we propose a novel deep learning-based architecture, named hybridShape eDeepCNN, for TFBS prediction which integrates DNA sequence and shape information in a spatially aligned manner. Our model utilizes the power of the multi-layer convolutional neural network and constructs an independent subnetwork to adapt for the distinct data distribution of heterogeneous features. Besides, we explore the usage of continuous embedding vectors as the representation of DNA sequences. Based on the experiments on 20 in-vitro datasets derived from universal protein binding microarrays (uPBMs), we demonstrate the superiority of our proposed method and validate the underlying design logic. Qinhu Zhang, Yindong Zhang, Siguo Wang, Valeriya V. Gribova, Vladimir F. Filaretov, De-Shuang Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2022 | FCNGRU: Locating Transcription Factor Binding Sites by Combing Fully Convolutional Neural Network With Gated Recurrent UnitabstractDeciphering the relationship between transcription factors (TFs) and DNA sequences is very helpful for computational inference of gene regulation and a comprehensive understanding of gene regulation mechanisms. Transcription factor binding sites (TFBSs) are specific DNA short sequences that play a pivotal role in controlling gene expression through interaction with TF proteins. Although recently many computational and deep learning methods have been proposed to predict TFBSs aiming to predict sequence specificity of TF-DNA binding, there is still a lack of effective methods to directly locate TFBSs. In order to address this problem, we propose FCNGRU combing a fully convolutional neural network (FCN) with the gated recurrent unit (GRU) to directly locate TFBSs in this paper. Furthermore, we present a two-task framework (FCNGRU-double): one is a classification task at nucleotide level which predicts the probability of each nucleotide and locates TFBSs, and the other is a regression task at sequence level which predicts the intensity of each sequence. A series of experiments are conducted on 45 in-vitro datasets collected from the UniPROBE database derived from universal protein binding microarrays (uPBMs). Compared with competing methods, FCNGRU-double achieves much better results on these datasets. Moreover, FCNGRU-double has an advantage over a single-task framework, FCNGRU-single, which only contains the branch of locating TFBSs. In addition, we combine with in vivo datasets to make a further analysis and discussion. Siguo Wang, Qinhu Zhang |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | A survey on deep learning in DNA/RNA motif miningabstractDNA/RNA motif mining is the foundation of gene function research. The DNA/RNA motif mining plays an extremely important role in identifying the DNA- or RNA-protein binding site, which helps to understand the mechanism of gene regulation and management. For the past few decades, researchers have been working on designing new efficient and accurate algorithms for mining motif. These algorithms can be roughly divided into two categories: the enumeration approach and the probabilistic method. In recent years, machine learning methods had made great progress, especially the algorithm represented by deep learning had achieved good performance. Existing deep learning methods in motif mining can be roughly divided into three types of models: convolutional neural network (CNN) based models, recurrent neural network (RNN) based models, and hybrid CNN-RNN based models. We introduce the application of deep learning in the field of motif mining in terms of data preprocessing, features of existing deep learning architectures and comparing the differences between the basic deep learning models. Through the analysis and comparison of existing deep learning methods, we found that the more complex models tend to perform better than simple ones when data are sufficient, and the current methods are relatively simple compared with other fields such as computer vision, language processing (NLP), computer games, etc. Therefore, it is necessary to conduct a summary in motif mining by deep learning, which can help researchers understand this field. Zhen Shen 0003, Qinhu Zhang, Siguo Wang, De-Shuang Huang |
Briefings Bioinform. | 3 |
| 2021 | Locating transcription factor binding sites by fully convolutional neural networkabstractTranscription factors (TFs) play an important role in regulating gene expression, thus identification of the regions bound by them has become a fundamental step for molecular and cellular biology. In recent years, an increasing number of deep learning (DL) based methods have been proposed for predicting TF binding sites (TFBSs) and achieved impressive prediction performance. However, these methods mainly focus on predicting the sequence specificity of TF-DNA binding, which is equivalent to a sequence-level binary classification task, and fail to identify motifs and TFBSs accurately. In this paper, we developed a fully convolutional network coupled with global average pooling (FCNA), which by contrast is equivalent to a nucleotide-level binary classification task, to roughly locate TFBSs and accurately identify motifs. Experimental results on human ChIP-seq datasets show that FCNA outperforms other competing methods significantly. Besides, we find that the regions located by FCNA can be used by motif discovery tools to further refine the prediction performance. Furthermore, we observe that FCNA can accurately identify TF-DNA binding motifs across different cell lines and infer indirect TF-DNA bindings. Qinhu Zhang, Siguo Wang, Qi Liu 0019, De-Shuang Huang |
Briefings Bioinform. | 1 |
| 2021 | Special issue: Advanced Intelligent Computing Theory and Applications in Big Data Era
Qinhu Zhang, Vitoantonio Bevilacqua, De-Shuang Huang |
Neurocomputing | 1 |
| 2021 | Predicting in-vitro Transcription Factor Binding Sites Using DNA Sequence + ShapeabstractDiscovery of transcription factor binding sites (TFBSs) is essential for understanding the underlying binding mechanisms and cellular functions. Recently, Convolutional neural network (CNN) has succeeded in predicting TFBSs from the primary DNA sequences. In addition to DNA sequences, several evidences suggest that protein-DNA binding is partly mediated by properties of DNA shape. Although many methods have been proposed to jointly account for DNA sequences and shape properties in predicting TFBSs, they ignore the power of the combination of deep learning and DNA sequence + shape. Therefore we develop a deep-learning-based sequence + shape framework (DLBSS) in this paper, which appropriately integrates DNA sequences and shape properties, to better understand protein-DNA binding preference. This method uses a shared CNN to find their common patterns from DNA sequences and their corresponding shape features, which are then concatenated to compute a predicted value. Using 66 in-vitro datasets derived from universal protein binding microarrays (uPBMs), we show that our proposed method DLBSS significantly improves the performance of predicting TFBSs. In addition, we explain the reason why we should use the shared CNN, and explore the performance of DLBSS when using a deeper CNN, through a series of experiments. Qinhu Zhang, Zhen Shen 0003, De-Shuang Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | Predicting TF-DNA Binding Motifs from ChIP-seq Datasets Using the Bag-Based Classifier Combined With a Multi-Fold Learning SchemeabstractThe rapid development of high-throughput sequencing technology provides unique opportunities for studying of transcription factor binding sites, but also brings new computational challenges. Recently, a series of discriminative motif discovery (DMD) methods have been proposed and offer promising solutions for addressing these challenges. However, because of the huge computation cost, most of them have to choose approximate schemes that either sacrifice the accuracy of motif representation or tune motif parameter indirectly. In this paper, we propose a bag-based classifier combined with a multi-fold learning scheme (BCMF) to discover motifs from ChIP-seq datasets. First, BCMF formulates input sequences as a labeled bag naturally. Then, a bag-based classifier, combining with a bag feature extracting strategy, is applied to construct the objective function, and a multi-fold learning scheme is used to solve it. Compared with the existing DMD tools, BCMF features three improvements: 1) Learning position weight matrix (PWM) directly in a continuous space; 2) Proposing to represent a positive bag with a feature fused by its k "most positive" patterns. 3) Applying a more advanced learning scheme. The experimental results on 134 ChIP-seq datasets show that BCMF substantially outperforms existing DMD methods (including DREME, HOMER, XXmotif, motifRG, EDCOD and our previous work). Qinhu Zhang, Dailun Wang, Kyungsook Han, De-Shuang Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | Multi-Scale Capsule Network for Predicting DNA-Protein Binding SitesabstractDiscovering DNA-protein binding sites, also known as motif discovery, is the foundation for further analysis of transcription factors (TFs). Deep learning algorithms such as convolutional neural networks (CNN) have been introduced to motif discovery task and have achieved state-of-art performance. However, due to the limitations of CNN, motif discovery methods based on CNN do not take full advantage of large-scale sequencing data generated by high-throughput sequencing technology. Hence, in this paper we propose multi-scale capsule network architecture (MSC) integrating multi-scale CNN, a variant of CNN able to extract motif features of different lengths, and capsule network, a novel type of artificial neural network architecture aimed at improving CNN. The proposed method is tested on real ChIP-seq datasets and the experimental results show a considerable improvement compared with two well-tested deep learning-based sequence model, DeepBind and Deepsea. Qinhu Zhang, Kyungsook Han, Asoke K. Nandi, De-Shuang Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | Three-Layer Dynamic Transfer Learning Language Model for E. Coli Promoter Classification
Qinhu Zhang, Siguo Wang, Chang-an Yuan 0001, Xiao Qin 0005, Hongjie Wu, Xingming Zhao |
ICIC (2) | 3 |
| 2020 | A New Method Combining DNA Shape Features to Improve the Prediction Accuracy of Transcription Factor Binding Sites
Siguo Wang, Qinhu Zhang, Chang-an Yuan 0001, Xiao Qin 0005, Hongjie Wu, Xingming Zhao |
ICIC (2) | 4 |
| 2020 | Predicting in-Vitro Transcription Factor Binding Sites with Deep Embedding Convolution Network
Yindong Zhang, Qinhu Zhang, Chang-an Yuan 0001, Xiao Qin 0005, Hongjie Wu, Xingming Zhao |
ICIC (2) | 2 |
| 2020 | Weakly-Supervised Convolutional Neural Network Architecture for Predicting Protein-DNA BindingabstractAlthough convolutional neural networks (CNN) have outperformed conventional methods in predicting the sequence specificities of protein-DNA binding in recent years, they do not take full advantage of the intrinsic weakly-supervised information of DNA sequences that a bound sequence may contain multiple TFBS(s). Here, we propose a weakly-supervised convolutional neural network architecture (WSCNN), combining multiple-instance learning (MIL) with CNN, to further boost the performance of predicting protein-DNA binding. WSCNN first divides each DNA sequence into multiple overlapping subsequences (instances) with a sliding window, and then separately models each instance using CNN, and finally fuses the predicted scores of all instances in the same bag using four fusion methods, including Max, Average, Linear Regression, and Top-Bottom Instances. The experimental results on in vivo and in vitro datasets illustrate the performance of the proposed approach. Moreover, models built on in vitro data using WSCNN can predict in vivo protein-DNA binding with good accuracy. In addition, we give a quantitative analysis of the importance of the reverse-complement mode in predicting in vivo protein-DNA binding, and explain why not directly use advanced pooling layers to combine MIL with CNN, through a series of experiments. Qinhu Zhang, Lin Zhu 0008, Wenzheng Bao, De-Shuang Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2019 | Motif Discovery via Convolutional Networks with K-mer Embedding
Dailun Wang, Qinhu Zhang, Chang-an Yuan 0001, Xiao Qin 0005, Zhi-Kai Huang |
ICIC (2) | 2 |
| 2019 | High-Order Convolutional Neural Network Architecture for Predicting DNA-Protein Binding SitesabstractAlthough Deep learning algorithms have outperformed conventional methods in predicting the sequence specificities of DNA-protein binding, they lack to consider the dependencies among nucleotides and the diverse binding lengths for different transcription factors (TFs). To address the above two limitations simultaneously, in this paper, we propose a high-order convolutional neural network architecture (HOCNN), which employs a high-order encoding method to build high-order dependencies among nucleotides, and a multi-scale convolutional layer to capture the motif features of different length. The experimental results on real ChIP-seq datasets show that the proposed method outperforms the state-of-the-art deep learning method (DeepBind) in the motif discovery task. In addition, we provide further insights about the importance of introducing additional convolutional kernels and the degeneration problem of importing high-order in the motif discovery task. Qinhu Zhang, Lin Zhu 0008, De-Shuang Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |