VLDB 2026 Research / reviewers in the wild / expert
Siguo Wang
dblp:260/2733
· DBLP profile ↗
18ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0002-3244-3629ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GraphLooper: predicting chromatin loops based on hierarchical multi-view graph pooling methodabstractChromatin loops serve as fundamental functional units of three-dimensional genome organization, playing pivotal roles in regulating gene expression and maintaining genomic spatial organization. Accurate identification of these fine-scale structures is crucial for advancing our understanding of cellular biological processes and the mechanisms underlying disease. However, due to the inherent complexity and dynamic of chromatin interactions, existing methods often fail to adequately characterize and capture multi-dimensional features. To address these limitations, we introduce GraphLooper, a novel framework using hierarchical multi-view graph pooling to enhance training and inference on large-scale data. GraphLooper transforms Hi-C data into a graph-structured representation, integrating multi-dimensional epigenomic features to construct a robust chromatin interaction model. Employing a hierarchical multi-view graph pooling mechanism, it effectively aggregates multi-scale features, enhancing representation learning. Evaluations across diverse cell lines demonstrate that GraphLooper outperforms state-of-the-art methods in prediction accuracy and generalization, particularly in capturing long-range chromatin interactions critical for precise spatial gene regulation. Siguo Wang, Zhipeng Li 0002, Hailin Feng, Zhen-Hao Guo, Zuquan Hu, Qinhu Zhang, De-Shuang Huang |
Briefings Bioinform. | 1 |
| 2025 | A brief survey of deep learning-based models for CircRNA-protein binding sites predictionabstractCircRNAs are a particular single-stranded, circular structure and “non-coding” RNA molecules, with various biological functions . Existing studies have demonstrated the fundamental role of circRNAs in gene expression regulation and their significant involvement in the development of diverse complex diseases. Predicting the protein binding sites in circRNA can aid in comprehending the regulation mechanism involved in circRNA-protein binding during gene expression and facilitate the investigation of potential diagnosis and treatment strategies for complex diseases. This review begins by introducing the concept and functions of circRNAs, as well as their involvement in gene expression regulation . Then, some critical and publicly accessible databases about circRNA annotation, protein annotation, circRNA-protein binding were listed. Next, we present a brief introduction to the computational model for predicting circRNA-protein binding, followed by model performance comparison and suggestions for non-computer science experts on model selection. Finally, we examine the problems, limitations, and advantages of computational models and explore the further direction of circRNA-protein prediction, such as developing new and complex computational models, introducing complex biological sequence encoding schemes, and integrating additional biological data related to circRNA-protein binding. Zhen Shen 0003, Lin Yuan 0001, Wenzheng Bao, Siguo Wang, Qinhu Zhang, De-Shuang Huang |
Neurocomputing | 4 |
| 2025 | NPENN: A Noise Perturbation Ensemble Neural Network for Microbiome Disease Phenotype PredictionabstractWith advances in microbiomics, the crucial role of microbes in disease progression is increasingly recognized. However, predicting disease phenotypes using microbiome data remains challenging due to data complexity, heterogeneity, and limited model generalization. Current methods often depend on specific datasets and are vulnerable to adversarial attacks. To address these issues, this paper introduces a novel Noise Perturbation Ensemble Neural Network model (NPENN), which combines noise mechanisms with Gradient Boosting (GB) techniques for robust neural network ensemble learning. NPENN, validated on multiple microbiome datasets, shows superior accuracy and generalization compared to traditional methods, effectively handling data complexity and variability. This approach enhances model robustness and feature learning by integrating GB prior knowledge. Additionally, the study explores microbial community roles in various diseases, providing insights into disease mechanisms and potential biomarkers for personalized precision diagnosis and treatment strategies. Yan Wu 0011, Qinhu Zhang, Siguo Wang, Zhen-Hao Guo |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | scCorrector: a robust method for integrating multi-study single-cell dataabstractThe advent of single-cell sequencing technologies has revolutionized cell biology studies. However, integrative analyses of diverse single-cell data face serious challenges, including technological noise, sample heterogeneity, and different modalities and species. To address these problems, we propose scCorrector, a variational autoencoder-based model that can integrate single-cell data from different studies and map them into a common space. Specifically, we designed a Study Specific Adaptive Normalization for each study in decoder to implement these features. scCorrector substantially achieves competitive and robust performance compared with state-of-the-art methods and brings novel insights under various circumstances (e.g. various batches, multi-omics, cross-species, and development stages). In addition, the integration of single-cell data and spatial data makes it possible to transfer information between different studies, which greatly expand the narrow range of genes covered by MERFISH technology. In summary, scCorrector can efficiently integrate multi-study single-cell datasets, thereby providing broad opportunities to tackle challenges emerging from noisy resources. Zhen-Hao Guo, Siguo Wang, Qinhu Zhang, De-Shuang Huang |
Briefings Bioinform. | 3 |
| 2023 | Computational prediction and characterization of cell-type-specific and shared binding sitesabstractMOTIVATION: Cell-type-specific gene expression is maintained in large part by transcription factors (TFs) selectively binding to distinct sets of sites in different cell types. Recent research works have provided evidence that such cell-type-specific binding is determined by TF's intrinsic sequence preferences, cooperative interactions with co-factors, cell-type-specific chromatin landscapes and 3D chromatin interactions. However, computational prediction and characterization of cell-type-specific and shared binding sites is rarely studied. RESULTS: In this article, we propose two computational approaches for predicting and characterizing cell-type-specific and shared binding sites by integrating multiple types of features, in which one is based on XGBoost and another is based on convolutional neural network (CNN). To validate the performance of our proposed approaches, ChIP-seq datasets of 10 binding factors were collected from the GM12878 (lymphoblastoid) and K562 (erythroleukemic) human hematopoietic cell lines, each of which was further categorized into cell-type-specific (GM12878- and K562-specific) and shared binding sites. Then, multiple types of features for these binding sites were integrated to train the XGBoost- and CNN-based models. Experimental results show that our proposed approaches significantly outperform other competing methods on three classification tasks. Moreover, we identified independent feature contributions for cell-type-specific and shared sites through SHAP values and explored the ability of the CNN-based model to predict cell-type-specific and shared binding sites by excluding or including DNase signals. Furthermore, we investigated the generalization ability of our proposed approaches to different binding factors in the same cellular environment. AVAILABILITY AND IMPLEMENTATION: The source code is available at: https://github.com/turningpoint1988/CSSBS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qinhu Zhang, Pengrui Teng, Siguo Wang, Zhenghao Guo, Chang-an Yuan 0001, Qi Liu 0019, De-Shuang Huang |
Bioinform. | 3 |
| 2023 | scInterpreter: a knowledge-regularized generative model for interpretably integrating scRNA-seq dataabstractBACKGROUND: The rapid emergence of single-cell RNA-seq (scRNA-seq) data presents remarkable opportunities for broad investigations through integration analyses. However, most integration models are black boxes that lack interpretability or are hard to train. RESULTS: To address the above issues, we propose scInterpreter, a deep learning-based interpretable model. scInterpreter substantially outperforms other state-of-the-art (SOTA) models in multiple benchmark datasets. In addition, scInterpreter is extensible and can integrate and annotate atlas scRNA-seq data. We evaluated the robustness of scInterpreter in a variety of situations. Through comparison experiments, we found that with a knowledge prior, the training process can be significantly accelerated. Finally, we conducted interpretability analysis for each dimension (pathway) of cell representation in the embedding space. CONCLUSIONS: The results showed that the cell representations obtained by scInterpreter are full of biological significance. Through weight sorting, we found several new genes related to pathways in PBMC dataset. In general, scInterpreter is an effective and interpretable integration tool. It is expected that scInterpreter will bring great convenience to the study of single-cell transcriptomics. Zhen-Hao Guo, Siguo Wang, Qinhu Zhang, Jin-Ming Shi |
BMC Bioinform. | 3 |
| 2023 | In silico prediction methods of self-interacting proteins: an empirical and academic survey
Zhu-Hong You, Qinhu Zhang, Zhen-Hao Guo, Siguo Wang |
Frontiers Comput. Sci. | 5 |
| 2023 | Predicting the Sequence Specificities of DNA-Binding Proteins by DNA Fine-Tuned Language Model With Decaying Learning RatesabstractDNA-binding proteins (DBPs) play vital roles in the regulation of biological systems. Although there are already many deep learning methods for predicting the sequence specificities of DBPs, they face two challenges as follows. Classic deep learning methods for DBPs prediction usually fail to capture the dependencies between genomic sequences since their commonly used one-hot codes are mutually orthogonal. Besides, these methods usually perform poorly when samples are inadequate. To address these two challenges, we developed a novel language model for mining DBPs using human genomic data and ChIP-seq datasets with decaying learning rates, named DNA Fine-tuned Language Model (DFLM). It can capture the dependencies between genome sequences based on the context of human genomic data and then fine-tune the features of DBPs tasks using different ChIP-seq datasets. First, we compared DFLM with the existing widely used methods on 69 datasets and we achieved excellent performance. Moreover, we conducted comparative experiments on complex DBPs and small datasets. The results show that DFLM still achieved a significant improvement. Finally, through visualization analysis of one-hot encoding and DFLM, we found that one-hot encoding completely cut off the dependencies of DNA sequences themselves, while DFLM using language models can well represent the dependency of DNA sequences. Source code are available at: https://github.com/Deep-Bioinfo/DFLM. Qinhu Zhang, Siguo Wang, Zhen-Hao Guo, De-Shuang Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | Using Fully Convolutional Network to Locate Transcription Factor Binding Sites Based on DNA Sequence and Conservation InformationabstractTranscription factors (TFs) play a part in gene expression. TFs can form complex gene expression regulation system by combining with DNA. Thereby, identifying the binding regions has become an indispensable step for understanding the regulatory mechanism of gene expression. Due to the great achievements of applying deep learning (DL) to computer vision and language processing in recent years, many scholars are inspired to use these methods to predict TF binding sites (TFBSs), achieving extraordinary results. However, these methods mainly focus on whether DNA sequences include TFBSs. In this paper, we propose a fully convolutional network (FCN) coupled with refinement residual block (RRB) and global average pooling layer (GAPL), namely FCNARRB. Our model could classify binding sequences at nucleotide level by outputting dense label for input data. Experimental results on human ChIP-seq datasets show that the RRB and GAPL structures are very useful for improving model performance. Adding GAPL improves the performance by 9.32% and 7.61% in terms of IoU (Intersection of Union) and PRAUC (Area Under Curve of Precision and Recall), and adding RRB improves the performance by 7.40% and 4.64%, respectively. In addition, we find that conservation information can help locate TFBSs. Qinhu Zhang, Youhong Xu, Siguo Wang, Yong Wu 0006, Yuan-Nong Ye, Chang-an Yuan 0001, Valeriya V. Gribova, Vladimir F. Filaretov, De-Shuang Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | DeepTPpred: A Deep Learning Approach With Matrix Factorization for Predicting Therapeutic Peptides by Integrating Length InformationabstractThe abuse of traditional antibiotics has led to increased resistance of bacteria and viruses. Efficient therapeutic peptide prediction is critical for peptide drug discovery. However, most of the existing methods only make effective predictions for one class of therapeutic peptides. It is worth noting that currently no predictive method considers sequence length information as a distinct feature of therapeutic peptides. In this article, a novel deep learning approach with matrix factorization for predicting therapeutic peptides (DeepTPpred) by integrating length information are proposed. The matrix factorization layer can learn the potential features of the encoded sequence through the mechanism of first compression and then restoration. And the length features of the sequence of therapeutic peptides are embedded with encoded amino acid sequences. To automatically learn therapeutic peptide predictions, these latent features are input into the neural networks with self-attention mechanism. On eight therapeutic peptide datasets, DeepTPpred achieved excellent prediction results. Based on these datasets, we first integrated eight datasets to obtain a full therapeutic peptide integration dataset. Then, we obtained two functional integration datasets based on the functional similarity of the peptides. Finally, we also conduct experiments on the latest versions of the ACP and CPP datasets. Overall, the experimental results show that our work is effective for the identification of therapeutic peptides. Siguo Wang, Qinhu Zhang |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | DLoopCaller: A deep learning approach for predicting genome-wide chromatin loops by integrating accessible chromatin landscapesabstractIn recent years, major advances have been made in various chromosome conformation capture technologies to further satisfy the needs of researchers for high-quality, high-resolution contact interactions. Discriminating the loops from genome-wide contact interactions is crucial for dissecting three-dimensional(3D) genome structure and function. Here, we present a deep learning method to predict genome-wide chromatin loops, called DLoopCaller, by combining accessible chromatin landscapes and raw Hi-C contact maps. Some available orthogonal data ChIA-PET/HiChIP and Capture Hi-C were used to generate positive samples with a wider contact matrix which provides the possibility to find more potential genome-wide chromatin loops. The experimental results demonstrate that DLoopCaller effectively improves the accuracy of predicting genome-wide chromatin loops compared to the state-of-the-art method Peakachu. Moreover, compared to two of most popular loop callers, such as HiCCUPS and Fit-Hi-C, DLoopCaller identifies some unique interactions. We conclude that a combination of chromatin landscapes on the one-dimensional genome contributes to understanding the 3D genome organization, and the identified chromatin loops reveal cell-type specificity and transcription factor motif co-enrichment across different cell lines and species. Siguo Wang, Qinhu Zhang, Zhen-Hao Guo, Kyungsook Han, De-Shuang Huang |
PLoS Comput. Biol. | 1 |
| 2022 | Base-resolution prediction of transcription factor binding signals by a deep learning frameworkabstractTranscription factors (TFs) play an important role in regulating gene expression, thus the identification of the sites bound by them has become a fundamental step for molecular and cellular biology. In this paper, we developed a deep learning framework leveraging existing fully convolutional neural networks (FCN) to predict TF-DNA binding signals at the base-resolution level (named as FCNsignal). The proposed FCNsignal can simultaneously achieve the following tasks: (i) modeling the base-resolution signals of binding regions; (ii) discriminating binding or non-binding regions; (iii) locating TF-DNA binding regions; (iv) predicting binding motifs. Besides, FCNsignal can also be used to predict opening regions across the whole genome. The experimental results on 53 TF ChIP-seq datasets and 6 chromatin accessibility ATAC-seq datasets show that our proposed framework outperforms some existing state-of-the-art methods. In addition, we explored to use the trained FCNsignal to locate all potential TF-DNA binding regions on a whole chromosome and predict DNA sequences of arbitrary length, and the results show that our framework can find most of the known binding regions and accept sequences of arbitrary length. Furthermore, we demonstrated the potential ability of our framework in discovering causal disease-associated single-nucleotide polymorphisms (SNPs) through a series of experiments. Qinhu Zhang, Siguo Wang, Zhen-Hao Guo, Qi Liu 0019, De-Shuang Huang |
PLoS Comput. Biol. | 3 |
| 2022 | Predicting In-Vitro DNA-Protein Binding With a Spatially Aligned Fusion of Sequence and ShapeabstractDiscovery of transcription factor binding sites (TFBSs) is of primary importance for understanding the underlying binding mechanic and gene regulation process. Growing evidence indicates that apart from the primary DNA sequences, DNA shape landscape has a significant influence on transcription factor binding preference. To effectively model the co-influence of sequence and shape features, we emphasize the importance of position information of sequence motif and shape pattern. In this paper, we propose a novel deep learning-based architecture, named hybridShape eDeepCNN, for TFBS prediction which integrates DNA sequence and shape information in a spatially aligned manner. Our model utilizes the power of the multi-layer convolutional neural network and constructs an independent subnetwork to adapt for the distinct data distribution of heterogeneous features. Besides, we explore the usage of continuous embedding vectors as the representation of DNA sequences. Based on the experiments on 20 in-vitro datasets derived from universal protein binding microarrays (uPBMs), we demonstrate the superiority of our proposed method and validate the underlying design logic. Qinhu Zhang, Yindong Zhang, Siguo Wang, Valeriya V. Gribova, Vladimir F. Filaretov, De-Shuang Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | FCNGRU: Locating Transcription Factor Binding Sites by Combing Fully Convolutional Neural Network With Gated Recurrent UnitabstractDeciphering the relationship between transcription factors (TFs) and DNA sequences is very helpful for computational inference of gene regulation and a comprehensive understanding of gene regulation mechanisms. Transcription factor binding sites (TFBSs) are specific DNA short sequences that play a pivotal role in controlling gene expression through interaction with TF proteins. Although recently many computational and deep learning methods have been proposed to predict TFBSs aiming to predict sequence specificity of TF-DNA binding, there is still a lack of effective methods to directly locate TFBSs. In order to address this problem, we propose FCNGRU combing a fully convolutional neural network (FCN) with the gated recurrent unit (GRU) to directly locate TFBSs in this paper. Furthermore, we present a two-task framework (FCNGRU-double): one is a classification task at nucleotide level which predicts the probability of each nucleotide and locates TFBSs, and the other is a regression task at sequence level which predicts the intensity of each sequence. A series of experiments are conducted on 45 in-vitro datasets collected from the UniPROBE database derived from universal protein binding microarrays (uPBMs). Compared with competing methods, FCNGRU-double achieves much better results on these datasets. Moreover, FCNGRU-double has an advantage over a single-task framework, FCNGRU-single, which only contains the branch of locating TFBSs. In addition, we combine with in vivo datasets to make a further analysis and discussion. Siguo Wang, Qinhu Zhang |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | A survey on deep learning in DNA/RNA motif miningabstractDNA/RNA motif mining is the foundation of gene function research. The DNA/RNA motif mining plays an extremely important role in identifying the DNA- or RNA-protein binding site, which helps to understand the mechanism of gene regulation and management. For the past few decades, researchers have been working on designing new efficient and accurate algorithms for mining motif. These algorithms can be roughly divided into two categories: the enumeration approach and the probabilistic method. In recent years, machine learning methods had made great progress, especially the algorithm represented by deep learning had achieved good performance. Existing deep learning methods in motif mining can be roughly divided into three types of models: convolutional neural network (CNN) based models, recurrent neural network (RNN) based models, and hybrid CNN-RNN based models. We introduce the application of deep learning in the field of motif mining in terms of data preprocessing, features of existing deep learning architectures and comparing the differences between the basic deep learning models. Through the analysis and comparison of existing deep learning methods, we found that the more complex models tend to perform better than simple ones when data are sufficient, and the current methods are relatively simple compared with other fields such as computer vision, language processing (NLP), computer games, etc. Therefore, it is necessary to conduct a summary in motif mining by deep learning, which can help researchers understand this field. Zhen Shen 0003, Qinhu Zhang, Siguo Wang, De-Shuang Huang |
Briefings Bioinform. | 4 |
| 2021 | Locating transcription factor binding sites by fully convolutional neural networkabstractTranscription factors (TFs) play an important role in regulating gene expression, thus identification of the regions bound by them has become a fundamental step for molecular and cellular biology. In recent years, an increasing number of deep learning (DL) based methods have been proposed for predicting TF binding sites (TFBSs) and achieved impressive prediction performance. However, these methods mainly focus on predicting the sequence specificity of TF-DNA binding, which is equivalent to a sequence-level binary classification task, and fail to identify motifs and TFBSs accurately. In this paper, we developed a fully convolutional network coupled with global average pooling (FCNA), which by contrast is equivalent to a nucleotide-level binary classification task, to roughly locate TFBSs and accurately identify motifs. Experimental results on human ChIP-seq datasets show that FCNA outperforms other competing methods significantly. Besides, we find that the regions located by FCNA can be used by motif discovery tools to further refine the prediction performance. Furthermore, we observe that FCNA can accurately identify TF-DNA binding motifs across different cell lines and infer indirect TF-DNA bindings. Qinhu Zhang, Siguo Wang, Qi Liu 0019, De-Shuang Huang |
Briefings Bioinform. | 2 |
| 2020 | Three-Layer Dynamic Transfer Learning Language Model for E. Coli Promoter Classification
Qinhu Zhang, Siguo Wang, Chang-an Yuan 0001, Xiao Qin 0005, Hongjie Wu, Xingming Zhao |
ICIC (2) | 4 |
| 2020 | A New Method Combining DNA Shape Features to Improve the Prediction Accuracy of Transcription Factor Binding Sites
Siguo Wang, Qinhu Zhang, Chang-an Yuan 0001, Xiao Qin 0005, Hongjie Wu, Xingming Zhao |
ICIC (2) | 1 |