Yongqing Zhang 0001

dblp:83/7153-1 · DBLP profile ↗
← Back
41ranked-venue papers
18as first author
36since 2021 · last 2026
0000-0003-3422-8305ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 22 · 10 first-author · 19 since 2021Artificial intelligence and machine learning · 18 · 7 first-author · 16 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Contrastive learning in both structure and function spaces improve drug-target interaction prediction
abstract
BACKGROUND : Identifying drug–target interactions (DTIs) is essential in drug discovery and repositioning. Recently, deep learning has become the mainstream methodology for DTI prediction. However, the scarcity of three-dimensional structural data has forced almost all methods to predict drug-target interactions with low-dimensional data, thereby constraining their overall performance. METHODS : In tackling this challenge, we introduce a novel approach, CLSF-DTI. CLSF-DTI incorporates high-dimensional structural and functional information into the drug and protein features through contrastive learning during the feature extraction stage. This ensures that the model no longer solely focuses on sequence information, leading to a more precise modeling outcome. RESULTS : Experiments on five benchmark datasets demonstrate that CLSF-DTI achieves the best overall performance among five state-of-the-art baselines. Through ablation studies, we further prove that the contrastive module enhances the predictive performance and generalization ability of CLSF-DTI. Moreover, CLSF-DTI successfully identified some ligands for the protein PKA-Cα in drug screening experiments. CONCLUSIONS : This study proposed a contrastive learning model CLSF-DTI that integrates structural and functional similarity. It outperforms existing methods in drug-target interaction prediction and has stronger generalization ability. However, the handling of unbalanced data and long-distance dependencies still needs to be improved in the future. The data and source code are available at https://github.com/ZhangLab312/CLSF_DTI .
Yongqing Zhang 0001, Shuwen Xiong, Zixuan Wang 0025, Quan Zou 0001
BMC Bioinform.1
2026 Explainable multiscale representation learning for anticancer peptide prediction
Yongqing Zhang 0001, Zhigan Zhou, Yugui Xu, Jin Wu 0002, Quan Zou 0001, Lei Xu 0047
Eng. Appl. Artif. Intell.1
2026 Deep generative framework for modeling single-cell drug perturbation response
Yongqing Zhang 0001, Chenpeng Wu, Zhigan Zhou, Zhengxiao Huang, Aochen Zhang, Zixuan Wang 0025
Neural Networks1
2026 Prediction of cancer drug response based on heterogeneous graph neural networks and multi-omics data
Shuwen Xiong, Yugui Xu, Yongqing Zhang 0001
Neural Networks4
2026 MCFusion-DDI: Multimodal cross-attention fusion of local-global features and latent drug associations for explainable DDI prediction
Yongqing Zhang 0001, Yugui Xu, Zhigan Zhou, Jin Wu 0002, Quan Zou 0001, Lei Xu 0047
Neural Networks1
2026 SAD: Sparse-Aware Diffusion Model for Single-Cell Gene Expression Completion
abstract
Single-cell RNA sequencing (scRNA-seq) is entering an era of foundation models that accept the complete gene atlas as input, yet most current datasets cover only 10-12 k genes and contain numerous technical zeros, severely limiting the generalization of these models in downstream tasks. To address this, we pioneer the gene-completion task for scRNA-seq and present SAD, a diffusion-based framework tailored to extremely sparse data, capable of completing genes and correcting sparsity bias under high missing rates. Unlike imputation or reconstruction methods that rely on the i.i.d. assumption, SAD's completion paradigm can generate gene entries originally absent from the expression profile, be aware of and rectify sparsity-distribution bias, and supply foundation models with consistent, reliable inputs of more than 30 k genes. Extensive benchmarks show that SAD significantly outperforms existing methods across multiple completion metrics, particularly in extreme scenarios with missing rates above 80%. This provides a data foundation for reusing missing scRNA-seq information and for precision-medicine applications.
Yixin Xiang, Zixuan Wang 0025, Mengwen Liu, Quan Zou 0001, Yongqing Zhang 0001
IEEE Trans. Comput. Biol. Bioinform.6
2025 Distribution and Knowledge Alignment-based Single-cell Multi-omics Cell Type Annotation
abstract
Single-cell type annotation is crucial for understanding the tumor microenvironment. With advances in single-cell sequencing, researchers can now analyze the functional characteristics of different cell types at the single-cell level. However, high-dimensional sparsity and technical noise in single-cell data pose challenges for existing annotation methods. Traditional approaches rely mainly on transcriptomic data, neglecting complementary multi-omics information, leading to information loss and reduced accuracy. To address these limitations, this study proposes a novel single-cell annotation method based on multimodal alignment. By designing two alignment strategies—inter-modal feature alignment and inter-model knowledge alignment—the method integrates scRNA-seq and scATAC-seq data. The inter-modal feature alignment strategy corrects distributional discrepancies between modalities, enabling seamless mapping in the latent space and more accurately capturing multi-omics characteristics. The inter-model knowledge alignment strategy transfers structured knowledge from a teacher model to a student model, balancing computational efficiency and annotation performance. This approach reduces reliance on single modalities in traditional methods and enhances adaptability to diverse experimental conditions. Experiments on public datasets show that the proposed method outperforms state-of-the-art single-omics models across multiple benchmarks, demonstrating strong annotation capabilities and superior generalization to rare cell types.
Yongqing Zhang 0001, Zhengxiao Huang, Zixuan Wang 0025
IJCNN2
2025 GGANet: A Geometry-Enhanced Gated Attention Network for Drug-Target Interaction Prediction
abstract
Drug-target interaction (DTI) prediction is a critical step in drug discovery. Despite significant advances in deep learning-based methods, data representation and feature alignment challenges remain. Specifically, previous methods often rely on low-dimensional representations and limited labeled data, overlooking the importance of high-dimensional spatial geometric information and unlabeled data, which restricts the extraction of crucial features. Additionally, most approaches align features between amino acid residues and drug atoms using dot-product similarity, ignoring their biochemical differences, which limits the effectiveness of the alignment. We propose GGANet, a geometry-enhanced gated attention network, to address these limitations. GGANet integrates low-dimensional pre-trained embeddings with high-dimensional geometric information to generate robust and generalizable feature representations while employing gated attention to ensure efficient feature alignment. We conducted experiments on four datasets, and the results show that GGANet outperforms baseline methods, especially in predicting unseen data, demonstrating superior robustness and generalization. The implementation details of GGANet are available at https://github.com/ZhangLab312/GGANet.
Yongqing Zhang 0001
IJCNN2
2025 Single Cell Mosaic Integration via Regulatory Graph Guided Variational Cross-modal Fusion
abstract
Multiomics integration of single-cell datasets generated from multiple omics technologies is crucial for defining cellular heterogeneity. Mosaic data contains any combination of different modalities and batches of data, which brings challenges such as incomplete reference samples, non-overlapping features, and batch effects, making it difficult to extract information from them. Most existing methods require a reference batch that contains all modalities, which limits the versatility of mosaic data. In this study, we proposed a single-cell mosaic data integration method that guides variational cross-modal fusion via regulation graphs (GVMOS). Based on mosaic data, GVMOS first generates a regulatory graph of cross-modal feature interactions. Then, a coupled-trained dual variational autoencoder module is used to extract the complementary and specific information in different modal data and the cross-modal feature topology information in the regulatory graph. By integrating modality information with feature information, GVMOS can guide modality integration from the perspective of spatial semantics. Experiments on three bimodal paired datasets and five bimodal mosaic combinations show that GVMOS significantly improves the integration performance of existing methods, with an average improvement of 3% in overall scores.
Jiaheng Lv, Zixuan Wang 0025, Zhigan Zhou, Yongqing Zhang 0001
IJCNN4
2025 Enhancing Drug Synergy Prediction via flexible Fusion of Multimodal Heterogeneous Data
abstract
Discovering effective combinations of anticancer drugs is crucial for improving cancer treatment strategies. The accumulation of drug information and cell line data contributes to the development of effective deep learning prediction models for drug synergy. However, selecting high-quality data sources and designing appropriate methods remains a challenge. This article proposes an attention-based multimodal heterogeneous information fusion network MHFSyn for predicting drug synergy. MHFSyn uses multiple feature extractors to extract different modality features, and captures cross modal interactions and structural information through an attention fusion network, thereby achieving effective fusion of multimodal data and improving the overall performance and generalization ability of the model. Five fold cross-validation and leave-one-out cross-validation experiments show that MHFSyn had the best overall performance compared to the six comparison methods.Our code is available at https://github.com/ZhangLab312/MHFSyn.
Yugui Xu, Zhigan Zhou, Shuwen Xiong, Zixuan Wang 0025, Yongqing Zhang 0001
IJCNN8
2025 Cancer Drug Response Prediction Based on Graph Neural Networks and Cross-Attention
abstract
Accurately predicting individual cancer drug response(CDR) remains a significant challenge. The accumulation of multi-omics data from cell lines and drug information has greatly facilitated the development of predictive models for CDR. However, effectively integrating multi-source data and improving the prediction accuracy of models remain critical challenges to be addressed. In this study, based on the Genomics of Drug Sensitivity in Cancer (GDSC) database, designed a CDR prediction model named AGCCK(Autoencoder-Graph Neural Networks-Convolutional Neural Networks-Cross-Attention-Kolmogorov-Arnold Network), which utilizes Graph Neural Networks (GNN) and Cross-Attention mechanisms. For multisource data, the model incorporates different feature extractors, employing GNN or Convolutional Neural Networks (CNN) to extract features. After extraction, the cell line and drug data are fused using a cross-attention module, followed by another cross-attention module to further integrate the cell line and drug features. Finally, a Kolmogorov-Arnold Network (KAN) module is used to predict the final outcome. Research indicates that this is the first attempt to combine cross-attention modules and KAN in the field of CDR prediction. A series of model evaluation experiments were conducted on the GDSC database. The results demonstrate that the model’s predictive performance surpasses existing state-of-the-art algorithms, with improvements of 1.8%, 2%, and 2.8% in the Pearson Correlation Coefficient (PCC), Spearman Correlation Coefficient (SCC), and Coefficient of Determination (R2), respectively, compared to the best baseline model. The model also exhibits strong predictive capabilities in missing value prediction tasks. This study holds significant importance for CDR prediction and new drug development. AGCCK is freely available at https://github.com/ZhangLab312/AGCCK.
Yongqing Zhang 0001
IJCNN2
2025 An overview of computational methods in single-cell transcriptomic cell type annotation
abstract
The rapid accumulation of single-cell RNA sequencing data has provided unprecedented computational resources for cell type annotation, significantly advancing our understanding of cellular heterogeneity. Leveraging gene expression profiles derived from transcriptomic data, researchers can accurately infer cell types, sparking the development of numerous innovative annotation methods. These methods utilize a range of strategies, including marker genes, correlation-based matching, and supervised learning, to classify cell types. In this review, we systematically examine these annotation approaches based on transcriptomics-specific gene expression profiles and provide a comprehensive comparison and categorization of these methods. Furthermore, we focus on the main challenges in the annotation process, especially the long-tail distribution problem arising from data imbalance in rare cell types. We discuss the potential of deep learning techniques to address these issues and enhance model capability in recognizing novel cell types within an open-world framework.
Zixuan Wang 0025, Quan Zou 0001, Yongqing Zhang 0001
Briefings Bioinform.6
2025 MMGCSyn: Explainable synergistic drug combination prediction based on multimodal fusion
Yongqing Zhang 0001, Shuwen Xiong, Zhigan Zhou, Yugui Xu, Meiqin Gong
Future Gener. Comput. Syst.1
2025 A Multi-Omics Data Integration Framework for Gene Regulatory Network Inference Based on Contrastive Learning
abstract
The Gene Regulatory Networks (GRNs) ensure the stability of cellular states, preserving specific phenotypes and functions throughout the differentiation process. However, current tools still need improvement to effectively integrate multi-omics data and infer GRNs for particular cell types. We introduce CLMOGRI, a multi-omics TF-gene regulatory network inference framework based on heterogeneous networks and contrastive learning, designed to integrate multi-omics data for GRN inference. Through random walk techniques, CLMOGRI embeds multi-omics data into a unified feature space and extracts similar features between nodes. It then measures node similarity and predicts node relationships by contrastive learning. Finally, it includes a regulatory network interpreter to identify critical nodes and modules in GRNs, offering an analytical method for understanding complex interactions within biological systems. CLMOGRI surpasses existing baseline methods in terms of Area Under the Precision-Recall Curve (AUPR) and F-Score metrics, indicating its efficacy in capturing multi-omics information for GRN inference. It also reveals vital nodes and modules within the gene regulatory network, improving the interpretability of CLMOGRI and the utility of GRNs.
Yongqing Zhang 0001, Zhigan Zhou, Maocheng Wang, Zixuan Wang 0025, Quan Zou 0001
IEEE Trans. Comput. Biol. Bioinform.1
2025 A Comprehensive Adaptive Interpretable Takagi-Sugeno-Kang Fuzzy Classifier for Fatigue Driving Detection
abstract
Electroencephalogram (EEG) signals, as a reliable biological indicator, have been widely used in fatigue driving detection due to their capacity to reflect a driver's cognitive and neural response state. However, EEG signals have problems such as imbalanced data distribution, significant differences between subjects, and complex scenes, which affect the detection effect. Small commonalities between input objects can be interpreted as important information about an entire sample. Therefore, to retain as much information as possible, We design a new approach for integrating fuzzy features, comprehensive adaptive interpretable TSK fuzzy classifier(CAI-TSK-FC). It not only captures the features of multiple subclassifiers more efficiently and alleviates the dataset imbalance problem. Also, it can reduce the accumulation of error information by randomly retaining fuzzy rules as well as normalization. Finally, we linearly combine the results of multiple subclassifiers to comprehensively consider the learning effect of multiple subclassifiers to adapt to different subjects and datasets. Experiments conducted on both self-made and public datasets (SEED-VIG) show that CAI-TSK-FC has good performance and interpretability on different EEG fatigue driving datasets. In comparison to existing methods, it achieves an accuracy improvement of 3.15% and 1.52%, respectively, as well as a specificity improvement of 4.72% and 0.91%, respectively.
Dongrui Gao, Shihong Liu, Yingxian Gao, Pengrui Li, Haokai Zhang, Manqing Wang, Yan Shen 0001, Lutao Wang, Yongqing Zhang 0001
IEEE Trans. Fuzzy Syst.9
2024 A Local-Ascending-Global Learning Strategy for Brain-Computer Interface
abstract
Neuroscience research indicates that the interaction among different functional regions of the brain plays a crucial role in driving various cognitive tasks. Existing studies have primarily focused on constructing either local or global functional connectivity maps within the brain, often lacking an adaptive approach to fuse functional brain regions and explore latent relationships between localization during different cognitive tasks. This paper introduces a novel approach called the Local-Ascending-Global Learning Strategy (LAG) to uncover higher-level latent topological patterns among functional brain regions. The strategy initiates from the local connectivity of individual brain functional regions and develops a K-Level Self-Adaptive Ascending Network (SALK) to dynamically capture strong connectivity patterns among brain regions during different cognitive tasks. Through the step-by-step fusion of brain regions, this approach captures higher-level latent patterns, shedding light on the progressively adaptive fusion of various brain functional regions under different cognitive tasks. Notably, this study represents the first exploration of higher-level latent patterns through progressively adaptive fusion of diverse brain functional regions under different cognitive tasks. The proposed LAG strategy is validated using datasets related to fatigue (SEED-VIG), emotion (SEED-IV), and motor imagery (BCI_C_IV_2a). The results demonstrate the generalizability of LAG, achieving satisfactory outcomes in independent-subject experiments across all three datasets. This suggests that LAG effectively characterizes higher-level latent patterns associated with different cognitive tasks, presenting a novel approach to understanding brain patterns in varying cognitive contexts.
Dongrui Gao, Haokai Zhang, Pengrui Li, Shihong Liu, Zhihong Zhou, Shaofei Ying, Yongqing Zhang 0001
AAAI9
2024 Cell-Specific Highly Correlated Network for Self-Supervised Distillation in Cell Type Annotation
abstract
Self-supervised learning has succeeded significantly in cell type annotation based on transcriptomic data. However, existing methods encode transcriptomic data as gene feature representations, limiting self-supervised learning to the gene level. This paper proposes a novel self-supervised distillation learning framework, scTCHCN, which designs a cell-specific highly correlated network. Firstly, this framework integrates the cell-specific highly correlated network with single-cell transcriptomic data to extract cell and gene features. ScTCHCN promotes interactions between the network and the transcriptomic data features during pre-training by constructing the cell-specific highly correlated network. Additionally, a self-supervised paradigm with distillation learning is proposed to enhance the global feature correlation between the two. Ultimately, the pre-trained student model aggregates deep feature constraints and complementary information at the single-cell and gene levels. It combines them with the prediction layer to construct the final cell type prediction model. Experimental results demonstrate that scTCHCN excels in cell type annotation and rare cell identification tasks, showcasing its potential for other applications. The source code and additional information for this study are available at https://github.com/ZhangLab312/scTCHCN.
Yugui Xu, Zixuan Wang 0025, Zhigan Zhou, Yongqing Zhang 0001, Quan Zou 0001
BIBM7
2024 An EEG-based cross-subject interpretable CNN for game player expertise level classification
Liqi Lin, Pengrui Li, Binnan Bai, Ruifang Cui, Zhenxia Yu, Dongrui Gao, Yongqing Zhang 0001
Expert Syst. Appl.8
2024 CSF-GTNet: A Novel Multi-Dimensional Feature Fusion Network Based on Convnext-GeLU- BiLSTM for EEG-Signals-Enabled Fatigue Driving Detection
abstract
Electroencephalography (EEG) signal has been recognized as an effective fatigue detection method, which can intuitively reflect the drivers' mental state. However, the research on multi-dimensional features in existing work could be much better. The instability and complexity of EEG signals will increase the difficulty of extracting data features. More importantly, most current work only treats deep learning models as classifiers. They ignored the features of different subjects learned by the model. Aiming at the above problems, this paper proposes a novel multi-dimensional feature fusion network, CSF-GTNet, based on time and space-frequency domains for fatigue detection. Specifically, it comprises Gaussian Time Domain Network (GTNet) and Pure Convolutional Spatial Frequency Domain Network (CSFNet). The experimental results show that the proposed method effectively distinguishes between alert and fatigue states. The accuracy rates are 85.16% and 81.48% on the self-made and SEED-VIG datasets, respectively, which are higher than the state-of-the-art methods. Moreover, we analyze the contribution of each brain region for fatigue detection through the brain topology map. In addition, we explore the changing trend of each frequency band and the significance between different subjects in the alert state and fatigue state through the heat map. Our research can provide new ideas in brain fatigue research and play a specific role in promoting the development of this field.
Dongrui Gao, Pengrui Li, Manqing Wang, Yujie Liang, Shihong Liu, Jiliu Zhou, Lutao Wang, Yongqing Zhang 0001
IEEE J. Biomed. Health Informatics8
2024 SFT-Net: A Network for Detecting Fatigue From EEG Signals by Combining 4D Feature Flow and Attention Mechanism
abstract
Fatigued driving is a leading cause of traffic accidents, and accurately predicting driver fatigue can significantly reduce their occurrence. However, modern fatigue detection models based on neural networks often face challenges such as poor interpretability and insufficient input feature dimensions. This article proposes a novel Spatial-Frequency-Temporal Network (SFT-Net) method for detecting driver fatigue using electroencephalogram (EEG) data. Our approach integrates EEG signals' spatial, frequency, and temporal information to improve recognition performance. We transform the differential entropy of five frequency bands of EEG signals into a 4D feature tensor to preserve these three types of information. An attention module is then used to recalibrate the spatial and frequency information of each input 4D feature tensor time slice. The output of this module is fed into a depthwise separable convolution (DSC) module, which extracts spatial and frequency features after attention fusion. Finally, long short-term memory (LSTM) is used to extract the temporal dependence of the sequence, and the final features are output through a linear layer. We validate the effectiveness of our model on the SEED-VIG dataset, and experimental results demonstrate that SFT-Net outperforms other popular models for EEG fatigue detection. Interpretability analysis supports the claim that our model has a certain level of interpretability. Our work addresses the challenge of detecting driver fatigue from EEG data and highlights the importance of integrating spatial, frequency, and temporal information.
Dongrui Gao, Kejie Wang, Manqing Wang, Jiliu Zhou, Yongqing Zhang 0001
IEEE J. Biomed. Health Informatics5
2023 Exploring Parameter-Efficient Fine-Tuning of a Large-Scale Pre-Trained Model for scRNA-seq Cell Type Annotation
abstract
Accurate identification of cell types is a pivotal and intricate task in scRNA-seq data analysis. Recently, significant strides have been made in cell type annotation of scRNA-seq data using pre-trained language models (PLMs). This method has surmounted the constraints of conventional approaches regarding precision, robustness, and generalization. However, the fine-tuning process of large-scale pre-trained models incurs substantial computational expenses. To tackle this issue, a promising avenue of research has emerged, proposing parameter-efficient fine-tuning techniques for PLMs. These techniques concentrate on fine-tuning only a small portion of the model parameters while attaining comparable performance. In this study, we extensively research parameter-efficient fine-tuning methods for scRNA-seq cell type annotation, employing scBERT as the backbone. We scrutinize the performance and compatibility of various parameter-efficient fine-tuning methodologies across multiple datasets. Through comprehensive analysis, we demonstrate the remarkable performance of parameter-efficient fine-tuning methods in cell type annotation. Hopefully, this study can inspire new thinking in analyzing scRNA-seq data.
Zixuan Wang 0025, Guiquan Zhu, Yongqing Zhang 0001, Quan Zou 0001
BIBM5
2023 KDProg: A Knowledge distillation graph neural network for cancer prognosis prediction and analysis
abstract
Accurately predicting cancer prognosis remains challenging, owing to the combination of computational and practical challenges. This study proposes KDProg, a knowledge distillation-based graph learning framework for predicting cancer prognosis and exploring downstream tasks. The framework includes a novel feature distillation paradigm that compresses a multi-layer complex teacher model to a single-layer simple student by using the teacher model’s middle-layer feature representations and outputs as supervision information to improve the student model’s performance. In addition, instead of introducing a unified temperature hyperparameter, KDProg adopts a novel strategy to parameterize the distillation temperature and combine it with the Cox partial log-likelihood function. So the model can learn the appropriate temperature. Furthermore, considering multi-omics data of patients are often complex to obtain in practical cancer prognosis, this paper uses different input data for the teacher and student models, respectively. The input data for the teacher model are multi-omics data (mRNA, CNV, and DNA methylation), clinical data, and KEGG pathways. The input data for the student model are mRNA, clinical data, and KEGG pathways. Extensive experiments on 15 real-world datasets from TCGA demonstrated the effectiveness and efficiency of the proposed method in predicting cancer prognosis. The results suggest that the proposed model can guide clinical decision-making.
Shuwen Xiong, Zixuan Wang 0025, Yongqing Zhang 0001, Quan Zou 0001
BIBM5
2023 HGTDG: An Interpretable Heterogeneous Graph Transformer Framework for Cancer Driver Gene Prediction
abstract
Accurately predicting cancer driver genes remains challenging due to the increasing size and complexity of cancer genomic data. In this study, HGTDG is proposed, a heterogeneous graph transformer framework for predicting cancer driver genes and exploring downstream tasks. The framework includes a heterogeneous graph construction module that constructs a gene-protein heterogeneous network based on KEGG pathways and the protein-protein interactions from the STRING database. In addition, the framework introduces a novel heterogeneous graph transformer module that uses multi-head attention mechanisms for gene node embedding. The transformer module can capture dedicated representations for genes and edges. Finally, the generated gene embeddings are fed into the classification module to classify genes into driver and non-driver genes. The experiment results show that HGTDG outperforms the state-of-the-art methods regarding the area under the receiver operating characteristic curves (AUROC) and the area under the precision-recall curves (AUPRC).
Shuwen Xiong, Zixuan Wang 0025, Guiquan Zhu, Yongqing Zhang 0001, Quan Zou 0001
BIBM5
2023 HEAP: a task adaptive-based explainable deep learning framework for enhancer activity prediction
abstract
Enhancers are crucial cis-regulatory elements that control gene expression in a cell-type-specific manner. Despite extensive genetic and computational studies, accurately predicting enhancer activity in different cell types remains a challenge, and the grammar of enhancers is still poorly understood. Here, we present HEAP (high-resolution enhancer activity prediction), an explainable deep learning framework for predicting enhancers and exploring enhancer grammar. The framework includes three modules that use grammar-based reasoning for enhancer prediction. The algorithm can incorporate DNA sequences and epigenetic modifications to obtain better accuracy. We use a novel two-step multi-task learning method, task adaptive parameter sharing (TAPS), to efficiently predict enhancers in different cell types. We first train a shared model with all cell-type datasets. Then we adapt to specific tasks by adding several task-specific subset layers. Experiments demonstrate that HEAP outperforms published methods and showcases the effectiveness of the TAPS, especially for those with limited training samples. Notably, the explainable framework HEAP utilizes post-hoc interpretation to provide insights into the prediction mechanisms from three perspectives: data, model architecture and algorithm, leading to a better understanding of model decisions and enhancer grammar. To the best of our knowledge, HEAP will be a valuable tool for insight into the complex mechanisms of enhancer activity.
Zixuan Wang 0025, Guiquan Zhu, Yongqing Zhang 0001
Briefings Bioinform.5
2023 Multiple sequence alignment based on deep reinforcement learning with self-attention and positional encoding
abstract
MOTIVATION: Multiple sequence alignment (MSA) is one of the hotspots of current research and is commonly used in sequence analysis scenarios. However, there is no lasting solution for MSA because it is a Nondeterministic Polynomially complete problem, and the existing methods still have room to improve the accuracy. RESULTS: We propose Deep reinforcement learning with Positional encoding and self-Attention for MSA, based on deep reinforcement learning, to enhance the accuracy of the alignment Specifically, inspired by the translation technique in natural language processing, we introduce self-attention and positional encoding to improve accuracy and reliability. Firstly, positional encoding encodes the position of the sequence to prevent the loss of nucleotide position information. Secondly, the self-attention model is used to extract the key features of the sequence. Then input the features into a multi-layer perceptron, which can calculate the insertion position of the gap according to the features. In addition, a novel reinforcement learning environment is designed to convert the classic progressive alignment into progressive column alignment, gradually generating each column's sub-alignment. Finally, merge the sub-alignment into the complete alignment. Extensive experiments based on several datasets validate our method's effectiveness for MSA, outperforming some state-of-the-art methods in terms of the Sum-of-pairs and Column scores. AVAILABILITY AND IMPLEMENTATION: The process is implemented in Python and available as open-source software from https://github.com/ZhangLab312/DPAMSA.
Zixuan Wang 0025, Shuwen Xiong, Naifeng Wen, Yongqing Zhang 0001
Bioinform.7
2023 HAMPLE: deciphering TF-DNA binding mechanism in different cellular environments by characterizing higher-order nucleotide dependency
abstract
MOTIVATION: Transcription factor (TF) binds to conservative DNA binding sites in different cellular environments and development stages by physical interaction with interdependent nucleotides. However, systematic computational characterization of the relationship between higher-order nucleotide dependency and TF-DNA binding mechanism in diverse cell types remains challenging. RESULTS: Here, we propose a novel multi-task learning framework HAMPLE to simultaneously predict TF binding sites (TFBS) in distinct cell types by characterizing higher-order nucleotide dependencies. Specifically, HAMPLE first represents a DNA sequence through three higher-order nucleotide dependencies, including k-mer encoding, DNA shape and histone modification. Then, HAMPLE uses the customized gate control and the channel attention convolutional architecture to further capture cell-type-specific and cell-type-shared DNA binding motifs and epigenomic languages. Finally, HAMPLE exploits the joint loss function to optimize the TFBS prediction for different cell types in an end-to-end manner. Extensive experimental results on seven datasets demonstrate that HAMPLE significantly outperforms the state-of-the-art approaches in terms of auROC. In addition, feature importance analysis illustrates that k-mer encoding, DNA shape, and histone modification have predictive power for TF-DNA binding in different cellular environments and are complementary to each other. Furthermore, ablation study, and interpretable analysis validate the effectiveness of the customized gate control and the channel attention convolutional architecture in characterizing higher-order nucleotide dependencies. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/ZhangLab312/Hample.
Zixuan Wang 0025, Shuwen Xiong, Jiliu Zhou, Yongqing Zhang 0001
Bioinform.5
2023 Emotion recognition based on convolutional gated recurrent units with attention
abstract
Studying brain activity and deciphering the information in electroencephalogram (EEG) signals has become an emerging research field, and substantial advances have been made in the EEG-based classification of emotions. However, using different EEG features and complementarity to discriminate other emotions is still challenging. Most existing models extract a single temporal feature from the EEG signal while ignoring the crucial temporal dynamic information, which, to a certain extent, constrains the classification capability of the model. To address this issue, we propose an Attention-Based Depthwise Parameterized Convolutional Gated Recurrent Unit (AB-DPCGRU) model and validate it with the mixed experiment on the SEED and SEED-IV datasets. The experimental outcomes revealed that the accuracy of the model outperforms the existing state-of-the-art methods, which confirmed the superiority of our approach over currently popular emotion recognition models.
Zhu Ye, Yuan Jing, Pengrui Li, Mingjing Yan, Yongqing Zhang 0001, Dongrui Gao
Connect. Sci.7
2023 SHNN: A single-channel EEG sleep staging model based on semi-supervised learning
Yongqing Zhang 0001, Wenpeng Cao, Lixiao Feng, Manqing Wang, Tianyu Geng, Jiliu Zhou, Dongrui Gao
Expert Syst. Appl.1
2022 Single-cell TF-DNA binding prediction and analysis based on transfer learning framework
abstract
Cell type-specific gene expressions during development or in disease are regulated by interactions between transcription factors (TFs) and their binding sites. Recently, many deep learning approaches have been developed to characterize TF-DNA binding within a population of cells. However, determining TF binding sites (TFBSs) in single cells remains challenging due to the sparsity of data. Here, we propose a multi-stage transfer learning framework called STAPLE for single-cell TF-DNA binding prediction and analysis. Specifically, we design the Cell Type Learning to capture the relationship between different TF-DNA binding events in the same cell type. Meanwhile, we present Individual Learning to extract common motif and chromatin accessibility features of a particular binding event in a cellular population. In addition, we leverage Single-cell Learning to annotate TFBSs in each cell without any supervised label. Extensive experiments based on 570 single-cell datasets validate the effectiveness of our framework for considering cellular heterogeneity, outperforming current methods. This work can provide new insight into the relationship between TF-DNA binding and cellular heterogeneity. The source code of STAPLE can be found at https://github.com/ZhangLab312/STAPLE.
Zixuan Wang 0025, Yongqing Zhang 0001, Maocheng Wang, Quan Zou 0001
BIBM2
2022 Predicting cell type-specific effects of variants on TF-DNA binding by meta-learning
abstract
Interpreting the regulatory code of gene expression and further understanding the functionality of noncoding variants on transcriptional effect is a crucial challenge. However, this remains difficult due to the complex association between SNPs and chromatin state. Here, we develop a meta-learning-based framework, U-TransNet, that can accurately predict the TF-DNA binding based on multiple chromatin features. Motivated by ab initio, our proposed framework contain two steps. (i) Meta-learning strategy is applied to predict chromatin profiles from DNA sequence. (ii) DNA sequence and all of these predicted chromatin features are used to predict TF-DNA binding affinity. Experiments demonstrate that U-TransNet has excellent performance, achieving significant improvements over existing methods in predicting base-resolution TF-DNA binding signals, TF binding sites, and motifs. We also demonstrate that integrating the more extended TFBS flank regions is a potential path to better understanding gene transcription. In addition, U-TransNet is applied to infer the effects of variants on TF-DNA binding affinity via in silico mutagenesis, and further to identify cell type-specific functional variants via comparing different cells. To the best of the authors’ knowledge, U-TransNet provides an efficient end-to-end computational framework for deciphering cis-regulator evolution.
Yongqing Zhang 0001, Zixuan Wang 0025, Maocheng Wang, Shuwen Xiong, Quan Zou 0001
BIBM1
2022 A novel convolution attention model for predicting transcription factor binding sites by combination of sequence and shape
abstract
The discovery of putative transcription factor binding sites (TFBSs) is important for understanding the underlying binding mechanism and cellular functions. Recently, many computational methods have been proposed to jointly account for DNA sequence and shape properties in TFBSs prediction. However, these methods fail to fully utilize the latent features derived from both sequence and shape profiles and have limitation in interpretability and knowledge discovery. To this end, we present a novel Deep Convolution Attention network combining Sequence and Shape, dubbed as D-SSCA, for precisely predicting putative TFBSs. Experiments conducted on 165 ENCODE ChIP-seq datasets reveal that D-SSCA significantly outperforms several state-of-the-art methods in predicting TFBSs, and justify the utility of channel attention module for feature refinements. Besides, the thorough analysis about the contribution of five shapes to TFBSs prediction demonstrates that shape features can improve the predictive power for transcription factors-DNA binding. Furthermore, D-SSCA can realize the cross-cell line prediction of TFBSs, indicating the occupancy of common interplay patterns concerning both sequence and shape across various cell lines. The source code of D-SSCA can be found at https://github.com/MoonLord0525/.
Yongqing Zhang 0001, Zixuan Wang 0025, Yuanqi Zeng, Shuwen Xiong, Maocheng Wang, Jiliu Zhou, Quan Zou 0001
Briefings Bioinform.1
2022 A survey on the algorithm and development of multiple sequence alignment
abstract
Multiple sequence alignment (MSA) is an essential cornerstone in bioinformatics, which can reveal the potential information in biological sequences, such as function, evolution and structure. MSA is widely used in many bioinformatics scenarios, such as phylogenetic analysis, protein analysis and genomic analysis. However, MSA faces new challenges with the gradual increase in sequence scale and the increasing demand for alignment accuracy. Therefore, developing an efficient and accurate strategy for MSA has become one of the research hotspots in bioinformatics. In this work, we mainly summarize the algorithms for MSA and its applications in bioinformatics. To provide a structured and clear perspective, we systematically introduce MSA's knowledge, including background, database, metric and benchmark. Besides, we list the most common applications of MSA in the field of bioinformatics, including database searching, phylogenetic analysis, genomic analysis, metagenomic analysis and protein analysis. Furthermore, we categorize and analyze classical and state-of-the-art algorithms, divided into progressive alignment, iterative algorithm, heuristics, machine learning and divide-and-conquer. Moreover, we also discuss the challenges and opportunities of MSA in bioinformatics. Our work provides a comprehensive survey of MSA applications and their relevant algorithms. It could bring valuable insights for researchers to contribute their knowledge to MSA and relevant studies.
Yongqing Zhang 0001, Jiliu Zhou, Quan Zou 0001
Briefings Bioinform.1
2021 By hybrid neural networks for prediction and interpretation of transcription factor binding sites based on multi-omics
abstract
Transcription factors (TFs) binding sites prediction and analysis are vital for comprehending cis-regulatory mechanisms. Recently, several deep learning-based methods have shown outstanding performance on TFs binding sites (TFBSs) recognition by leveraging solely base-pair arrangement of regulatory sequences. Except for the aforementioned genomic features, the epigenomics signature represented by the histone modification is also a critical factor related to TFs-DNA binding. We present a multi-omics based hybrid neural network, dubbed as BHSite, for TFBSs prediction by adaptively integrating base-pair arrangements and histone modification signatures. Experiments over 196 ChIP-seq datasets demonstrate that BHSite significantly outperforms several state-of-the-art methods in TFBSs prediction. Besides, studies of the relative importance of histone modification signatures prove that diverse signatures complement each other. Furthermore, visualization analysis of Squeeze-and-Excitation Network reveals the contribution of multi-omics latent features concerning different cell types to TFBS prediction. Thus, BHSite improves both performance and interpretability by combining the multi-omic features into deep learning architecture.
Yongqing Zhang 0001, Zixuan Wang 0025, Libo Lu, Xiaoyao Tan, Quan Zou 0001
BIBM1
2021 High-resolution transcription factor binding sites prediction improved performance and interpretability by deep learning method
abstract
Transcription factors (TFs) are essential proteins in regulating the spatiotemporal expression of genes. It is crucial to infer the potential transcription factor binding sites (TFBSs) with high resolution to promote biology and realize precision medicine. Recently, deep learning-based models have shown exemplary performance in the prediction of TFBSs at the base-pair level. However, the previous models fail to integrate nucleotide position information and semantic information without noisy responses. Thus, there is still room for improvement. Moreover, both the inner mechanism and prediction results of these models are challenging to interpret. To this end, the Deep Attentive Encoder-Decoder Neural Network (D-AEDNet) is developed to identify the location of TFs-DNA binding sites in DNA sequences. In particular, our model adopts Skip Architecture to leverage the nucleotide position information in the encoder and removes noisy responses in the information fusion process by Attention Gate. Simultaneously, the Transcription Factor Motif Discovery based on Sliding Window (TF-MoDSW), an approach to discover TFs-DNA binding motifs by utilizing the output of neural networks, is proposed to understand the biological meaning of the predicted result. On ChIP-exo datasets, experimental results show that D-AEDNet has better performance than competing methods. Besides, we authenticate that Attention Gate can improve the interpretability of our model by ways of visualization analysis. Furthermore, we confirm that ability of D-AEDNet to learn TFs-DNA binding motifs outperform the state-of-the-art methods and availability of TF-MoDSW to discover biological sequence motifs in TFs-DNA interaction by conducting experiment on ChIP-seq datasets.
Yongqing Zhang 0001, Zixuan Wang 0025, Yuanqi Zeng, Jiliu Zhou, Quan Zou 0001
Briefings Bioinform.1
2021 MFFNet: Multi-dimensional Feature Fusion Network based on attention mechanism for sEMG analysis to detect muscle fatigue
Yongqing Zhang 0001, Wenpeng Cao, Dongrui Gao, Manqing Wang, Jiliu Zhou, Ting Wang 0046
Expert Syst. Appl.1
2021 CAE-CNN: Predicting transcription factor binding site with convolutional autoencoder and convolutional neural network
Yongqing Zhang 0001, Shaojie Qiao, Yuanqi Zeng, Dongrui Gao, Nan Han, Jiliu Zhou
Expert Syst. Appl.1
2020 GRRFNet: Guided Regularized Random Forest-based Gene Regulatory Network Inference Using Data Integration
abstract
Gene regulatory network (GRN) inference based on gene expression data is still a huge challenge in systems biology. Genomic data, including time-series expression data, steady-state data, knockout data, and other biological data, such as Gene Ontology (GO) annotations, provide information on potential gene regulation. However, most existing methods continue to use only a single dataset for GRN inference. To integrate these types of data and improve the accuracy of inference, we propose a new data-integration strategy based on guided regular random forest (GRRF) for GRN inference, dubbed GRRFNet. Specifically, first, time-series data and steady-state data as main datasets are integrated to generate learning samples; simultaneously, other datasets are processed to design penalty coefficients for guiding feature selection; then, a GRRF model is applied to integrate the prior information with a main dataset to learn the transcription function and evaluate the importance of feature; finally, the score of the feature's importance is used as the possibility of the gene regulatory relationships to construct the GRN. To evaluate the performance of GRRFNet, we compare it with GENIE3, dynGENIE3, and GRIEF at the artificial DREAM4 dataset and the real Escherichia coli dataset. Although GRRFNet does not yield the best performance on every network, its competitiveness is still reflected herein.
Yongqing Zhang 0001, Qingyuan Chen, Dongrui Gao, Quan Zou 0001
BIBM1
2020 A point-of-interest suggestion algorithm in Multi-source geo-social networks
Shaojie Qiao, Nan Han, Guan Yuan, Yongqing Zhang 0001
Eng. Appl. Artif. Intell.6
2019 How to balance the bioinformatics data: pseudo-negative sampling
abstract
BACKGROUND: Imbalanced datasets are commonly encountered in bioinformatics classification problems, that is, the number of negative samples is much larger than that of positive samples. Particularly, the data imbalance phenomena will make us underestimate the performance of the minority class of positive samples. Therefore, how to balance the bioinformatic data becomes a very challenging and difficult problem. RESULTS: In this study, we propose a new data sampling approach, called pseudo-negative sampling, which can be effectively applied to handle the case that: negative samples greatly dominate positive samples. Specifically, we design a supervised learning method based on a max-relevance min-redundancy criterion beyond Pearson correlation coefficient (MMPCC), which is used to choose pseudo-negative samples from the negative samples and view them as positive samples. In addition, MMPCC uses an incremental searching technique to select optimal pseudo-negative samples to reduce the computation cost. Consequently, the discovered pseudo-negative samples have strong relevance to positive samples and less redundancy to negative ones. CONCLUSIONS: To validate the performance of our method, we conduct experiments base on four UCI datasets and three real bioinformatics datasets. According to the experimental results, we clearly observe the performance of MMPCC is better than other sampling methods in terms of Sensitivity, Specificity, Accuracy and the Mathew's Correlation Coefficient. This reveals that the pseudo-negative samples are particularly helpful to solve the imbalance dataset problem. Moreover, the gain of Sensitivity from the minority samples with pseudo-negative samples grows with the improvement of prediction accuracy on all dataset.
Yongqing Zhang 0001, Shaojie Qiao, Rongzhao Lu, Nan Han, Dingxiang Liu, Jiliu Zhou
BMC Bioinform.1
2019 Identification of DNA-protein binding sites by bootstrap multiple convolutional neural networks on sequence information
Yongqing Zhang 0001, Shaojie Qiao, Shengjie Ji, Nan Han, Dingxiang Liu, Jiliu Zhou
Eng. Appl. Artif. Intell.1
2018 ENSEMBLE-CNN: Predicting DNA Binding Sites in Protein Sequences by an Ensemble Deep Learning Method
Yongqing Zhang 0001, Shaojie Qiao, Shengjie Ji, Jiliu Zhou
ICIC (2)1