EDBT 2026 Demo / reviewers in the wild / expert
Shuwen Xiong
dblp:314/3192
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-0935-9787ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Contrastive learning in both structure and function spaces improve drug-target interaction predictionabstractBACKGROUND : Identifying drug–target interactions (DTIs) is essential in drug discovery and repositioning. Recently, deep learning has become the mainstream methodology for DTI prediction. However, the scarcity of three-dimensional structural data has forced almost all methods to predict drug-target interactions with low-dimensional data, thereby constraining their overall performance. METHODS : In tackling this challenge, we introduce a novel approach, CLSF-DTI. CLSF-DTI incorporates high-dimensional structural and functional information into the drug and protein features through contrastive learning during the feature extraction stage. This ensures that the model no longer solely focuses on sequence information, leading to a more precise modeling outcome. RESULTS : Experiments on five benchmark datasets demonstrate that CLSF-DTI achieves the best overall performance among five state-of-the-art baselines. Through ablation studies, we further prove that the contrastive module enhances the predictive performance and generalization ability of CLSF-DTI. Moreover, CLSF-DTI successfully identified some ligands for the protein PKA-Cα in drug screening experiments. CONCLUSIONS : This study proposed a contrastive learning model CLSF-DTI that integrates structural and functional similarity. It outperforms existing methods in drug-target interaction prediction and has stronger generalization ability. However, the handling of unbalanced data and long-distance dependencies still needs to be improved in the future. The data and source code are available at https://github.com/ZhangLab312/CLSF_DTI . Yongqing Zhang 0001, Shuwen Xiong, Zixuan Wang 0025, Quan Zou 0001 |
BMC Bioinform. | 5 |
| 2026 | Prediction of cancer drug response based on heterogeneous graph neural networks and multi-omics data
Shuwen Xiong, Yugui Xu, Yongqing Zhang 0001 |
Neural Networks | 2 |
| 2026 | A Multi-Modal Contrastive Learning Framework for Cyclic Peptide Permeability PredictionabstractCyclic peptides represent a rapidly growing class of therapeutics, yet their development is often hindered by the challenge of predicting cell membrane permeability, a critical determinant of drug efficacy. Existing computational methods often struggle to integrate the diverse structural information inherent in these complex molecules, resulting in suboptimal predictive accuracy. Here, we introduce MCPerm, a multi-modal deep learning framework that synergistically integrates 1D SMILES, 2D topological, and 3D geometric information through a novel modality share and contrastive learning strategy to accurately predict cyclic peptide permeability. MCPerm fine-tunes a pretrained peptide language model for SMILES encoding and uses a parameter-sharing graph transformer for structural representation, while a dual contrastive learning mechanism enforces representational consistency both within and between modalities. On the benchmark PAMPA dataset, MCPerm achieves state-of-the-art performance, significantly outperforming leading methods. We further demonstrate its robustness and competitive transferability across three independent assays (Caco-2, MDCK, and RRCK). Our work presents a robust in silico framework that holds potential to accelerate the rational design and discovery of cell-permeable cyclic peptide drugs. Furthermore, to move beyond predictive accuracy, we introduced an attention-based visualization analysis. The results demonstrate that our model is not a "black box"; it has learned key chemical principles governing cyclic peptide permeability. Shuwen Xiong, Feifei Cui, Rao Zeng, Ran Su, Leyi Wei |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2025 | Enhancing Drug Synergy Prediction via flexible Fusion of Multimodal Heterogeneous DataabstractDiscovering effective combinations of anticancer drugs is crucial for improving cancer treatment strategies. The accumulation of drug information and cell line data contributes to the development of effective deep learning prediction models for drug synergy. However, selecting high-quality data sources and designing appropriate methods remains a challenge. This article proposes an attention-based multimodal heterogeneous information fusion network MHFSyn for predicting drug synergy. MHFSyn uses multiple feature extractors to extract different modality features, and captures cross modal interactions and structural information through an attention fusion network, thereby achieving effective fusion of multimodal data and improving the overall performance and generalization ability of the model. Five fold cross-validation and leave-one-out cross-validation experiments show that MHFSyn had the best overall performance compared to the six comparison methods.Our code is available at https://github.com/ZhangLab312/MHFSyn. Yugui Xu, Zhigan Zhou, Shuwen Xiong, Zixuan Wang 0025, Yongqing Zhang 0001 |
IJCNN | 6 |
| 2025 | MMGCSyn: Explainable synergistic drug combination prediction based on multimodal fusion
Yongqing Zhang 0001, Shuwen Xiong, Zhigan Zhou, Yugui Xu, Meiqin Gong |
Future Gener. Comput. Syst. | 4 |
| 2023 | KDProg: A Knowledge distillation graph neural network for cancer prognosis prediction and analysisabstractAccurately predicting cancer prognosis remains challenging, owing to the combination of computational and practical challenges. This study proposes KDProg, a knowledge distillation-based graph learning framework for predicting cancer prognosis and exploring downstream tasks. The framework includes a novel feature distillation paradigm that compresses a multi-layer complex teacher model to a single-layer simple student by using the teacher model’s middle-layer feature representations and outputs as supervision information to improve the student model’s performance. In addition, instead of introducing a unified temperature hyperparameter, KDProg adopts a novel strategy to parameterize the distillation temperature and combine it with the Cox partial log-likelihood function. So the model can learn the appropriate temperature. Furthermore, considering multi-omics data of patients are often complex to obtain in practical cancer prognosis, this paper uses different input data for the teacher and student models, respectively. The input data for the teacher model are multi-omics data (mRNA, CNV, and DNA methylation), clinical data, and KEGG pathways. The input data for the student model are mRNA, clinical data, and KEGG pathways. Extensive experiments on 15 real-world datasets from TCGA demonstrated the effectiveness and efficiency of the proposed method in predicting cancer prognosis. The results suggest that the proposed model can guide clinical decision-making. Shuwen Xiong, Zixuan Wang 0025, Yongqing Zhang 0001, Quan Zou 0001 |
BIBM | 1 |
| 2023 | HGTDG: An Interpretable Heterogeneous Graph Transformer Framework for Cancer Driver Gene PredictionabstractAccurately predicting cancer driver genes remains challenging due to the increasing size and complexity of cancer genomic data. In this study, HGTDG is proposed, a heterogeneous graph transformer framework for predicting cancer driver genes and exploring downstream tasks. The framework includes a heterogeneous graph construction module that constructs a gene-protein heterogeneous network based on KEGG pathways and the protein-protein interactions from the STRING database. In addition, the framework introduces a novel heterogeneous graph transformer module that uses multi-head attention mechanisms for gene node embedding. The transformer module can capture dedicated representations for genes and edges. Finally, the generated gene embeddings are fed into the classification module to classify genes into driver and non-driver genes. The experiment results show that HGTDG outperforms the state-of-the-art methods regarding the area under the receiver operating characteristic curves (AUROC) and the area under the precision-recall curves (AUPRC). Shuwen Xiong, Zixuan Wang 0025, Guiquan Zhu, Yongqing Zhang 0001, Quan Zou 0001 |
BIBM | 1 |
| 2023 | Multiple sequence alignment based on deep reinforcement learning with self-attention and positional encodingabstractMOTIVATION: Multiple sequence alignment (MSA) is one of the hotspots of current research and is commonly used in sequence analysis scenarios. However, there is no lasting solution for MSA because it is a Nondeterministic Polynomially complete problem, and the existing methods still have room to improve the accuracy. RESULTS: We propose Deep reinforcement learning with Positional encoding and self-Attention for MSA, based on deep reinforcement learning, to enhance the accuracy of the alignment Specifically, inspired by the translation technique in natural language processing, we introduce self-attention and positional encoding to improve accuracy and reliability. Firstly, positional encoding encodes the position of the sequence to prevent the loss of nucleotide position information. Secondly, the self-attention model is used to extract the key features of the sequence. Then input the features into a multi-layer perceptron, which can calculate the insertion position of the gap according to the features. In addition, a novel reinforcement learning environment is designed to convert the classic progressive alignment into progressive column alignment, gradually generating each column's sub-alignment. Finally, merge the sub-alignment into the complete alignment. Extensive experiments based on several datasets validate our method's effectiveness for MSA, outperforming some state-of-the-art methods in terms of the Sum-of-pairs and Column scores. AVAILABILITY AND IMPLEMENTATION: The process is implemented in Python and available as open-source software from https://github.com/ZhangLab312/DPAMSA. Zixuan Wang 0025, Shuwen Xiong, Naifeng Wen, Yongqing Zhang 0001 |
Bioinform. | 5 |
| 2023 | HAMPLE: deciphering TF-DNA binding mechanism in different cellular environments by characterizing higher-order nucleotide dependencyabstractMOTIVATION: Transcription factor (TF) binds to conservative DNA binding sites in different cellular environments and development stages by physical interaction with interdependent nucleotides. However, systematic computational characterization of the relationship between higher-order nucleotide dependency and TF-DNA binding mechanism in diverse cell types remains challenging. RESULTS: Here, we propose a novel multi-task learning framework HAMPLE to simultaneously predict TF binding sites (TFBS) in distinct cell types by characterizing higher-order nucleotide dependencies. Specifically, HAMPLE first represents a DNA sequence through three higher-order nucleotide dependencies, including k-mer encoding, DNA shape and histone modification. Then, HAMPLE uses the customized gate control and the channel attention convolutional architecture to further capture cell-type-specific and cell-type-shared DNA binding motifs and epigenomic languages. Finally, HAMPLE exploits the joint loss function to optimize the TFBS prediction for different cell types in an end-to-end manner. Extensive experimental results on seven datasets demonstrate that HAMPLE significantly outperforms the state-of-the-art approaches in terms of auROC. In addition, feature importance analysis illustrates that k-mer encoding, DNA shape, and histone modification have predictive power for TF-DNA binding in different cellular environments and are complementary to each other. Furthermore, ablation study, and interpretable analysis validate the effectiveness of the customized gate control and the channel attention convolutional architecture in characterizing higher-order nucleotide dependencies. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/ZhangLab312/Hample. Zixuan Wang 0025, Shuwen Xiong, Jiliu Zhou, Yongqing Zhang 0001 |
Bioinform. | 2 |
| 2022 | Predicting cell type-specific effects of variants on TF-DNA binding by meta-learningabstractInterpreting the regulatory code of gene expression and further understanding the functionality of noncoding variants on transcriptional effect is a crucial challenge. However, this remains difficult due to the complex association between SNPs and chromatin state. Here, we develop a meta-learning-based framework, U-TransNet, that can accurately predict the TF-DNA binding based on multiple chromatin features. Motivated by ab initio, our proposed framework contain two steps. (i) Meta-learning strategy is applied to predict chromatin profiles from DNA sequence. (ii) DNA sequence and all of these predicted chromatin features are used to predict TF-DNA binding affinity. Experiments demonstrate that U-TransNet has excellent performance, achieving significant improvements over existing methods in predicting base-resolution TF-DNA binding signals, TF binding sites, and motifs. We also demonstrate that integrating the more extended TFBS flank regions is a potential path to better understanding gene transcription. In addition, U-TransNet is applied to infer the effects of variants on TF-DNA binding affinity via in silico mutagenesis, and further to identify cell type-specific functional variants via comparing different cells. To the best of the authors’ knowledge, U-TransNet provides an efficient end-to-end computational framework for deciphering cis-regulator evolution. Yongqing Zhang 0001, Zixuan Wang 0025, Maocheng Wang, Shuwen Xiong, Quan Zou 0001 |
BIBM | 5 |
| 2022 | A novel convolution attention model for predicting transcription factor binding sites by combination of sequence and shapeabstractThe discovery of putative transcription factor binding sites (TFBSs) is important for understanding the underlying binding mechanism and cellular functions. Recently, many computational methods have been proposed to jointly account for DNA sequence and shape properties in TFBSs prediction. However, these methods fail to fully utilize the latent features derived from both sequence and shape profiles and have limitation in interpretability and knowledge discovery. To this end, we present a novel Deep Convolution Attention network combining Sequence and Shape, dubbed as D-SSCA, for precisely predicting putative TFBSs. Experiments conducted on 165 ENCODE ChIP-seq datasets reveal that D-SSCA significantly outperforms several state-of-the-art methods in predicting TFBSs, and justify the utility of channel attention module for feature refinements. Besides, the thorough analysis about the contribution of five shapes to TFBSs prediction demonstrates that shape features can improve the predictive power for transcription factors-DNA binding. Furthermore, D-SSCA can realize the cross-cell line prediction of TFBSs, indicating the occupancy of common interplay patterns concerning both sequence and shape across various cell lines. The source code of D-SSCA can be found at https://github.com/MoonLord0525/. Yongqing Zhang 0001, Zixuan Wang 0025, Yuanqi Zeng, Shuwen Xiong, Maocheng Wang, Jiliu Zhou, Quan Zou 0001 |
Briefings Bioinform. | 5 |