Shengli Zhang 0002

dblp:41/4882-2 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0001-8786-0940ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 7 since 2021
YearPublicationVenuePosition
2025 An innovative peptide toxicity prediction model based on multi-scale convolutional neural network and residual connection
abstract
MOTIVATION: Peptide toxicity is a critical concern in the development of peptide-based therapeutics, as toxic peptides can lead to severe side effects, including organ damage, immune reactions, and cytotoxicity. Predicting peptide toxicity accurately is essential to ensure the safety and efficacy of these drugs. RESULTS: In this study, we propose a novel model, ToxMSRC, to predict peptide toxicity using a combination of the continuous bag of words (CBOW) method from word2vec, synthetic minority over-sampling technique (SMOTE), multi-scale convolutional neural networks (CNN), and bidirectional long short-term memory (BiLSTM). This approach addresses the challenge of data imbalance by augmenting positive samples and improves feature extraction through multi-scale convolution. Furthermore, the model incorporates a residual connection that helps prevent overfitting and enhances generalization ability, improving classification performance. The model is evaluated on benchmark and independent test sets, achieving BACC scores of 92.17% on independent test1 and 86.89% on independent test2, outperforming existing state-of-the-art models. Additionally, ToxMSRC provides valuable insights into the relationship between peptide toxicity and amino acid sequences, demonstrating its potential and practical value in peptide-based drug development. AVAILABILITY AND IMPLEMENTATION: The complete datasets, source code, and pre-trained models are made available at https://github.com/Renjingyi123/ToxMSRC and https://doi.org/10.5281/zenodo.15668530.
Shengli Zhang 0002, Jingyi Ren, Yunyun Liang
Bioinform.1
2024 AACFlow: an end-to-end model based on attention augmented convolutional neural network and flow-attention mechanism for identification of anticancer peptides
abstract
MOTIVATION: Anticancer peptides (ACPs) have natural cationic properties and can act on the anionic cell membrane of cancer cells to kill cancer cells. Therefore, ACPs have become a potential anticancer drug with good research value and prospect. RESULTS: In this article, we propose AACFlow, an end-to-end model for identification of ACPs based on deep learning. End-to-end models have more room to automatically adjust according to the data, making the overall fit better and reducing error propagation. The combination of attention augmented convolutional neural network (AAConv) and multi-layer convolutional neural network (CNN) forms a deep representation learning module, which is used to obtain global and local information on the sequence. Based on the concept of flow network, multi-head flow-attention mechanism is introduced to mine the deep features of the sequence to improve the efficiency of the model. On the independent test dataset, the ACC, Sn, Sp, and AUC values of AACFlow are 83.9%, 83.0%, 84.8%, and 0.892, respectively, which are 4.9%, 1.5%, 8.0%, and 0.016 higher than those of the baseline model. The MCC value is 67.85%. In addition, we visualize the features extracted by each module to enhance the interpretability of the model. Various experiments show that our model is more competitive in predicting ACPs.
Shengli Zhang 0002, Yunyun Liang
Bioinform.1
2024 MMD-DTA: A Multi-Modal Deep Learning Framework for Drug-Target Binding Affinity and Binding Region Prediction
abstract
The prediction of drug-target affinity (DTA) plays a crucial role in drug development and the identification of potential drug targets. In recent years, computer-assisted DTA prediction has emerged as a significant approach in this field. In this study, we propose a multi-modal deep learning framework called MMD-DTA for predicting drug-target binding affinity and binding regions. The model can predict DTA while simultaneously learning the binding regions of drug-target interactions through unsupervised learning. To achieve this, MMD-DTA first uses graph neural networks and target structural feature extraction network to extract multi-modal information from the sequences and structures of drugs and targets. It then utilizes the feature interaction and fusion modules to generate interaction descriptors for predicting DTA and interaction strength for binding region prediction. Our experimental results demonstrate that MMD-DTA outperforms existing models based on key evaluation metrics. Furthermore, external validation results indicate that MMD-DTA enhances the generalization capability of the model by integrating sequence and structural information of drugs and targets. The model trained on the benchmark dataset can effectively generalize to independent virtual screening tasks. The visualization of drug-target binding region prediction showcases the interpretability of MMD-DTA, providing valuable insights into the functional regions of drug molecules that interact with proteins.
Qi Zhang 0126, Yuxiao Wei, Liwei Liu 0001, Shengli Zhang 0002
IEEE ACM Trans. Comput. Biol. Bioinform.5
2023 TNFIPs-Net: A deep learning model based on multi-feature fusion for prediction of TNF-α inducing epitopes
abstract
Tumor necrosis factor alpha (TNF-α) is a cytokine belonging to the tumor necrosis factor family. It plays a crucial regulatory role in the immune system and is involved in various biological processes. TNF-α inducing epitopes are specific antigenic epitopes capable of stimulating or inducing the production of TNF-α in cells. By identifying and studying TNF-α inducing epitopes, we gain a better understanding of their relevance to diseases, offering novel targets and intervention strategies for drug development and treatment. Furthermore, in-depth research on TNF-α inducing epitopes provides important insights and guidance for personalized medicine and precision immune therapy. In this study, we propose a novel deep learning model based on multi-feature fusion for predicting TNF-α inducing epitopes (TNFIPs-Net). Our model utilizes a dual-branch architecture guided by adaptive features and hand-crafted features. In the encoder layers, the adaptive word embedding features and hand-crafted features are separately input to different encoders (transformer and BiGRU, respectively). Finally, an attention mechanism is employed in the output layer to fuse the features from both branches. Through experimental validation, the fusion of multiple features and the use of self-attention mechanism enable the model to capture complex feature information more effectively, thereby improving the predictive performance of TNF-α inducing epitopes. The datasets and code used in this research are available at https://github.com/yujiexu321/TNFIPs-Net.
Shengli Zhang 0002, Yunyun Liang, Yuanyuan Jing
BIBM1
2023 PreVFs-RG: A Deep Hybrid Model for Identifying Virulence Factors Based on Residual Block and Gated Recurrent Unit
abstract
Many infectious diseases are caused by bacterial pathogens. The pathogenic mechanisms of bacterial pathogens are complex and it is usually caused by virulence factors (VFs) in many cases. Whether VFs exist is the main difference between the genomes of pathogenic and non-pathogenic bacteria. Therefore, identification of VFs is of great significance for exploring the pathogenesis of infectious diseases. In this paper, we develop a new model for predicting VFs based on multiple features and deep learning. Firstly, we used kmer, Dipeptide deviation from expected mean (DDE) and Amino acid composition (AAC) to extract features. And the deep model is constructed by integrating residual block and gated recurrent unit (GRU). Residual neural network is a kind of jump connection network to avoid gradient explosion or gradient disappearance, which is used to train deeper networks. Two residual blocks, which are separated by Relu activation function and Batch Normalization(BN) layer are used. Finally, the combined features processed by resnet are input into the recurrent structure GRU, and important information is filtered out through reset gate and update gate. The results of the proposed model are better than those of the existing methods by 10-fold cross-validation, and the accuracy is 93.81% and 95.43%, respectively. It shows that the deep hybrid model for identifying virulence factors based on residual block and gate recurrent unit is effective and feasible. The codes and datasets are accessible at https://github.com/VFs625/PreVFs-RG.git.
Shengli Zhang 0002, Yuanyuan Jing
IEEE ACM Trans. Comput. Biol. Bioinform.1
2022 R5hmCFDV: computational identification of RNA 5-hydroxymethylcytosine based on deep feature fusion and deep voting
abstract
RNA 5-hydroxymethylcytosine (5hmC) is a kind of RNA modification, which is related to the life activities of many organisms. Studying its distribution is very important to reveal its biological function. Previously, high-throughput sequencing was used to identify 5hmC, but it is expensive and inefficient. Therefore, machine learning is used to identify 5hmC sites. Here, we design a model called R5hmCFDV, which is mainly divided into feature representation, feature fusion and classification. (i) Pseudo dinucleotide composition, dinucleotide binary profile and frequency, natural vector and physicochemical property are used to extract features from four aspects: nucleotide composition, coding, natural language and physical and chemical properties. (ii) To strengthen the relevance of features, we construct a novel feature fusion method. Firstly, the attention mechanism is employed to process four single features, stitch them together and feed them to the convolution layer. After that, the output data are processed by BiGRU and BiLSTM, respectively. Finally, the features of these two parts are fused by the multiply function. (iii) We design the deep voting algorithm for classification by imitating the soft voting mechanism in the Python package. The base classifiers contain deep neural network (DNN), convolutional neural network (CNN) and improved gated recurrent unit (GRU). And then using the principle of soft voting, the corresponding weights are assigned to the predicted probabilities of the three classifiers. The predicted probability values are multiplied by the corresponding weights and then summed to obtain the final prediction results. We use 10-fold cross-validation to evaluate the model, and the evaluation indicators are significantly improved. The prediction accuracy of the two datasets is as high as 95.41% and 93.50%, respectively. It demonstrates the stronger competitiveness and generalization performance of our model. In addition, all datasets and source codes can be found at https://github.com/HongyanShi026/R5hmCFDV.
Hongyan Shi, Shengli Zhang 0002
Briefings Bioinform.2
2022 An improved residual network using deep fusion for identifying RNA 5-methylcytosine sites
abstract
MOTIVATION: 5-Methylcytosine (m5C) is a crucial post-transcriptional modification. With the development of technology, it is widely found in various RNAs. Numerous studies have indicated that m5C plays an essential role in various activities of organisms, such as tRNA recognition, stabilization of RNA structure, RNA metabolism and so on. Traditional identification is costly and time-consuming by wet biological experiments. Therefore, computational models are commonly used to identify the m5C sites. Due to the vast computing advantages of deep learning, it is feasible to construct the predictive model through deep learning algorithms. RESULTS: In this study, we construct a model to identify m5C based on a deep fusion approach with an improved residual network. First, sequence features are extracted from the RNA sequences using Kmer, K-tuple nucleotide frequency component (KNFC), Pseudo dinucleotide composition (PseDNC) and Physical and chemical property (PCP). Kmer and KNFC extract information from a statistical point of view. PseDNC and PCP extract information from the physicochemical properties of RNA sequences. Then, two parts of information are fused with new features using bidirectional long- and short-term memory and attention mechanisms, respectively. Immediately after, the fused features are fed into the improved residual network for classification. Finally, 10-fold cross-validation and independent set testing are used to verify the credibility of the model. The results show that the accuracy reaches 91.87%, 95.55%, 92.27% and 95.60% on the training sets and independent test sets of Arabidopsis thaliana and M.musculus, respectively. This is a considerable improvement compared to previous studies and demonstrates the robust performance of our model. AVAILABILITY AND IMPLEMENTATION: The data and code related to the study are available at https://github.com/alivelxj/m5c-DFRESG.
Shengli Zhang 0002, Hongyan Shi
Bioinform.2