Naifeng Wen

dblp:198/5613 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021
YearPublicationVenuePosition
2024 Drug-Target Binding Affinity Prediction in a Continuous Latent Space Using Variational Autoencoders
abstract
Accurate prediction of Drug-Target binding Affinity (DTA) is a daunting yet pivotal task in the sphere of drug discovery. Over the years, a plethora of deep learning-based DTA models have emerged, rendering promising results in predicting the binding affinities between drugs and their target proteins. However, in contrast to the conventional approach of modeling binding affinity in vector spaces, we propose a more nuanced modeling process in a continuous space to account for the diversity of input samples. Initially, the drug is encoded using the Simplified Molecular Input Line Entry System (SMILES), while the target sequences are characterized via a pretrained language model. Subsequently, highly correlative information is extracted utilizing residual gated convolutional neural networks. In a departure from existing deep learning-based models, our model learns the hidden representations of the drugs and targets jointly. Instead of employing two vectors, our hidden representations consist of two Gaussian distributions. To validate the effectiveness of our proposal, we conducted evaluations on commonly utilized benchmark datasets. The experimental outcomes corroborated that our method surpasses the state-of-the-art vectorial representation methods in terms of performance. This approach, therefore, offers potential enhancements in the precision of DTA predictions, potentially contributing to more efficient drug discovery processes.
Lingling Zhao, Yan Zhu 0006, Naifeng Wen, Chunyu Wang 0002, Junjie Wang 0005, Yongfeng Yuan
IEEE ACM Trans. Comput. Biol. Bioinform.3
2023 Multiple sequence alignment based on deep reinforcement learning with self-attention and positional encoding
abstract
MOTIVATION: Multiple sequence alignment (MSA) is one of the hotspots of current research and is commonly used in sequence analysis scenarios. However, there is no lasting solution for MSA because it is a Nondeterministic Polynomially complete problem, and the existing methods still have room to improve the accuracy. RESULTS: We propose Deep reinforcement learning with Positional encoding and self-Attention for MSA, based on deep reinforcement learning, to enhance the accuracy of the alignment Specifically, inspired by the translation technique in natural language processing, we introduce self-attention and positional encoding to improve accuracy and reliability. Firstly, positional encoding encodes the position of the sequence to prevent the loss of nucleotide position information. Secondly, the self-attention model is used to extract the key features of the sequence. Then input the features into a multi-layer perceptron, which can calculate the insertion position of the gap according to the features. In addition, a novel reinforcement learning environment is designed to convert the classic progressive alignment into progressive column alignment, gradually generating each column's sub-alignment. Finally, merge the sub-alignment into the complete alignment. Extensive experiments based on several datasets validate our method's effectiveness for MSA, outperforming some state-of-the-art methods in terms of the Sum-of-pairs and Column scores. AVAILABILITY AND IMPLEMENTATION: The process is implemented in Python and available as open-source software from https://github.com/ZhangLab312/DPAMSA.
Zixuan Wang 0025, Shuwen Xiong, Naifeng Wen, Yongqing Zhang 0001
Bioinform.6
2023 DataDTA: a multi-feature and dual-interaction aggregation framework for drug-target binding affinity prediction
abstract
MOTIVATION: Accurate prediction of drug-target binding affinity (DTA) is crucial for drug discovery. The increase in the publication of large-scale DTA datasets enables the development of various computational methods for DTA prediction. Numerous deep learning-based methods have been proposed to predict affinities, some of which only utilize original sequence information or complex structures, but the effective combination of various information and protein-binding pockets have not been fully mined. Therefore, a new method that integrates available key information is urgently needed to predict DTA and accelerate the drug discovery process. RESULTS: In this study, we propose a novel deep learning-based predictor termed DataDTA to estimate the affinities of drug-target pairs. DataDTA utilizes descriptors of predicted pockets and sequences of proteins, as well as low-dimensional molecular features and SMILES strings of compounds as inputs. Specifically, the pockets were predicted from the three-dimensional structure of proteins and their descriptors were extracted as the partial input features for DTA prediction. The molecular representation of compounds based on algebraic graph features was collected to supplement the input information of targets. Furthermore, to ensure effective learning of multiscale interaction features, a dual-interaction aggregation neural network strategy was developed. DataDTA was compared with state-of-the-art methods on different datasets, and the results showed that DataDTA is a reliable prediction tool for affinities estimation. Specifically, the concordance index (CI) of DataDTA is 0.806 and the Pearson correlation coefficient (R) value is 0.814 on the test dataset, which is higher than other methods. AVAILABILITY AND IMPLEMENTATION: The codes and datasets of DataDTA are available at https://github.com/YanZhu06/DataDTA.
Yan Zhu 0006, Lingling Zhao, Naifeng Wen, Junjie Wang 0005, Chunyu Wang 0002
Bioinform.3
2022 Learning representations for gene ontology terms by jointly encoding graph structure and textual node descriptors
abstract
Measuring the semantic similarity between Gene Ontology (GO) terms is a fundamental step in numerous functional bioinformatics applications. To fully exploit the metadata of GO terms, word embedding-based methods have been proposed recently to map GO terms to low-dimensional feature vectors. However, these representation methods commonly overlook the key information hidden in the whole GO structure and the relationship between GO terms. In this paper, we propose a novel representation model for GO terms, named GT2Vec, which jointly considers the GO graph structure obtained by graph contrastive learning and the semantic description of GO terms based on BERT encoders. Our method is evaluated on a protein similarity task on a collection of benchmark datasets. The experimental results demonstrate the effectiveness of using a joint encoding graph structure and textual node descriptors to learn vector representations for GO terms.
Lingling Zhao, Huiting Sun, Xinyi Cao, Naifeng Wen, Junjie Wang 0005, Chunyu Wang 0002
Briefings Bioinform.4
2021 SeqGO-CPA: Improving Compound-Protein Binding Affinity Prediction with Sequence Information and Gene Ontology Knowledge
abstract
The compound-protein binding affinity (CPA) pre-diction is vital for drug discovery and drug repurposing. Deep learning methods have been developed to model the complicated relationship between CPA and the sequences or structures of proteins and molecules. This study proposes a novel deep learning method, SeqGO-CPA, integrating protein function knowledge represented by Gene Ontology (GO) annotations in the CPA prediction. To capture the semantic information of GO annotations, a fine-tuned natural language processing model for biomedical domains is utilized to encode the set of GO terms. Meanwhile, based on the observation that CPA often occurs in sub-structures, our method uses the tokenization algorithm to learn sub-structure information of proteins and compounds from a large number of unlabeled sequences. Further, a deep neural network architecture involving the jointly-feature representation and a highway block is developed to enhance the CPA prediction ability. The proposed model was evaluated on two public benchmark datasets in both standard cross-validation and blinding split settings. The experimental results demonstrate our method outperforms the deep learning-based baselines, meanwhile the incorporating of GO information further improves the prediction performance.
Chunyu Wang 0002, Yan Zhu 0006, Naifeng Wen, Lingling Zhao, Junjie Wang 0005
BIBM3