Yanjing Wang 0003

dblp:283/6579 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0001-6826-1635ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021
YearPublicationVenuePosition
2025 SageTCR: a structure-based model integrating residue- and atom-level representations for enhanced TCR-pMHC binding prediction
abstract
T-cell receptors (TCRs) recognize peptide-MHC (pMHC) complexes through intricate structural interactions, which is a core component of adaptive immunity. However, the diverse and cross-reactive nature of TCRs poses great challenges for accurate prediction of TCR-epitope interactions, hampering the advancement and broad application of TCR-related therapies. Here, we present SageTCR, a bi-level graph neural network (GNN) framework that leverages structural data to predict TCR-pMHC binding possibilities. Harnessing the pretrained language models, SageTCR encodes detailed structural arrangement at both residue-level and atomic-level and effectively integrates the bimodal representations via attention mechanisms. To tackle the deficiency of experimental structures, we explore comprehensive data augmentation strategies to enrich the training and increase the generalizability while concurrently preserving the characteristic TCR-pMHC diagonal binding mode. SageTCR demonstrates superior performance compared to six methods with different deep learning architectures. Furthermore, SageTCR offers the interpretability by identifying and focusing on the conformational features of pivotal contact residues on the interface, which can provide valuable insights for TCR engineering and immunotherapy design.
Xiangyi Li, Chuance Sun, Weiran Huang 0001, Yanjing Wang 0003, Buyong Ma
Briefings Bioinform.4
2022 MDF-SA-DDI: predicting drug-drug interaction events based on multi-source drug fusion, multi-source feature fusion and transformer self-attention mechanism
abstract
One of the main problems with the joint use of multiple drugs is that it may cause adverse drug interactions and side effects that damage the body. Therefore, it is important to predict potential drug interactions. However, most of the available prediction methods can only predict whether two drugs interact or not, whereas few methods can predict interaction events between two drugs. Accurately predicting interaction events of two drugs is more useful for researchers to study the mechanism of the interaction of two drugs. In the present study, we propose a novel method, MDF-SA-DDI, which predicts drug-drug interaction (DDI) events based on multi-source drug fusion, multi-source feature fusion and transformer self-attention mechanism. MDF-SA-DDI is mainly composed of two parts: multi-source drug fusion and multi-source feature fusion. First, we combine two drugs in four different ways and input the combined drug feature representation into four different drug fusion networks (Siamese network, convolutional neural network and two auto-encoders) to obtain the latent feature vectors of the drug pairs, in which the two auto-encoders have the same structure, and their main difference is the number of neurons in the input layer of the two auto-encoders. Then, we use transformer blocks that include self-attention mechanism to perform latent feature fusion. We conducted experiments on three different tasks with two datasets. On the small dataset, the area under the precision-recall-curve (AUPR) and F1 scores of our method on task 1 reached 0.9737 and 0.8878, respectively, which were better than the state-of-the-art method. On the large dataset, the AUPR and F1 scores of our method on task 1 reached 0.9773 and 0.9117, respectively. In task 2 and task 3 of two datasets, our method also achieved the same or better performance as the state-of-the-art method. More importantly, the case studies on five DDI events are conducted and achieved satisfactory performance. The source codes and data are available at https://github.com/ShenggengLin/MDF-SA-DDI.
Shenggeng Lin, Yanjing Wang 0003, Yanyi Chu, Yatong Liu, Yitian Fang, Yi Xiong 0002
Briefings Bioinform.2
2021 DTI-MLCD: predicting drug-target interactions using multi-label learning with community detection method
abstract
Identifying drug-target interactions (DTIs) is an important step for drug discovery and drug repositioning. To reduce the experimental cost, a large number of computational approaches have been proposed for this task. The machine learning-based models, especially binary classification models, have been developed to predict whether a drug-target pair interacts or not. However, there is still much room for improvement in the performance of current methods. Multi-label learning can overcome some difficulties caused by single-label learning in order to improve the predictive performance. The key challenge faced by multi-label learning is the exponential-sized output space, and considering label correlations can help to overcome this challenge. In this paper, we facilitate multi-label classification by introducing community detection methods for DTI prediction, named DTI-MLCD. Moreover, we updated the gold standard data set by adding 15,000 more positive DTI samples in comparison to the data set, which has widely been used by most of previously published DTI prediction methods since 2008. The proposed DTI-MLCD is applied to both data sets, demonstrating its superiority over other machine learning methods and several existing methods. The data sets and source code of this study are freely available at https://github.com/a96123155/DTI-MLCD.
Yanyi Chu, Xiaoqi Shan, Tianhang Chen, Yanjing Wang 0003, Dennis R. Salahub, Yi Xiong 0002
Briefings Bioinform.5
2021 MDA-GCNFTG: identifying miRNA-disease associations based on graph convolutional networks via graph sampling through the feature and topology graph
abstract
Accurate identification of the miRNA-disease associations (MDAs) helps to understand the etiology and mechanisms of various diseases. However, the experimental methods are costly and time-consuming. Thus, it is urgent to develop computational methods towards the prediction of MDAs. Based on the graph theory, the MDA prediction is regarded as a node classification task in the present study. To solve this task, we propose a novel method MDA-GCNFTG, which predicts MDAs based on Graph Convolutional Networks (GCNs) via graph sampling through the Feature and Topology Graph to improve the training efficiency and accuracy. This method models both the potential connections of feature space and the structural relationships of MDA data. The nodes of the graphs are represented by the disease semantic similarity, miRNA functional similarity and Gaussian interaction profile kernel similarity. Moreover, we considered six tasks simultaneously on the MDA prediction problem at the first time, which ensure that under both balanced and unbalanced sample distribution, MDA-GCNFTG can predict not only new MDAs but also new diseases without known related miRNAs and new miRNAs without known related diseases. The results of 5-fold cross-validation show that the MDA-GCNFTG method has achieved satisfactory performance on all six tasks and is significantly superior to the classic machine learning methods and the state-of-the-art MDA prediction methods. Moreover, the effectiveness of GCNs via the graph sampling strategy and the feature and topology graph in MDA-GCNFTG has also been demonstrated. More importantly, case studies for two diseases and three miRNAs are conducted and achieved satisfactory performance.
Yanyi Chu, Xuhong Wang, Qiuying Dai, Yanjing Wang 0003, Shaoliang Peng, Xiaoyong Wei, Jingfei Qiu, Dennis R. Salahub, Yi Xiong 0002
Briefings Bioinform.4
2021 NeuroPpred-Fuse: an interpretable stacking model for prediction of neuropeptides by fusing sequence information and feature selection methods
abstract
Neuropeptides acting as signaling molecules in the nervous system of various animals play crucial roles in a wide range of physiological functions and hormone regulation behaviors. Neuropeptides offer many opportunities for the discovery of new drugs and targets for the treatment of neurological diseases. In recent years, there have been several data-driven computational predictors of various types of bioactive peptides, but the relevant work about neuropeptides is little at present. In this work, we developed an interpretable stacking model, named NeuroPpred-Fuse, for the prediction of neuropeptides through fusing a variety of sequence-derived features and feature selection methods. Specifically, we used six types of sequence-derived features to encode the peptide sequences and then combined them. In the first layer, we ensembled three base classifiers and four feature selection algorithms, which select non-redundant important features complementarily. In the second layer, the output of the first layer was merged and fed into logistic regression (LR) classifier to train the model. Moreover, we analyzed the selected features and explained the feasibility of the selected features. Experimental results show that our model achieved 90.6% accuracy and 95.8% AUC on the independent test set, outperforming the state-of-the-art models. In addition, we exhibited the distribution of selected features by these tree models and compared the results on the training set to that on the test set. These results fully showed that our model has a certain generalization ability. Therefore, we expect that our model would provide important advances in the discovery of neuropeptides as new drugs for the treatment of neurological diseases.
Shenggan Luo, Yanyi Chu, Tianhang Chen, Xueying Mao, Yatong Liu, Yanjing Wang 0003, Yi Xiong 0002
Briefings Bioinform.9
2021 Irinotecan and vandetanib create synergies for treatment of pancreatic cancer patients with concomitant TP53 and KRAS mutations
abstract
BACKGROUND: The most frequently mutated gene pairs in pancreatic adenocarcinoma (PAAD) are KRAS and TP53, and our goal is to illustrate the multiomics and molecular dynamics landscapes of KRAS/TP53 mutation and also to obtain prospective novel drugs for KRAS- and TP53-mutated PAAD patients. Moreover, we also made an attempt to discover the probable link amid KRAS and TP53 on the basis of the abovementioned multiomics data. METHOD: We utilized TCGA & Cancer Cell Line Encyclopedia data for the analysis of KRAS/TP53 mutation in a multiomics manner. In addition to that, we performed molecular dynamics analysis of KRAS and TP53 to produce mechanistic descriptions of particular mutations and carcinogenesis. RESULT: We discover that there is a significant difference in the genomics, transcriptomics, methylomics, and molecular dynamics pattern of KRAS and TP53 mutation from the matching wild type in PAAD, and the prognosis of pancreatic cancer is directly linked with a particular mutation of KRAS and protein stability. Screened drugs are potentially effective in PAAD patients. CONCLUSIONS: KRAS and TP53 prognosis of PAAD is directly associated with a specific mutation of KRAS. Irinotecan and vandetanib are prospective drugs for PAAD patients with KRASG12Dmutation and TP53 mutation.
Aman Chandra Kaushik, Yanjing Wang 0003, Xiangeng Wang
Briefings Bioinform.2