Weihe Dong

dblp:349/8002 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0001-8022-9793ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 8 since 2021
YearPublicationVenuePosition
2025 TCRdesign: an antigen-specific generative language model for de novo design of T-cell receptors
abstract
T-cell receptors (TCR), which are heterodimers of $\alpha $ and $\beta $ chains that recognize foreign antigens, are of great significance to current immunotherapy. Although artificial intelligence (AI) has explosively accelerated de novo protein design, the challenge of therapeutic TCR design has been overlooked by most researchers. Existing TCR engineering relies heavily on isolating antigen-specific TCRs from tumor tissues, which requires a large amount of labor resources and wet experimental verification. To mitigate this issue, we present TCRdesign, a pretrained generative protein language model (PLM) for the de novo design of artificial TCR $\beta $-chain complementarity-determining region 3 sequences conditioned on antigen-binding specificity (BS). In parallel, we develop a high-accuracy binding predictor (TCRBinder) that couples paired $\alpha $/$\beta $ chain information with antigen sequences to assess BS. Our in silico comparisons demonstrate that (i) TCRdesign surpasses state-of-the-art baselines in generating antigen-specific TCR sequences. The model leverages paired-chain coherence to refine amino-acid level interaction patterns. (ii) TCRdesign-generated TCR sequences exhibit better antigen binding capability to diverse oncogenic hotspots compared with natural counterparts. (iii) TCRdesign inherits the intrinsic properties of large PLMs, enabling effectively identify the determinant residues in TCR-antigen binding, which enhances its interpretability. These results highlight the significant capability of TCRdesign in understanding and generating TCR sequences with an antigen-specific interaction pattern, charting a versatile path toward AI-driven T-cell engineering for precision immunotherapy.
Xiaokun Li, Qiang Yang 0015, Weihe Dong, Kuanquan Wang, Suyu Dong, Wei Wang 0169, Gongning Luo, Xianyu Zhang 0004, Tiansong Yang, Xin Gao 0001, Guohua Wang 0001
Briefings Bioinform.4
2025 AdaSemb: an adaptive knowledge-driven deep learning framework integrating cancer protein assemblies for predicting PI3Kα inhibitor response and resistance
abstract
Protein kinases regulate diverse cellular functions, including cell cycle progression, metabolism, differentiation, and survival, with their dysregulation implicated in multiple carcinogenic processes. Phosphatidylinositol 3-kinase alpha inhibitors (PI3K$ \alpha $is) have revolutionized breast cancer treatment, but acquired resistance remains a major clinical challenge, with around 40% of patients experiencing progression within 4-6 months. Current drug response prediction (DRP) methods typically rely on individual pathways or biomarkers, limiting their ability to capture complex cancer-specific molecular interactions and predict resistance mechanisms. To overcome these limitations, we present AdaSemb, an adaptive, knowledge-driven deep learning framework that uses a multi-protein assembly map to predict responses and resistance to PI3K$ \alpha $i. AdaSemb comprises two modules: the AdaSemb-PA module incorporates tumor genomic variations into a biological structural neural network, while the AdaSemb-DRP module uses conditional domain adversarial networks to enhance gene-drug distribution generalization. By combining genomic data with drug molecular structures, AdaSemb identifies critical protein combinations linked to drug resistance. In validation with 1244 cancer cell lines and patient-derived xenografts (PDX), AdaSemb outperformed existing DRP models. In a cohort of 116 breast cancer patients from the Cancer Genome Atlas (TCGA), it predicted significantly longer survival for sensitive patients, surpassing traditional biomarkers in precision. Furthermore, we identified seven key assemblages that integrate mutations from 93 genes, which distinguish alpelisib sensitive and resistant cell lines. These results are applicable to breast cancer patient samples and PDX models, demonstrating AdaSemb's significant clinical potential in personalized treatment and prediction of resistance for breast cancer.
Zaiduo Li, Qiang Yang 0001, Weihe Dong, Xiaochuan Yang, Xianyu Zhang 0004, Tiansong Yang, Xiaokun Li
Briefings Bioinform.4
2025 Meta learning for mutant HLA class I epitope immunogenicity prediction to accelerate cancer clinical immunotherapy
abstract
Accurate prediction of binding between human leukocyte antigen (HLA) class I molecules and antigenic peptide segments is a challenging task and a key bottleneck in personalized immunotherapy for cancer. Although existing prediction tools have demonstrated significant results using established datasets, most can only predict the binding affinity of antigenic peptides to HLA and do not enable the immunogenic interpretation of new antigenic epitopes. This limitation results from the training data for the computational models relying heavily on a large amount of peptide-HLA (pHLA) eluting ligand data, in which most of the candidate epitopes lack immunogenicity. Here, we propose an adaptive immunogenicity prediction model, named MHLAPre, which is trained on the large-scale MS-derived HLA I eluted ligandome (mostly presented by epitopes) that are immunogenic. Allele-specific and pan-allelic prediction models are also provided for endogenous peptide presentation. Using a meta-learning strategy, MHLAPre rapidly assessed HLA class I peptide affinities across the whole pHLA pairs and accurately identified tumor-associated endogenous antigens. During the process of adaptive immune response of T-cells, pHLA-specific binding in the antigen presentation is only a pre-task for CD8+ T-cell recognition. The key factor in activating the immune response is the interaction between pHLA complexes and T-cell receptors (TCRs). Therefore, we performed transfer learning on the pHLA model using the pHLA-TCR dataset. In pHLA binding task, MHLAPre demonstrated significant improvement in identifying neoepitope immunogenicity compared with five state-of-the-art models, proving its effectiveness and robustness. After transfer learning of the pHLA-TCR data, MHLAPre also exhibited relatively superior performance in revealing the mechanism of immunotherapy. MHLAPre is a powerful tool to identify neoepitopes that can interact with TCR and induce immune responses. We believe that the proposed method will greatly contribute to clinical immunotherapy, such as anti-tumor immunity, tumor-specific T-cell engineering, and personalized tumor vaccine.
Qiang Yang 0015, Weihe Dong, Xiaokun Li, Kuanquan Wang, Suyu Dong, Xianyu Zhang 0004, Tiansong Yang, Gongning Luo, Xingyu Liao, Xin Gao 0001, Guohua Wang 0001
Briefings Bioinform.3
2025 THLANet: A deep learning framework for predicting TCR-pHLA binding in immunotherapy applications
abstract
Adaptive immunity is a targeted immune response that enables the body to identify and eliminate foreign pathogens, playing a critical role in the anti-tumor immune response. Tumor cell expression of antigens forms the foundation for inducing this adaptive response. However, the human leukocyte antigens (HLA)-restricted recognition of antigens by T-cell receptors (TCR) limits their ability to detect all neoantigens, with only a small subset capable of activating T-cells. Accurately predicting neoantigen binding to TCR is, therefore, crucial for assessing their immunogenic potential in clinical settings. We present THLANet, a deep learning model designed to predict the binding specificity of TCR to neoantigens presented by class I HLAs. THLANet employs evolutionary scale modeling-2 (ESM-2), replacing the traditional embedding methods to enhance sequence feature representation. Using scTCR-seq data, we obtained the TCR immune repertoire and constructed a TCR-pHLA binding database to validate THLANet's clinical potential. The model's performance was further evaluated using clinical cancer data across various cancer types. Additionally, by analyzing divided complementarity-determining region (CDR3) sequences and simulating alanine scanning of antigen sequences, we provided new insights into the 3D binding interactions of TCRs and antigens. Predicting TCR-neoantigen pairing remains a significant challenge in immunology, THLANet provides accurate predictions using only the TCR sequence (CDR3β), antigen sequence, and class I HLA, offering novel insights into TCR-antigen interactions.
Qiang Yang 0015, Weihe Dong, Xiaokun Li, Kuanquan Wang, Suyu Dong, Gongning Luo, Xianyu Zhang 0004, Tiansong Yang, Xin Gao 0001, Guohua Wang 0001
PLoS Comput. Biol.3
2024 HLAIImaster: a deep learning method with adaptive domain knowledge predicts HLA II neoepitope immunogenic responses
abstract
While significant strides have been made in predicting neoepitopes that trigger autologous CD4+ T cell responses, accurately identifying the antigen presentation by human leukocyte antigen (HLA) class II molecules remains a challenge. This identification is critical for developing vaccines and cancer immunotherapies. Current prediction methods are limited, primarily due to a lack of high-quality training epitope datasets and algorithmic constraints. To predict the exogenous HLA class II-restricted peptides across most of the human population, we utilized the mass spectrometry data to profile >223 000 eluted ligands over HLA-DR, -DQ, and -DP alleles. Here, by integrating these data with peptide processing and gene expression, we introduce HLAIImaster, an attention-based deep learning framework with adaptive domain knowledge for predicting neoepitope immunogenicity. Leveraging diverse biological characteristics and our enhanced deep learning framework, HLAIImaster is significantly improved against existing tools in terms of positive predictive value across various neoantigen studies. Robust domain knowledge learning accurately identifies neoepitope immunogenicity, bridging the gap between neoantigen biology and the clinical setting and paving the way for future neoantigen-based therapies to provide greater clinical benefit. In summary, we present a comprehensive exploitation of the immunogenic neoepitope repertoire of cancers, facilitating the effective development of "just-in-time" personalized vaccines.
Qiang Yang 0015, Weihe Dong, Xiaokun Li, Kuanquan Wang, Suyu Dong, Xianyu Zhang 0004, Tiansong Yang, Feng Jiang 0001, Bin Zhang 0042, Gongning Luo, Xin Gao 0001, Guohua Wang 0001
Briefings Bioinform.3
2024 DrugMGR: a deep bioactive molecule binding method to identify compounds targeting proteins
abstract
MOTIVATION: Understanding the intermolecular interactions of ligand-target pairs is key to guiding the optimization of drug research on cancers, which can greatly mitigate overburden workloads for wet labs. Several improved computational methods have been introduced and exhibit promising performance for these identification tasks, but some pitfalls restrict their practical applications: (i) first, existing methods do not sufficiently consider how multigranular molecule representations influence interaction patterns between proteins and compounds; and (ii) second, existing methods seldom explicitly model the binding sites when an interaction occurs to enable better prediction and interpretation, which may lead to unexpected obstacles to biological researchers. RESULTS: To address these issues, we here present DrugMGR, a deep multigranular drug representation model capable of predicting binding affinities and regions for each ligand-target pair. We conduct consistent experiments on three benchmark datasets using existing methods and introduce a new specific dataset to better validate the prediction of binding sites. For practical application, target-specific compound identification tasks are also carried out to validate the capability of real-world compound screen. Moreover, the visualization of some practical interaction scenarios provides interpretable insights from the results of the predictions. The proposed DrugMGR achieves excellent overall performance in these datasets, exhibiting its advantages and merits against state-of-the-art methods. Thus, the downstream task of DrugMGR can be fine-tuned for identifying the potential compounds that target proteins for clinical treatment. AVAILABILITY AND IMPLEMENTATION: https://github.com/lixiaokun2020/DrugMGR.
Xiaokun Li, Qiang Yang 0015, Weihe Dong, Gongning Luo, Wei Wang 0169, Suyu Dong, Kuanquan Wang, Ping Xuan, Xianyu Zhang 0004, Xin Gao 0001
Bioinform.4
2023 Graph Convolutional Network with Neural Inductive Matrix Completion for Predicting Disease-Related LncRNA Genes
abstract
Numerous researches emphasized that long non-coding RNA (lncRNA) plays a vital factor in various biological processes, and its mismatched expression and dysfunction are tightly linked with the occurrence of human diseases. Thus, computational models were designed to identify lncRNA-disease interactions by merging heterogeneous biological data. However, most of them neglected the intrinsic structure of multi-source information, which limits the performance for potential lncRNA-disease association prediction. Here, GCN-NIMC is introduced to alleviate the dilemma for disease-associated lncRNA genes identification based on the graph convolutional network with neural inductive matrix. This method builds a feature matrix with multi-source heterogeneous data and then learn the various information contained in the feature matrix for the sake of acquiring better feature expressions of the lncRNA-disease interactions. Experimental results on 10-repeated 5-fold cross-validation demonstrated that our proposed GCN-NIMC is superior to existing cutting-edge methods for identifying disease-related lncRNA genes. Furthermore, case studies confirmed our computational method as a practical tool with clinical benefits to develop the therapeutic schedule at lncRNA-level.
Qiang Yang 0015, Suyu Dong, Weihe Dong, Xiaokun Li, Pengzhong Sun, Feng Jiang 0001, Xianyu Zhang 0004, Gongning Luo
BIBM5
2023 Multi-modality attribute learning-based method for drug-protein interaction prediction based on deep neural network
abstract
Identification of active candidate compounds for target proteins, also called drug-protein interaction (DPI) prediction, is an essential but time-consuming and expensive step, which leads to fostering the development of drug discovery. In recent years, deep network-based learning methods were frequently proposed in DPIs due to their powerful capability of feature representation. However, the performance of existing DPI methods is still limited by insufficiently labeled pharmacological data and neglected intermolecular information. Therefore, overcoming these difficulties to perfect the performance of DPIs is an urgent challenge for researchers. In this article, we designed an innovative 'multi-modality attributes' learning-based framework for DPIs with molecular transformer and graph convolutional networks, termed, multi-modality attributes (MMA)-DPI. Specifically, intermolecular sub-structural information and chemical semantic representations were extracted through an augmented transformer module from biomedical data. A tri-layer graph convolutional neural network module was applied to associate the neighbor topology information and learn the condensed dimensional features by aggregating a heterogeneous network that contains multiple biological representations of drugs, proteins, diseases and side effects. Then, the learned representations were taken as the input of a fully connected neural network module to further integrate them in molecular and topological space. Finally, the attribute representations were fused with adaptive learning weights to calculate the interaction score for the DPIs tasks. MMA-DPI was evaluated in different experimental conditions and the results demonstrate that the proposed method achieved higher performance than existing state-of-the-art frameworks.
Weihe Dong, Qiang Yang 0015, Xiaokun Li, Gongning Luo, Xin Gao 0001
Briefings Bioinform.1