Yang Tan 0001

dblp:131/1387-1 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
0009-0004-7261-1705ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Meta-Learning Inspired Single-Step Generative Model for Expensive Multitask Optimization Problems
abstract
In expensive multitask optimization problems (ExMTOPs), multiple complex tasks must be optimized simultaneously under limited computational budgets. Existing approaches, often based on surrogate models, aim to approximate objective functions but struggle to generalize across heterogeneous tasks, depend on task-specific sampling, and require frequent retraining. To address these challenges, we propose the Multifactorial Evolutionary Algorithm–Single Step Generative Model (MFEA-SSG), a meta-learning-inspired framework that learns to generate high-quality solutions across tasks. Inspired by meta-learning, we treat each random shuffle of the decision variables as a unique pseudo-task, training the model on a distribution of these tasks to learn a task-agnostic prior about the structure of elite solutions. This process disrupts task-specific dependencies, allowing the model to learn transferable structures from recomposed samples. We then adopt a diffusion-based generative model to learn the distribution of optimal solutions, enabling knowledge transfer across tasks without directly approximating objective functions. To reduce inference cost, we introduce a student model distilled from the diffusion process. Unlike conventional diffusion models that denoise iteratively, the student generates solutions in a single forward pass, significantly reducing inference time. Comprehensive experiments on both general multitask benchmarks and a real-world protein mutation prediction scenario demonstrate that MFEA-SSG achieves high-quality solutions with fast convergence and low computational cost under limited evaluation budgets, outperforming state-of-the-art general and ExMTOPs algorithms.
Xiang Feng 0002, Huiqun Yu, Yang Tan 0001, Edmund M.-K. Lai
IEEE Trans. Evol. Comput.4
2025 Immunogenicity Prediction with Dual Attention Enables Vaccine Target Selection
abstract
Immunogenicity prediction is a central topic in reverse vaccinology for finding candidate vaccines that can trigger protective immune responses. Existing approaches typically rely on highly compressed features and simple model architectures, leading to limited prediction accuracy and poor generalizability. To address these challenges, we introduce VenusVaccine, a novel deep learning solution with a dual attention mechanism that integrates pre-trained latent vector representations of protein sequences and structures. We also compile the most comprehensive immunogenicity dataset to date, encompassing over 7000 antigen sequences, structures, and immunogenicity labels from bacteria, viruses, and tumors. Extensive experiments demonstrate that VenusVaccine outperforms existing methods across a wide range of evaluation metrics. Furthermore, we establish a post-hoc validation protocol to assess the practical significance of deep learning models in tackling vaccine design challenges. Our work provides an effective tool for vaccine design and sets valuable benchmarks for future research. The implementation is at \url{https://github.com/songleee/VenusVaccine}.
Yang Tan 0001, Song Ke, Bingxin Zhou
ICLR2
2025 From high-throughput evaluation to wet-lab studies: advancing mutation effect prediction with a retrieval-enhanced model
abstract
MOTIVATION: Enzyme engineering is a critical approach for producing enzymes that meet industrial and research demands by modifying wild-type proteins to enhance properties such as catalytic activity and thermostability. Beyond traditional directed evolution and rational design, recent advancements in deep learning offer cost-effective and high-performance alternatives. By encoding implicit coevolutionary patterns, these pretrained models have become powerful tools, with the central challenge being to uncover the intricate relationships among protein sequence, structure, and function. RESULTS: We present VenusREM, a retrieval-enhanced protein language model designed to capture local amino acid interactions in both spatial and temporal scales. VenusREM achieves state-of-the-art performance on 217 assays from the ProteinGym benchmark. Beyond high-throughput open benchmark validations, we conducted a low-throughput post hoc analysis on more than 30 mutants to verify the model's ability to improve the stability and binding affinity of a VHH antibody. We also validated the effectiveness of VenusREM by designing 10 novel mutants of a DNA polymerase and performing wet-lab experiments to evaluate their enhanced activity at elevated temperatures. Both in silico and experimental evaluations not only confirm the reliability of VenusREM as a computational tool for enzyme engineering but also demonstrate a comprehensive evaluation framework for future computational studies in mutation effect prediction. AVAILABILITY AND IMPLEMENTATION: The implementation is available at https://github.com/tyang816/VenusREM.
Yang Tan 0001, Banghao Wu, Bingxin Zhou
Bioinform.1
2025 Sequence-only prediction of binding affinity changes: a robust and interpretable model for antibody engineering
abstract
MOTIVATION: A pivotal area of research in antibody engineering is to find effective modifications that enhance antibody-antigen binding affinity. Traditional wet-lab experiments assess mutants in a costly and time-consuming manner. Emerging deep learning solutions offer an alternative by modeling antibody structures to predict binding affinity changes. However, they heavily depend on high-quality complex structures, which are frequently unavailable in practice. Therefore, we propose ProtAttBA, a deep learning model that predicts binding affinity changes based solely on the sequence information of antibody-antigen complexes. RESULTS: ProtAttBA employs a pre-training phase to learn protein sequence patterns, following a supervised training phase using labeled antibody-antigen complex data to train a cross-attention-based regressor for predicting binding affinity changes. We evaluated ProtAttBA on three open benchmarks under different conditions. Compared to both sequence- and structure-based prediction methods, our approach achieves competitive performance, demonstrating notable robustness, especially with uncertain complex structures. Notably, our method possesses interpretability from the attention mechanism. We show that the learned attention scores can identify critical residues with impacts on binding affinity. This work introduces a rapid and cost-effective computational tool for antibody engineering, with the potential to accelerate the development of novel therapeutic antibodies. AVAILABILITY AND IMPLEMENTATION: Source codes and data are available at https://github.com/code4luck/ProtAttBA.
Yang Tan 0001, Wenrui Gou, Guisheng Fan, Bingxin Zhou
Bioinform.3
2025 SRPM-Sol: A Structure Robust Protein Multimodal Model for Solubility Prediction
abstract
The solubility of natural proteins is closely linked to their expression and purification processes. Accurate computational prediction of protein solubility not only aids in functional assessment but also reduces the cost of preliminary wet-lab experiments. The current mainstream deep learning prediction methods have begun to explore the multimodal framework. However, existing multimodal models mainly focus on sequences and structures, overlooking other influential factors. Additionally, inherent errors in predicted structure information pose a significant challenge to model robustness. To solve the above issues, we introduce SRPM-Sol, a novel multimodal protein solubility prediction model. Built upon the state-of-the-art ESM3 model, this framework combines amino acid sequences, structure information, secondary structure sequences, and physicochemical properties for more accurate prediction. This is the most diverse-input model to date in the protein multimodal field for solubility prediction. In order to verify the effectiveness of our method, we have constructed the first hierarchical dataset, PDE-Sol, by organizing data based on Predicted Local Distance Difference Test (pLDDT) scores. The experimental results demonstrate that compared to the baselines, SRPM-Sol achieves stronger robustness and higher accuracy on different levels of PDE-Sol, even in the presence of uncertain structure information.
Wenhui Ge, Yang Tan 0001, Huiqun Yu, Guisheng Fan
IEEE Trans. Comput. Biol. Bioinform.2
2024 Secondary Structure-Guided Novel Protein Sequence Generation with Latent Graph Diffusion
abstract
Designing protein sequences with restrictions or conditions is an important research topic in biology. Many powerful deep generative models have been proposed to create proteins belonging to specific families or with determined backbone structures. However, the amount of homologous data is not always sufficient for any proteins to train a model, and proteins from the same family may lack the necessary structural similarity, posing challenges in ensuring the presence of crucial structures in the generated proteins. On the other hand, when generating proteins with fixed backbone, there exists a trade-off between reliability and flexibility of sequence generation, necessitating prior specification of protein length and precise positions of amino acids. This work introduces a flexible protein generation method for amino acid sequence generation with latent diffusion models and protein language models. The generation is conditioned on protein secondary structures to address the practical considerations in bioengineering better. It enables the imposition of structural constraints on generated proteins while ensuring an adequate level of novelty and diversity at the sequence level. We compare the performance of our method against popular language models and structure-based methods using quantifiable metrics, demonstrating its superiority in generating diverse and novel sequences that exhibit high foldability. Furthermore, we provide case studies of generating proteins with specific secondary structures to analyze the biological significance of our method. The source code is publicly available at https://github.com/riacd/CPDiffusion-SS.
Yutong Hu 0008, Yang Tan 0001, Andi Han, Lirong Zheng 0005, Bingxin Zhou
BIBM2
2024 Protein Representation Learning with Sequence Information Embedding: Does it Always Lead to a Better Performance?
abstract
Deep learning has become a crucial tool in studying proteins. While the significance of modeling protein structure has been discussed extensively in the literature, amino acid types are typically included in the input as a default operation for many inference tasks. This study demonstrates with structure alignment task that embedding amino acid types in some cases may not help a deep learning model learn better representation. To this end, we propose ProtLOCA, a local geometry alignment method based solely on amino acid structure representation. The effectiveness of ProtLOCA is examined by a global structure-matching task on protein pairs with an independent test dataset based on CATH labels. Our method outperforms existing sequence-and structure-based representation learning methods by more quickly and accurately matching structurally consistent protein domains. Furthermore, in local structure pairing tasks, ProtLOCA for the first time provides a valid solution to highlight common local structures among proteins with different overall structures but the same function. This suggests a new possibility for using deep learning methods to analyze protein structure to infer function.
Yang Tan 0001, Lirong Zheng 0005, Bozitao Zhong, Bingxin Zhou
BIBM1
2024 PROTSOLM: Protein Solubility Prediction with Multi-modal Features
Yang Tan 0001, Bingxin Zhou
BIBM1
2024 ProSST: Protein Language Modeling with Quantized Structure and Disentangled Attention
abstract
Protein language models (PLMs) have shown remarkable capabilities in various protein function prediction tasks. However, while protein function is intricately tied to structure, most existing PLMs do not incorporate protein structure information. To address this issue, we introduce ProSST, a Transformer-based protein language model that seamlessly integrates both protein sequences and structures. ProSST incorporates a structure quantization module and a Transformer architecture with disentangled attention. The structure quantization module translates a 3D protein structure into a sequence of discrete tokens by first serializing the protein structure into residue-level local structures and then embeds them into dense vector space. These vectors are then quantized into discrete structure tokens by a pre-trained clustering model. These tokens serve as an effective protein structure representation. Furthermore, ProSST explicitly learns the relationship between protein residue token sequences and structure token sequences through the sequence-structure disentangled attention. We pre-train ProSST on millions of protein structures using a masked language model objective, enabling it to learn comprehensive contextual representations of proteins. To evaluate the proposed ProSST, we conduct extensive experiments on the zero-shot mutation effect prediction and several supervised downstream tasks, where ProSST achieves the state-of-the-art performance among all baselines. Our code and pre-trained models are publicly available.
Yang Tan 0001, Xinzhu Ma, Bozitao Zhong, Huiqun Yu, Ziyi Zhou 0002, Wanli Ouyang, Bingxin Zhou, Pan Tan
NeurIPS2