Sisi Yuan

dblp:146/9312 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0005-7557-9453ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A systematic review of molecular representation learning foundation models
abstract
Molecular representation learning (MRL) is afoundation in leveraging computational methods for drug discovery, enabling the transformation of molecular structure and properties into numerical vectors. These vectors serve as input for machine learning models and facilitate the prediction and analysis of molecular attributes, functions, and reactions. The advent of foundation models has introduced both new opportunities and challenges to MRL. These models have improved generalizability and migration in scarce data. Through pretraining and fine-tuning, foundation models can be adapted to various domains. Their robust encoding and generative abilities also allow the transformation of molecular data into more expressive forms. This paper provides a detailed review of current mainstream molecular descriptors and datasets, focusing primarily on the representation of small molecules while excluding larger molecules such as proteins and peptides. It classifies foundation models into two primary categories based on the form of input: unimodal-based and multimodal-based models. For each category, representative models are identified and their advantages and disadvantages evaluated. Moreover, we systematically summarize four core pretraining strategies for MRL foundation models, analyzing their task designs, applicable scenarios, and impacts on downstream performance. In addition, the application of molecular representation foundation models in drug discovery and development is discussed, together with the current status of model interpretability. The paper concludes with insights into the future directions of MRL foundation models.
Bosheng Song, Yuansheng Liu, Sisi Yuan, Xia Zhen
Briefings Bioinform.6
2026 Empowering chemical structures with biological insights for scalable phenotypic virtual screening
abstract
MOTIVATION: The scalable identification of bioactive compounds is essential for contemporary drug discovery. This process faces a key trade-off: structural screening offers scalability but lacks biological context, whereas high-content phenotypic profiling provides deep biological insights but is resource-intensive. The primary challenge is to extract robust biological signals from noisy data and encode them into representations that do not require biological data at inference. RESULTS: This study presents DECODE (DEcomposing Cellular Observations of Drug Effects), a framework that bridges this gap by empowering chemical representations with intrinsic biological semantics to enable structure-based in silico biological profiling. DECODE leverages limited paired transcriptomic and morphological data as supervisory signals during training, enabling the extraction of a measurement-invariant biological fingerprint from chemical structures and explicit filtering of modality-specific variation. Across held-out retrieval, scaffold-split and UMAP-clustering virtual-screening benchmarks, DECODE improves functional retrieval and early active-compound prioritization over baselines. AVAILABILITY AND IMPLEMENTATION: The codes and datasets of DECODE are available at https://github.com/lian-xiao/DECODE.
Xiaoqing Lian, Pengsen Ma, Tengfeng Ma, Zhong-Hao Ren, Xibao Cai, Zhixiang Cheng, Bosheng Song, Sisi Yuan, Chen Lin 0001
Bioinform.11
2025 GRAPE: graph-regularized protein language modeling unlocks TCR-epitope binding specificity
abstract
T-cell receptor (TCR)-epitope binding prediction is critical for immunotherapies but remains challenged by sparse interaction networks and severe class imbalance in training data. Current graph neural network (GNN) approaches for predicting TCR-epitope binding (TEB) fail to address two key limitations: over-smoothing during message propagation in sparse TCR-epitope graphs and biased predictions toward dominant epitope-TCR pairs. Here, we present GRAPE (Graph-Regularized Attentive Protein Embeddings), a framework unifying spectral graph regularization and imbalance-aware learning. GRAPE first leverages protein language models (ESM-2) to generate evolutionary-informed TCR/epitope embeddings, constructing a topology-aware interaction graph. To mitigate over-smoothing, we introduce spectral graph regularization, explicitly constraining node feature smoothness to preserve discriminative patterns in sparse neighborhoods. Simultaneously, a dynamic edge reweighting module prioritizes unobserved TCR-epitope edges during graph propagation, coupled with a differentiable area under the ROC curve-maximization objective that directly optimizes for imbalance resilience. Extensive benchmarking on public datasets demonstrates that GRAPE significantly outperforms state-of-the-art methods in TEB prediction. This work establishes GRAPE as a robust framework for elucidating TCR-epitope interactions, with broad applications in immunology research and therapeutic design.
Xiangzheng Fu, Mingqiang Rong, Dong-Sheng Cao 0001, Sisi Yuan, Aiping Lu
Briefings Bioinform.7
2025 An image-based protein-ligand binding representation learning framework via multi-level flexible dynamics trajectory pre-training
abstract
MOTIVATION: Accurate prediction of protein-ligand binding (PLB) relationships plays a crucial role in drug discovery, which helps identify drugs that modulate the activity of specific targets. Traditional biological assays for measuring PLB relationships are time consuming and costly. In addition, models for predicting PLB relationships have been developed and widely used in drug discovery tasks. However, learning more accurate PLB representations is essential to meet the stringent standards required for drug discovery. RESULTS: We propose an image-based PLB representation learning framework, called ImagePLB, which equips ligand representation learner (LRL) and protein representation learner (PRL) to accept 3D multi-view ligand images and protein graphs as input, respectively, and learns rich interaction information between ligand and protein through a binding representation learner (BRL). Considering the scarcity of protein-ligand pairs, we further propose a multi-level next trajectory prediction (MLNTP) task to pre-train ImagePLB on the 4D flexible dynamics trajectory of 16 972 complexes, including ligand level, protein level, and complex level, to learn information related to trajectories. Besides, by introducing trajectory regularization (TR), we effectively alleviate the problem of high (even almost identical) feature similarity caused by adjacent trajectories. Compared with the current state-of-the-art methods, ImagePLB has achieved competitive improvements on PLB-related prediction tasks, including protein-ligand affinity and efficacy prediction tasks. This study opens the door to the image-based PLB learning paradigm. AVAILABILITY AND IMPLEMENTATION: All data and implementation details of code can be obtained from https://github.com/HongxinXiang/ImagePLB.
Hongxin Xiang, Mingquan Liu, Linlin Hou, Shuting Jin, Jianmin Wang 0016, Jun Xia 0001, Wenjie Du 0003, Sisi Yuan, Xiangzheng Fu, Lei Xu 0047
Bioinform.8
2025 Semantic substructure guided multiple objective molecular generation with discrete diffusion probabilistic model
Shugao Chen, Bosheng Song, Ruizhe Chen, Guifei Zhou, Sisi Yuan
Neurocomputing6
2025 Enhancing Herbal Medicine-Drug Interaction Prediction Using Large Language Models
abstract
Investigating potential interactions between drugs and herbal medicines helps optimize combined treatment strategies and supports personalized and precision medicine. Deep learning-based methods have been successful in predicting drug-related interactions. However, these methods face challenges such as low data quality and uneven distribution. Large language models (LLMs) effectively address these challenges through their extensive knowledge bases. Motivated by this, we integrate LLMs, one-hot encoding, and variational graph autoencoders (VGAEs) to propose a herbal medicine-drug interaction (HDI) prediction model. First, LLMs are employed to extract features from drug SMILES, generating high-quality molecular representations. Second, one-hot encoding is applied to herbal medicines with multiple natural products to construct feature vectors and improve model interpretability. Finally, VGAEs are utilized to reconstruct herbal medicine-drug graphs and predict unknown HDIs. Additionally, we differentiate between herbal medicine-drug similarity and the degree of individual drug or herbal medicine nodes to mitigate the dominance of high-degree nodes in VGAE message flow. Multiple experiments were conducted to validate the significance of the proposed model and its key components. This method shows great potential for applications in traditional Chinese medicine formulation optimization, new drug development, and precision medicine.
Sisi Yuan, Zhecheng Zhou, Xinyuan Jin, Linlin Zhuo, Keqin Li 0001
IEEE J. Biomed. Health Informatics1
2015 Identification of cytokine via an improved genetic algorithm
Xiangxiang Zeng, Sisi Yuan, Xianxian Huang, Quan Zou 0001
Frontiers Comput. Sci.2