Yahan Li

dblp:251/9311 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hierarchical Structure-Property Alignment for Data-Efficient Molecular Generation and Editing
abstract
Property-constrained molecular generation and editing are crucial in AI-driven drug discovery but remain hindered by two factors: (i) capturing the complex relationships between molecular structures and multiple properties remains challenging, and (ii) the narrow coverage and incomplete annotations of molecular properties weaken the effectiveness of property-based models. To tackle these limitations, we propose HSPAG, a data-efficient framework featuring hierarchical structure–property alignment. By treating SMILES and molecular properties as complementary modalities, the model learns their relationships at atom, substructure, and whole-molecule levels. Moreover, we select representative samples through scaffold clustering and hard samples via an auxiliary variational auto-encoder (VAE), substantially reducing the required pre-training data. In addition, we incorporate a property relevance-aware masking mechanism and diversified perturbation strategies to enhance generation quality under sparse annotations. Experiments demonstrate that HSPAG captures fine-grained structure–property relationships and supports controllable generation under multiple property constraints. Two real-world case studies further validate the editing capabilities of HSPAG.
Ziyu Fan, Zhijian Huang 0001, Yahan Li, Yunliang Wang, Zeyu Zhong, Shuhong Liu, Shuning Yang, Shangqian Wu, Min Wu 0008, Lei Deng 0002
AAAI3
2026 A multi-objective multi-stage genetic algorithm for community detection in biological networks
Mingyuan Bi, Junliang Shang, Yahan Li, Feng Li 0033, Jin-Xing Liu 0001
Future Gener. Comput. Syst.3
2025 ASRSMA: Atomic-Scale and Structure-Based Modeling for RNA-Small Molecule Binding Affinity Prediction via Contrastive Pretraining
abstract
RNA is intricately involved in aberrant cellular functions and a wide range of disease processes, playing pivotal roles in gene regulation, viral replication, and innate immunity. Consequently, it has emerged as a highly promising therapeutic target. To accelerate the discovery of such drugs, it is essential to develop an effective computational method for predicting RNA-small molecule affinity. Therefore, we propose ASRSMA as an atomiclevel structure-aware model with contrastive pre-training for RNA-small molecule binding affinity prediction. ASRSMA represents RNA and small molecules at atomic resolution, enabling the capture of fine-grained structural features. It incorporates an atom-pair encoding module and an inter-molecular interaction module to model intra- and inter-molecular interactions. In addition, a self-supervised contrastive learning strategy is employed during pre-training, maximising the similarity between different views of the same complex while minimising similarities across different complexes, thereby yielding deep representations with strong generalisation capacity. Experimental results demonstrate that ASRSMA significantly outperforms state-of-the-art baseline models in predicting RNA-small molecule binding affinities.
Yahan Li, Zhijian Huang 0001, Yucheng Wang 0001, Min Wu 0008, Lei Deng 0002
BIBM1
2025 MVFDSP: A Multi-View Fusion Framework for Drug Side-Effect Frequency Prediction
abstract
Accurate prediction of drug side effect frequencies is critical for drug safety evaluation and clinical decision-making. Current methods primarily emphasize the associations between drugs and side effects, yet they often neglect the underlying structural and semantic features of both, which limits further advancements in prediction accuracy. In this study, we propose a novel multi-view fusion framework, MVFDSP, which integrates pre-trained molecular representation of 1D and 2D views with graph-based side effect information for side effect frequency prediction. Firstly, we obtain both 1D and 2D molecular representations from the pretrained molecular language model, and combine them using an adaptive fusion strategy. Subsequently, we construct a similarity network based on the side effect frequency matrix using K-Nearest Neighbors (KNN), and incorporate semantic embeddings derived from the terminology system of MedDRA to construct a side effect information graph. A multi-head graph attention network is then employed to capture the multi-dimensional information within this graph, allowing the model to attend to diverse aspects of the semantic and structural relationships among side effects. The final frequency prediction matrix is derived from the inner product between the learned drug and side effect embeddings. Experimental results on the SIDER 4.1 dataset demonstrate that MVFDSP outperforms existing methods, highlighting its effectiveness in capturing complex relationships of drugs and side effects. The code and data are available at https://github.com/Sonder-Echo/MVFDSP.
Zhengkang Wang, Zhijian Huang 0001, Yurong Qian, Yuanpeng Zhang 0004, Yahan Li, Qahtan Adnan Aljanabi, Jinmiao Song, Lei Deng 0002
BIBM5
2025 MVRBind: multi-view learning for RNA-small molecule binding site prediction
abstract
RNA plays a critical role in cellular processes, and its dysregulation is linked to many diseases, positioning RNA-targeted drugs as an important area of research. Accurate prediction of RNA-small molecule binding sites is crucial for advancing RNA-targeted therapies. Although deep learning has shown promise in this area, challenges remain in integrating and processing multi-dimensional data, such as RNA sequences and structural features, particularly given the inherent flexibility of RNA structures. In this study, we present MVRBind, a multi-view graph convolutional network designed to predict RNA-small molecule binding sites. MVRBind generates feature representations of RNA nucleotides across different structural levels. To effectively integrate these features, we developed a multi-view feature fusion module that constructs graphs based on RNA's primary, secondary, and tertiary structural views, enabling the model to capture diverse aspects of RNA structure. In addition, we fuse embeddings from multi-scale to obtain a comprehensive representation of RNA nucleotides, which is then used to predict RNA-small molecule binding sites. Extensive experiments demonstrate that MVRBind consistently outperforms baseline methods in various experimental settings. Our MVRBind shows exceptional performance in predicting binding sites for both the holo and apo forms of RNA, even when RNA adopts multiple conformations. These results suggest that MVRBind offers a robust model for structure-based RNA analysis, contributing toward accurate prediction and analysis of RNA-small molecule binding sites. All datasets and resource codes are available at https://github.com/cschen-y/MVRBind.
Zhijian Huang 0001, Yucheng Wang 0001, Yahan Li, Yaw Sing Tan, Lei Deng 0002, Min Wu 0008
Briefings Bioinform.4
2025 Selective fine-tuning for large language models via matrix nuclear norm
Tingyu Xia, Yahan Li, Yuan Wu 0002, Yi Chang 0001
Inf. Process. Manag.2
2024 LSNSCDA: Unraveling CircRNA-Drug Sensitivity via Local Smoothing Graph Neural Network and Credible Negative Samples
abstract
This study investigates the role of circular RNAs (circRNAs) in drug sensitivity, with a focus on their potential to inform personalized medicine. While current methods for identifying circRNA-drug sensitivity associations are resource-intensive, we propose LSNSCDA, a novel prediction algorithm that integrates Local Smoothing Graph Neural Networks (LS-GNN) and Credible Negative Sampling (CNS) to improve prediction accuracy. Our approach overcomes the challenges of fixed-length propagation in graph neural networks and the unreliability of randomly sampled negative instances. Experimental results show that LSNSCDA outperforms existing models, providing more reliable predictions and valuable insights into cancer treatment. Extensive evaluation confirms the effectiveness of each component of our model, while case studies further demonstrate its practical applicability. The source code and dataset are available at https://github.com/ZiyuFanCSU/LSNSCDA.
Ziyu Fan, Yuanpeng Zhang 0004, Yahan Li, Zeyu Zhong, Lei Deng 0002
BIBM3
2024 A Particle Swarm Optimization Algorithm Based on Multi-Population Mutual Learning for SNP-SNP Interaction Detection
abstract
Single nucleotide polymorphism (SNPs) data have become abundant thanks to the quick advancement of high-throughput sequencing technology, which provides convenience for genome-wide association studies. Single SNPs have been proven to be the cause of some diseases, and the emergence of complex diseases is often thought to be the result of the interaction of multiple SNPs. However, the possible interaction of millions of SNPs imposes a heavy computational burden for uncovering complex disease mechanisms. The existing SNP-SNP interaction detection algorithms frequently have flaws including high computation complexity and poor optimization effectiveness. In this study, a particle swarm optimization algorithm based on multi-population mutual learning (PSOMPML) is proposed to detect SNP-SNP interactions. In this algorithm, the mutual learning strategy is introduced to deal with different particles in different sub-populations to facilitate knowledge exchange. In addition, the elite preservation mechanism is incorporated into PSOMPML, to better preserve the good SNPs in the elite particles. The promising region local search strategy searches the optimal solution along the target solution and its near space to increase the convergence speed of the proposed algorithm. Experiments on simulated data sets and real data also demonstrate the effectiveness of the proposed algorithm.
Linqian Zhao, Yahan Li, Junliang Shang, Qianqian Ren, Yuanyuan Zhang 0008, Jin-Xing Liu 0001
BIBM2
2024 CPSORCL: A Cooperative Particle Swarm Optimization Method with Random Contrastive Learning for Interactive Feature Selection
Junliang Shang, Yahan Li, Feng Li 0033, Yuanyuan Zhang 0008, Jin-Xing Liu 0001
ISBRA (2)2
2023 idenLD-AREL: identifying lncRNA-disease associations by random forests based on an ensemble learning framework
abstract
Identification of disease-associated long non-coding RNAs (lncRNAs) facilitates the understanding of the pathogenesis of complex diseases. Many different types of computational models have been proposed. Although some of them have achieved encouraging results in predicting disease-associated lncRNAs, how to obtain stable results is still a challenge. In this paper, we propose a computational model based on an ensemble learning framework via the adaptive random forests, in short, idenLD-AREL. The idenLD-AREL integrates multiple random forest predictors and adaptive strategies to predict the scores of potential lncRNA-disease associations (LDAs), which ensure the stability and accuracy of the prediction results. In addition, there are a large number of false negative samples in the association datasets. For this reason, the resampling strategy is applied to idenLD-AREL to balance the samples. The idenLD-AREL is assessed by five-fold cross-validation in both the benchmark dataset and independent test set, showing excellent performance. Besides, the experimental results of the case study further demonstrate the effectiveness of the idenLD-AREL in predicting potential LDAs. The demo codes of the iLncDA-RSN are available online at https://github.com/CDMBlab/idenLD-AREL.
Yahan Li, Junliang Shang, Feng Li 0033, Jin-Xing Liu 0001
BIBM1
2023 ABCAE: Artificial Bee Colony Algorithm with Adaptive Exploitation for Epistatic Interaction Detection
Qianqian Ren, Yahan Li, Feng Li 0033, Jin-Xing Liu 0001, Junliang Shang
ISBRA2
2019 Modeling allele-specific expression at the gene and SNP levels simultaneously by a Bayesian logistic mixed regression model
abstract
BACKGROUND: High-throughput sequencing experiments, which can determine allele origins, have been used to assess genome-wide allele-specific expression. Despite the amount of data generated from high-throughput experiments, statistical methods are often too simplistic to understand the complexity of gene expression. Specifically, existing methods do not test allele-specific expression (ASE) of a gene as a whole and variation in ASE within a gene across exons separately and simultaneously. RESULTS: We propose a generalized linear mixed model to close these gaps, incorporating variations due to genes, single nucleotide polymorphisms (SNPs), and biological replicates. To improve reliability of statistical inferences, we assign priors on each effect in the model so that information is shared across genes in the entire genome. We utilize Bayesian model selection to test the hypothesis of ASE for each gene and variations across SNPs within a gene. We apply our method to four tissue types in a bovine study to de novo detect ASE genes in the bovine genome, and uncover intriguing predictions of regulatory ASEs across gene exons and across tissue types. We compared our method to competing approaches through simulation studies that mimicked the real datasets. The R package, BLMRM, that implements our proposed algorithm, is publicly available for download at https://github.com/JingXieMIZZOU/BLMRM . CONCLUSIONS: We will show that the proposed method exhibits improved control of the false discovery rate and improved power over existing methods when SNP variation and biological variation are present. Besides, our method also maintains low computational requirements that allows for whole genome analysis.
Jing Xie 0008, Tieming Ji, Marco A. R. Ferreira, Yahan Li, Bhaumik N. Patel, Rocio M. Rivera
BMC Bioinform.4