Yifan Wu 0008

dblp:25/7019-8 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 14 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 NumMolFormer: an explicit functional group number-guided framework for structure-based drug design
abstract
MOTIVATION: Rational molecule generation that balances binding affinity with favorable physicochemical properties remains a formidable challenge in structure-based drug design. The number of functional groups is a key determinant, as over-functionalization compromises physicochemical properties, whereas under-functionalization reduces binding affinity. However, current methods are limited in their capacity to incorporate the constraint. RESULTS: To address this, we present NumMolFormer, a Transformer-based framework designed to explicitly model functional group numbers. NumMolFormer adopts a dual-sequence input strategy, integrated with a numerical embedding module and a dual-stream differential attention mechanism, allowing molecular structures and functional group numbers to be encoded separately. This formulation alleviates the inherent limitations of standard Transformer in handling numerical information. In addition, we construct a large-scale dataset of 18 million molecules with functional group annotations for molecular pre-training, and further fine-tune the model using a combination of self-supervised learning and reinforcement learning under protein pocket constraints. The results demonstrate that NumMolFormer effectively leverages functional group information to generate molecules with improved binding affinity, synthetic accessibility, and drug-likeness compared to baseline methods. AVAILABILITY AND IMPLEMENTATION: The source code and datasets are available at http://www.github.com/zengzhicun/nummolformer.
Zhicun Zeng, Yifan Wu 0008, Zhangli Lu, Min Li 0007
Bioinform.2
2025 RNA3D-SSCL: Improving RNA Tertiary Structure Prediction via a Secondary Structure-Constrained Loss Function
abstract
The tertiary structure of RNA plays a crucial role in determining its biological functions, stability, and interactions with other molecules. Accurate prediction of RNA tertiary structure is essential for understanding RNA's functional roles in cellular processes. Although the accuracy of RNA secondary structure prediction is currently considered acceptable, using these predictions as explicit geometric constraints in tertiary structure modeling remains challenges. In the study, we propose RNA3D-SSCL, an end-to-end deep learning framework for RNA tertiary structure prediction. RNA3D-SSCL leverages deep learning techniques to predict the three-dimensional folding of RNAs, incorporating a novel loss function that integrates secondary structure constraints to improve prediction performance. By utilizing a combination of sequence and secondary structure features, RNA3D-SSCL is capable of generating more accurate RNA tertiary structures. RNA3D-SSCL was evaluated on an independent test set, revealing significant improvements compared to existing methods. Ablation studies confirm that the secondary structure-constrained loss function led to a notable reduction in RMSD and an improvement in TM-score, indicating higher prediction performance. The source code can be obtained at https://github.com/CSUBioGroup/RNA3D-SSCL.
Jingwei Lu, Yifan Wu 0008, Qianpei Liu, Yang Gao 0030, Min Zeng 0004
BIBM2
2025 Enhancing ICD classification with semantic embedding rectification and long-tail refinement
Yuhao Wu 0002, Yifan Wu 0008, Wei Fan 0010, Min Li 0007
Knowl. Based Syst.2
2023 DILM-ICD: A Deep Iterative Learning Model for Automatic ICD Coding
abstract
Automatic International Classification of Disease (ICD) coding plays a crucial role in assigning ICD codes to electronic medical records. This task presents a challenging multi-label text classification problem due to the vast number of ICD codes and the imbalanced label distribution. However, accurately predicting all labels simultaneously is extremely difficult for such the large label space. In this paper, we propose a novel model called Deep Iterative Learning Model (DILM-ICD), which uses an iterative learning framework to perform automatic ICD coding task. The iterative learning framework can refine the prediction results by repeating the iteration modules, which simulates the human-like coding process. In addition, we propose a multi-head text-label matching mechanism, which combines the embedding ICD description information to better match the relationship between text and label. The combination of the iterative learning framework with the multi-head text-label matching mechanism enables the model pay attention to lowfrequency ICD codes. DILM-ICD is evaluated on the MIMIC-III-full dataset and MIMIC-III-50 dataset. The experimental results show that DILM-ICD achieves state-of-the-art results across multiple evaluation metrics, which demonstrates the effectiveness of our proposed model.
Weiyan Qiu, Yifan Wu 0008, Kunying Niu, Min Zeng 0004, Min Li 0007
BIBM2
2023 Singularformer: Learning to Decompose Self-Attention to Linearize the Complexity of Transformer
abstract
Transformers achieve excellent performance in a variety of domains since they can capture long-distance dependencies through the self-attention mechanism. However, self-attention is computationally costly due to its quadratic complexity and high memory consumption. In this paper, we propose a novel Transformer variant (Singularformer) that uses neural networks to learn the singular value decomposition process of the attention matrix to design a linear-complexity and memory-efficient global self-attention mechanism. Specifically, we decompose the attention matrix into the product of three matrix factors based on singular value decomposition and design neural networks to learn these matrix factors, then the associative law of matrix multiplication is used to linearize the calculation of self-attention. The above procedure allows us to compute self-attention as two-dimensional reduction processes in the first and second token dimensional spaces, followed by a multi-head self-attention computational process on the first dimensional reduced token features. Experimental results on 8 real-world datasets demonstrate that Singularformer performs favorably against the other Transformer variants with lower time and space complexity. Our source code is publicly available at https://github.com/CSUBioGroup/Singularformer.
Yifan Wu 0008, Shichao Kan, Min Zeng 0004, Min Li 0007
IJCAI1
2023 RLBind: a deep learning method to predict RNA-ligand binding sites
abstract
Identification of RNA-small molecule binding sites plays an essential role in RNA-targeted drug discovery and development. These small molecules are expected to be leading compounds to guide the development of new types of RNA-targeted therapeutics compared with regular therapeutics targeting proteins. RNAs can provide many potential drug targets with diverse structures and functions. However, up to now, only a few methods have been proposed. Predicting RNA-small molecule binding sites still remains a big challenge. New computational model is required to better extract the features and predict RNA-small molecule binding sites more accurately. In this paper, a deep learning model, RLBind, was proposed to predict RNA-small molecule binding sites from sequence-dependent and structure-dependent properties by combining global RNA sequence channel and local neighbor nucleotides channel. To our best knowledge, this research was the first to develop a convolutional neural network for RNA-small molecule binding sites prediction. Furthermore, RLBind also can be used as a potential tool when the RNA experimental tertiary structure is not available. The experimental results show that RLBind outperforms other state-of-the-art methods in predicting binding sites. Therefore, our study demonstrates that the combination of global information for full-length sequences and local information for limited local neighbor nucleotides in RNAs can improve the model's predictive performance for binding sites prediction. All datasets and resource codes are available at https://github.com/KailiWang1/RLBind.
Renyi Zhou, Yifan Wu 0008, Min Li 0007
Briefings Bioinform.3
2023 LncLocFormer: a Transformer-based deep learning model for multi-label lncRNA subcellular localization prediction by using localization-specific attention mechanism
abstract
MOTIVATION: There is mounting evidence that the subcellular localization of lncRNAs can provide valuable insights into their biological functions. In the real world of transcriptomes, lncRNAs are usually localized in multiple subcellular localizations. Furthermore, lncRNAs have specific localization patterns for different subcellular localizations. Although several computational methods have been developed to predict the subcellular localization of lncRNAs, few of them are designed for lncRNAs that have multiple subcellular localizations, and none of them take motif specificity into consideration. RESULTS: In this study, we proposed a novel deep learning model, called LncLocFormer, which uses only lncRNA sequences to predict multi-label lncRNA subcellular localization. LncLocFormer utilizes eight Transformer blocks to model long-range dependencies within the lncRNA sequence and shares information across the lncRNA sequence. To exploit the relationship between different subcellular localizations and find distinct localization patterns for different subcellular localizations, LncLocFormer employs a localization-specific attention mechanism. The results demonstrate that LncLocFormer outperforms existing state-of-the-art predictors on the hold-out test set. Furthermore, we conducted a motif analysis and found LncLocFormer can capture known motifs. Ablation studies confirmed the contribution of the localization-specific attention mechanism in improving the prediction performance. AVAILABILITY AND IMPLEMENTATION: The LncLocFormer web server is available at http://csuligroup.com:9000/LncLocFormer. The source code can be obtained from https://github.com/CSUBioGroup/LncLocFormer.
Min Zeng 0004, Yifan Wu 0008, Rui Yin 0002, Chengqian Lu, Junwen Duan, Min Li 0007
Bioinform.2
2023 Retrieve and rerank for automated ICD coding via Contrastive Learning
Kunying Niu, Yifan Wu 0008, Yaohang Li, Min Li 0007
J. Biomed. Informatics2
2022 MERAS: A method for entity recognition, alignment, and structuration from Chinese Electronic Medical Records
abstract
The description of entity information in Chinese Electronic Medical Records (EMRs) is not always isolated, and it often has rich attribute constraints, especially the symptom entities. Moreover, these entity descriptions are unstructured and may be colloquial, non-terminological, and misspelled. This paper proposed a new method, which is called MERAS for entity recognition, alignment, and structuration from Chinese EMRs. Different from the traditional entity extraction tasks, a series of complex entities with attribute constraints are annotated and recognized by MERAS, and then the entities are further converted to the standardized and structured entity sequences in JSON format based on the end-to-end translation idea. BiLSTM-CRF is applied to recognize the complex entities, and Transformer is applied to realize the entity alignment and structuration in this paper. With the manually annotated data set as the reference, 9 types of entities are extracted from the EMRs, and the corresponding corpus-feature-based enhanced strategies are proposed to support the model training. The experimental results show that MERAS can effectively extract standardized and structured entities from Chinese EMRs, and it can have great potential in practical applications. (Abstract)
Yifan Wu 0008, Min Li 0007
BIBM2
2022 DeepLncLoc: a deep learning framework for long non-coding RNA subcellular localization prediction based on subsequence embedding
abstract
Long non-coding RNAs (lncRNAs) are a class of RNA molecules with more than 200 nucleotides. A growing amount of evidence reveals that subcellular localization of lncRNAs can provide valuable insights into their biological functions. Existing computational methods for predicting lncRNA subcellular localization use k-mer features to encode lncRNA sequences. However, the sequence order information is lost by using only k-mer features. We proposed a deep learning framework, DeepLncLoc, to predict lncRNA subcellular localization. In DeepLncLoc, we introduced a new subsequence embedding method that keeps the order information of lncRNA sequences. The subsequence embedding method first divides a sequence into some consecutive subsequences and then extracts the patterns of each subsequence, last combines these patterns to obtain a complete representation of the lncRNA sequence. After that, a text convolutional neural network is employed to learn high-level features and perform the prediction task. Compared with traditional machine learning models, popular representation methods and existing predictors, DeepLncLoc achieved better performance, which shows that DeepLncLoc could effectively predict lncRNA subcellular localization. Our study not only presented a novel computational model for predicting lncRNA subcellular localization but also introduced a new subsequence embedding method which is expected to be applied in other sequence-based prediction tasks. The DeepLncLoc web server is freely accessible at http://bioinformatics.csu.edu.cn/DeepLncLoc/, and source code and datasets can be downloaded from https://github.com/CSUBioGroup/DeepLncLoc.
Min Zeng 0004, Yifan Wu 0008, Chengqian Lu, Fuhao Zhang, Fang-Xiang Wu, Min Li 0007
Briefings Bioinform.2
2022 BACPI: a bi-directional attention neural network for compound-protein interaction and binding affinity prediction
abstract
MOTIVATION: The identification of compound-protein interactions (CPIs) is an essential step in the process of drug discovery. The experimental determination of CPIs is known for a large amount of funds and time it consumes. Computational model has therefore become a promising and efficient alternative for predicting novel interactions between compounds and proteins on a large scale. Most supervised machine learning prediction models are approached as a binary classification problem, which aim to predict whether there is an interaction between the compound and the protein or not. However, CPI is not a simple binary on-off relationship, but a continuous value reflects how tightly the compound binds to a particular target protein, also called binding affinity. RESULTS: In this study, we propose an end-to-end neural network model, called BACPI, to predict CPI and binding affinity. We employ graph attention network and convolutional neural network (CNN) to learn the representations of compounds and proteins and develop a bi-directional attention neural network model to integrate the representations. To evaluate the performance of BACPI, we use three CPI datasets and four binding affinity datasets in our experiments. The results show that, when predicting CPIs, BACPI significantly outperforms other available machine learning methods on both balanced and unbalanced datasets. This suggests that the end-to-end neural network model that predicts CPIs directly from low-level representations is more robust than traditional machine learning-based methods. And when predicting binding affinities, BACPI achieves higher performance on large datasets compared to other state-of-the-art deep learning methods. This comparison result suggests that the proposed method with bi-directional attention neural network can capture the important regions of compounds and proteins for binding affinity prediction. AVAILABILITY AND IMPLEMENTATION: Data and source codes are available at https://github.com/CSUBioGroup/BACPI.
Min Li 0007, Zhangli Lu, Yifan Wu 0008, Yaohang Li
Bioinform.3
2022 BridgeDPI: a novel Graph Neural Network for predicting drug-protein interactions
abstract
MOTIVATION: Exploring drug-protein interactions (DPIs) provides a rapid and precise approach to assist in laboratory experiments for discovering new drugs. Network-based methods usually utilize a drug-protein association network and predict DPIs by the information of its associated proteins or drugs, called 'guilt-by-association' principle. However, the 'guilt-by-association' principle is not always true because sometimes similar proteins cannot interact with similar drugs. Recently, learning-based methods learn molecule properties underlying DPIs by utilizing existing databases of characterized interactions but neglect the network-level information. RESULTS: We propose a novel method, namely BridgeDPI. We devise a class of virtual nodes to bridge the gap between drugs and proteins and construct a learnable drug-protein association network. The network is optimized based on the supervised signals from the downstream task-the DPI prediction. Through information passing on this drug-protein association network, a Graph Neural Network can capture the network-level information among diverse drugs and proteins. By combining the network-level information and the learning-based method, BridgeDPI achieves significant improvement in three real-world DPI datasets. Moreover, the case study further verifies the effectiveness and reliability of BridgeDPI. AVAILABILITY AND IMPLEMENTATION: The source code of BridgeDPI can be accessed at https://github.com/SenseTime-Knowledge-Mining/BridgeDPI. The source data used in this study is available on the https://github.com/IBM/InterpretableDTIP (for the BindingDB dataset), https://github.com/masashitsubaki/CPI_prediction (for the C.ELEGANS and HUMAN) datasets, http://dude.docking.org/ (for the DUD-E dataset), repectively.
Yifan Wu 0008, Min Zeng 0004, Jie Zhang 0122, Min Li 0007
Bioinform.1
2022 KAICD: A knowledge attention-based deep learning framework for automatic ICD coding
Yifan Wu 0008, Min Zeng 0004, Zhihui Fei, Fang-Xiang Wu, Min Li 0007
Neurocomputing1
2022 Accurate Prediction of Human Essential Proteins Using Ensemble Deep Learning
abstract
Essential proteins are considered the foundation of life as they are indispensable for the survival of living organisms. Computational methods for essential protein discovery provide a fast way to identify essential proteins. But most of them heavily rely on various biological information, especially protein-protein interaction networks, which limits their practical applications. With the rapid development of high-throughput sequencing technology, sequencing data has become the most accessible biological data. However, using only protein sequence information to predict essential proteins has limited accuracy. In this paper, we propose EP-EDL, an ensemble deep learning model using only protein sequence information to predict human essential proteins. EP-EDL integrates multiple classifiers to alleviate the class imbalance problem and to improve prediction accuracy and robustness. In each base classifier, we employ multi-scale text convolutional neural networks to extract useful features from protein sequence feature matrices with evolutionary information. Our computational results show that EP-EDL outperforms the state-of-the-art sequence-based methods. Furthermore, EP-EDL provides a more practical and flexible way for biologists to accurately predict essential proteins. The source code and datasets can be downloaded from https://github.com/CSUBioGroup/EP-EDL.
Min Zeng 0004, Yifan Wu 0008, Yaohang Li, Min Li 0007
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 A Pseudo Label-Wise Attention Network for Automatic ICD Coding
abstract
Automatic International Classification of Diseases (ICD) coding is defined as a kind of text multi-label classification problem, which is difficult because the number of labels is very large and the distribution of labels is unbalanced. The label-wise attention mechanism is widely used in automatic ICD coding because it can assign weights to every word in full Electronic Medical Records (EMR) for different ICD codes. However, the label-wise attention mechanism is redundant and costly in computing. In this paper, we propose a pseudo label-wise attention mechanism to tackle the problem. Instead of computing different attention modes for different ICD codes, the pseudo label-wise attention mechanism automatically merges similar ICD codes and computes only one attention mode for the similar ICD codes, which greatly compresses the number of attention modes and improves the predicted accuracy. In addition, we apply a more convenient and effective way to obtain the ICD vectors, and thus our model can predict new ICD codes by calculating the similarities between EMR vectors and ICD vectors. Our model demonstrates effectiveness in extensive computational experiments. On the public MIMIC-III dataset and private Xiangya dataset, our model achieves the best performance on micro F1 (0.583 and 0.806), micro AUC (0.986 and 0.994), P@8 (0.756 and 0.413), and costs much smaller GPU memory (about 26.1% of the models with label-wise attention). Furthermore, we verify the ability of our model in predicting new ICD codes. The interpretablility analysis and case study show the effectiveness and reliability of the patterns obtained by the pseudo label-wise attention mechanism.
Yifan Wu 0008, Min Zeng 0004, Yaohang Li, Min Li 0007
IEEE J. Biomed. Health Informatics1
2021 A Hybrid Pooling Based Deep Learning Framework For Automated ICD Coding
abstract
ICD coding is the practice of allocating diagnostic and procedure codes to the clinical records following the International Classification of Diseases. The manual allocation of ICD codes to clinical notes is a very tedious job which has become costly, time-consuming and error-prone. Up to now, various methods for automated ICD coding have been devised, ranging from machine learning to deep learning methodologies. Earlier cutting-edge models relied on CNN’s with one or several fixed window widths. However, the length and dependency of text fragments linked to ICD labels in clinical literature differ considerably, posing a difficulty in determining the optimal window size. Apart from that, in prior models that utilized CNN architecture, the average features of the clinical notes have been ignored, resulting in a lack of preparation of rich features for the classifier. In this research, we present a Deep Recurrent Convolutional Neural Network with Hybrid Pooling (DRCNN-HP), which addresses all of the above mentioned issues. DRCNN-HP takes into account the different lengths as well as the dependency of the ICD code-related text chunks. Furthermore, we applied a powerful hybrid pooling layer in DRCNN-HP to capture the rich feature representation (i.e., by concatenating maximum and average features of text) for the classifier, which resulted in giving state-of-the-art results on the MIMIC-III top 50 dataset as compared to the prior competitive models.
Sajida Raz Bhutto, Yifan Wu 0008, Akhtar Hussain 0001, Min Li 0007
BIBM2
2021 Improving human essential protein prediction using only protein sequences via ensemble learning
abstract
Accurate prediction of essential proteins by using computational methods can effectively reduce the cost of wet-lab experiments. Existing computational methods usually rely on constructed protein-protein interaction (PPI) networks with different kinds of biological data. However, high-quality PPI networks and other biological data are not available for all proteins. Thus, it is very necessary and valuable to develop accurate methods for fast and effective prediction of essential proteins by using only protein sequences. We propose EPGBDT, a machine learning ensemble model, to improve the performance of essential protein prediction by using only protein sequences. EP-GBDT has an ensemble structure that combines multiple Gradient Boosting Decision Tree (GBDT) base classifiers. In addition, to reduce the effects of imbalanced dataset, EP-GBDT uses a sampling technique. The results show that EP-GBDT outperforms state-of-the-art sequence-based methods and network-based centrality measures. The source code and datasets can be downloaded from https://github.com/CSUBioGroup/EP-GBDT.
Min Zeng 0004, Yifan Wu 0008, Fang-Xiang Wu, Min Li 0007
BIBM3