Yongxian Fan

dblp:270/7217 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0003-0120-8092ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 7 first-author · 14 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Protein-nucleic acid binding site prediction using interpretable Kolmogorov-Arnold networks with hypergraph representation learning
abstract
MOTIVATION: In recent years, protein language models (pLMs) and graph neural networks (GNNs) have demonstrated powerful expressive and reasoning capabilities in modeling protein-RNA/DNA interactions. However, existing methods, which use simple graphs to describe the relationships between residues, struggle to effectively capture the high-order, multi-body residue interactions present in protein-nucleic acid complex structures. In fact, spatially continuous but sequence-wise discontinuous residues often cooperatively determine nucleic acid binding capacity. RESULTS: In this study, we present IKANbind, a computational approach that combines hypergraph representation learning and interpretable Kolmogorov-Arnold Networks (KANs), for identifying nucleic acid binding residues (NBRs) in proteins. By combining the advantages of pLM, hypergraph neural networks and symbolic KAN, IKANbind outperforms existing methods on multiple NBR benchmark datasets. We also demonstrated that the pLM used in IKANbind can implicitly learn the physicochemical properties of binding residues, such as charge and hydrophobicity. In addition, the symbolic KAN, which uses a unique weighted mechanism of decomposable basis functions, can accurately identify the features with the greatest contribution to NBR recognition. We found that polarity and charge make greater contributions to NBR prediction than other physicochemical properties or evolutionary information. Finally, IKANbind achieves promising performance when extended to other ligand-binding residue prediction tasks. AVAILABILITY AND IMPLEMENTATION: IKANbind is freely available at https://github.com/yangfengzhuguet/IKANBind.
Yangfeng Zhu, Guicong Sun, Weimin Zhu, Yongxian Fan, Zeheng Wu, Xianchen Zheng, Xiaoyong Pan
Bioinform.4
2026 MSGSNet: sequence learning enhanced with multi-scale De Bruijn graphs for transcription factor binding site prediction
Yongxian Fan, Zeheng Wu, Yangfeng Zhu, Guicong Sun
Neurocomputing2
2026 FGAIM: Identifying Drug-Target Activation and Inhibition Mechanisms via Inductive Graph Neural Networks Based on Fine-Grained Interaction Strategies
abstract
Distinguishing the activation and inhibition mechanisms between drugs and targets can reveal the potential regulatory pathways of target functions, which is crucial in drug discovery and development. Although numerous deep learning-based computational methods have been proposed, most of them extract drug and target features independently while neglecting interactions between drug molecules and protein residues. In addition, existing methods approach drugs in a relatively simple way and predominantly rely on protein sequence information. Deep learning, particularly graph neural networks (GNNs), has demonstrated unique advantages in processing data with complex graph-structured relationships. Motivated by these limitations, this study proposes a novel computational method named FGAIM, which designed to effectively identify activation and inhibition mechanisms between drugs and targets. First, a multi-scale GNN module in FGAIM is employed to learn expressive drug molecular embeddings. At the same time, protein representations are constructed by integrating pre-trained language model (PLM) embeddings and structural information derived from 3D conformations. Subsequently, FGAIM leverages a GraphSAGE module to extract features from both the primary drug graph and the fine-grained drug-protein interaction graph. Finally, the resulting drug and target embeddings are fused and fed into a Multilayer Perceptron (MLP) for classification prediction. Comparative experiments on two public datasets demonstrate that FGAIM significantly outperforms existing computational approaches and exhibits strong generalization capabilities. Further case studies reveal the advantages of FGAIM in mining previously unrecognized activation/inhibition relationships. Building upon the fine-grained interactions, we further explored the model's interpretability at the molecular structural level through analysis of attention weights. Additionally, visualization analyses confirm that FGAIM effectively captures underlying structural patterns in the data and identifies discriminative features across different categories.
Yongxian Fan, Guicong Sun, Mengxin Zheng
IEEE Trans. Comput. Biol. Bioinform.2
2026 iDRKAN: Interpretable miRNA-Disease Association Prediction Based on Dual-Graph Representation Learning and Kolmogorov-Arnold Network
abstract
Accurately identifying miRNA-disease association (MDA) is of great importance in biomedical research and clinical applications. However, most existing computational methods rely on similarity, making it difficult to effectively capture the deep semantic information among heterogeneous nodes in complex networks. In addition, the inherent "closed box" nature of traditional deep learning models leads to a lack of transparency in their decision-making process. Therefore, we propose an interpretable MDA prediction method (iDRKAN) based on dual-graph representation learning and Kolmogorov-Arnold Network. First, iDRKAN constructs similarity views and meta-path views based on the similarity matrix and association matrix respectively, and learns higher-order feature representation of each view by graph convolutional network (GCN). Subsequently, the multi-channel attention (MCA) mechanism is introduced to adaptively fuse the contextual information of each similarity view in different convolutional layers, and the semantic layer attention (SLA) mechanism is used to integrate the semantic information contained in different meta-path views. Next, the contrastive learning strategy is used to optimize the consistency between the dual-graph representation. Finally, the dual-graph features are weighted fused, and fed into the interpretable Kolmogorov-Arnold Network (KAN) for prediction. Experimental results on two public datasets show that iDRKAN significantly outperforms existing computational approaches in multiple performance indicators. In addition, experiments with different classifiers verify that iDRKAN achieves a good balance between prediction performance and interpretability. The case study further demonstrates the effectiveness of iDRKAN in mining potential associations.
Yangfeng Zhu, Yongxian Fan, Guicong Sun
IEEE Trans. Comput. Biol. Bioinform.2
2025 EGCPPIS: learning hierarchical equivariant graph representations with contrastive integration for protein-protein interaction site identification
abstract
BACKGROUND: Protein-protein interactions regulate the dynamic operation of intracellular molecular networks, serving as the molecular basis for revealing protein functions and disease mechanisms. Recently, several computational methods for predicting protein-protein interaction sites (PPIs) have been presented as alternatives to costly and labor-intensive traditional experiments. However, existing methods generally ignore the inherent hierarchical structure of protein chains. Furthermore, the equivariance of graph structure during spatial transformations is often neglected when applying graph neural networks to modeling. Therefore, accurately identifying PPIs remains a challenging task. RESULTS: In this work, we propose an end-to-end GNN-based computational method, EGCPPIS, for efficiently identifying protein-protein interaction sites. First, we construct a hierarchical graph representation of the protein chain, including residue-level graph and atom-level graph. Next, EGCPPIS designs an E(n) Equivariant Graph Neural Network (EGNN) module to learn residue-level embeddings with equivariant features. After further extracting atom-level embeddings using the GraphSAGE module, we introduce the contrastive learning strategy to integrate hierarchical graph features. This strategy enables us to learn consistent embeddings between residue-level and atom-level representations. Finally, the fused embeddings are weighted using an improved gated multi-head attention mechanism. CONCLUSION: Comprehensive evaluation results on multiple datasets demonstrate that EGCPPIS significantly outperforms state-of-the-art methods. Extensive comparative experiments and case studies further confirm that EGCPPIS can reveal the decision-making patterns in PPIs prediction, facilitating the discovery of potential PPIs. The original datasets and code of EGCPPIS are available at https://github.com/GuicongSun/EGCPPIS .
Guicong Sun, Yongxian Fan, Yangfeng Zhu, Mengxin Zheng
BMC Bioinform.2
2025 MLC-DTA: Drug-target affinity prediction based on multi-level contrastive learning and equivariant graph neural networks
Mengxin Zheng, Guicong Sun, Yongxian Fan
Neurocomputing3
2024 EGPDI: identifying protein-DNA binding sites based on multi-view graph embedding fusion
abstract
Mechanisms of protein-DNA interactions are involved in a wide range of biological activities and processes. Accurately identifying binding sites between proteins and DNA is crucial for analyzing genetic material, exploring protein functions, and designing novel drugs. In recent years, several computational methods have been proposed as alternatives to time-consuming and expensive traditional experiments. However, accurately predicting protein-DNA binding sites still remains a challenge. Existing computational methods often rely on handcrafted features and a single-model architecture, leaving room for improvement. We propose a novel computational method, called EGPDI, based on multi-view graph embedding fusion. This approach involves the integration of Equivariant Graph Neural Networks (EGNN) and Graph Convolutional Networks II (GCNII), independently configured to profoundly mine the global and local node embedding representations. An advanced gated multi-head attention mechanism is subsequently employed to capture the attention weights of the dual embedding representations, thereby facilitating the integration of node features. Besides, extra node features from protein language models are introduced to provide more structural information. To our knowledge, this is the first time that multi-view graph embedding fusion has been applied to the task of protein-DNA binding site prediction. The results of five-fold cross-validation and independent testing demonstrate that EGPDI outperforms state-of-the-art methods. Further comparative experiments and case studies also verify the superiority and generalization ability of EGPDI.
Mengxin Zheng, Guicong Sun, Yongxian Fan
Briefings Bioinform.4
2024 iProL: identifying DNA promoters from sequence information based on Longformer pre-trained model
abstract
Promoters are essential elements of DNA sequence, usually located in the immediate region of the gene transcription start sites, and play a critical role in the regulation of gene transcription. Its importance in molecular biology and genetics has attracted the research interest of researchers, and it has become a consensus to seek a computational method to efficiently identify promoters. Still, existing methods suffer from imbalanced recognition capabilities for positive and negative samples, and their recognition effect can still be further improved. We conducted research on E. coli promoters and proposed a more advanced prediction model, iProL, based on the Longformer pre-trained model in the field of natural language processing. iProL does not rely on prior biological knowledge but simply uses promoter DNA sequences as plain text to identify promoters. It also combines one-dimensional convolutional neural networks and bidirectional long short-term memory to extract both local and global features. Experimental results show that iProL has a more balanced and superior performance than currently published methods. Additionally, we constructed a novel independent test set following the previous specification and compared iProL with three existing methods on this independent test set.
Binchao Peng, Guicong Sun, Yongxian Fan
BMC Bioinform.3
2024 An Improved COVID-19 Classification Model on Chest Radiography by Dual-Ended Multiple Attention Learning
abstract
As a highly contagious disease, COVID-19 has not only had a great impact on the life, study and work of hundreds of millions of people around the world, but also had a huge impact on the global health care system. Therefore, any technical tool that allows for rapid screening and high-precision diagnosis of COVID-19 infections can be of vital help. In order to reduce the burden on health care system, the computer-aided diagnosis of COVID-19 has become a current research hotspot. X-ray imaging is a common and low-cost tool that can help with the COVID-19 diagnosis. The data used for this study has 15,153 CXR images, containing 10,192 normal lungs, 3,631 COVID-19 positive cases and 1,345 images of viral pneumonia. For this computer-aided task, we propose the dual-ended multiple attention learning model (DMAL). The model incorporates multiple attention learning into both networks, and the two networks are linked using an integration module. Specifically, in both networks, the backbone network is used to extract global features and the branch network captures local area information; the integration module combines multi-stage features; and the attention module containing element, channel and spatial attention prompts the model to focus on multi-scale information relevant to the disease. We evaluate the proposed DMAL network using relevant competitive methods as well as ten advanced deep learning models in the image domain and obtain the best performance with 99.67%, 99.53%, 99.66%, 99.60% and 99.76% in terms of Accuracy, Precision, Sensitivity, F1 Scores and Specificity. The proposed method will help in the rapid screening and high-precision diagnosis of COVID-19, given the general trend of such severe global infections. Our code and model are available in [https://github.com/Graziagh/DMALNet].
Yongxian Fan, Hao Gong 0006
IEEE J. Biomed. Health Informatics1
2023 IHCP: interpretable hepatitis C prediction system based on black-box machine learning models
abstract
BACKGROUND: Hepatitis C is a prevalent disease that poses a high risk to the human liver. Early diagnosis of hepatitis C is crucial for treatment and prognosis. Therefore, developing an effective medical decision system is essential. In recent years, many computational methods have been proposed to identify hepatitis C patients. Although existing hepatitis prediction models have achieved good results in terms of accuracy, most of them are black-box models and cannot gain the trust of doctors and patients in clinical practice. As a result, this study aims to use various Machine Learning (ML) models to predict whether a patient has hepatitis C, while also using explainable models to elucidate the prediction process of the ML models, thus making the prediction process more transparent. RESULT: We conducted a study on the prediction of hepatitis C based on serological testing and provided comprehensive explanations for the prediction process. Throughout the experiment, we modeled the benchmark dataset, and evaluated model performance using fivefold cross-validation and independent testing experiments. After evaluating three types of black-box machine learning models, Random Forest (RF), Support Vector Machine (SVM), and AdaBoost, we adopted Bayesian-optimized RF as the classification algorithm. In terms of model interpretation, in addition to using common SHapley Additive exPlanations (SHAP) to provide global explanations for the model, we also utilized the Local Interpretable Model-Agnostic Explanations with stability (LIME_stabilitly) to provide local explanations for the model. CONCLUSION: Both the fivefold cross-validation and independent testing show that our proposed method significantly outperforms the state-of-the-art method. IHCP maintains excellent model interpretability while obtaining excellent predictive performance. This helps uncover potential predictive patterns of the model and enables clinicians to better understand the model's decision-making process.
Yongxian Fan, Xiqian Lu, Guicong Sun
BMC Bioinform.1
2023 DeepASDPred: a CNN-LSTM-based deep learning method for Autism spectrum disorders risk RNA identification
abstract
BACKGROUND: Autism spectrum disorders (ASD) are a group of neurodevelopmental disorders characterized by difficulty communicating with society and others, behavioral difficulties, and a brain that processes information differently than normal. Genetics has a strong impact on ASD associated with early onset and distinctive signs. Currently, all known ASD risk genes are able to encode proteins, and some de novo mutations disrupting protein-coding genes have been demonstrated to cause ASD. Next-generation sequencing technology enables high-throughput identification of ASD risk RNAs. However, these efforts are time-consuming and expensive, so an efficient computational model for ASD risk gene prediction is necessary. RESULTS: In this study, we propose DeepASDPerd, a predictor for ASD risk RNA based on deep learning. Firstly, we use K-mer to feature encode the RNA transcript sequences, and then fuse them with corresponding gene expression values to construct a feature matrix. After combining chi-square test and logistic regression to select the best feature subset, we input them into a binary classification prediction model constructed by convolutional neural network and long short-term memory for training and classification. The results of the tenfold cross-validation proved our method outperformed the state-of-the-art methods. Dataset and source code are available at https://github.com/Onebear-X/DeepASDPred is freely available. CONCLUSIONS: Our experimental results show that DeepASDPred has outstanding performance in identifying ASD risk RNA genes.
Yongxian Fan, Guicong Sun
BMC Bioinform.1
2023 ELMo4m6A: A Contextual Language Embedding-Based Predictor for Detecting RNA N6-Methyladenosine Sites
abstract
N6-methyladenosine (m6A) is a universal post-transcriptional modification of RNAs, and it is widely involved in various biological processes. Identifying m6A modification sites accurately is indispensable to further investigate m6A-mediated biological functions. How to better represent RNA sequences is crucial for building effective computational methods for detecting m6A modification sites. However, traditional encoding methods require complex biological prior knowledge and are time-consuming. Furthermore, most of the existing m6A sites prediction methods are limited to single species, and few methods are able to predict m6A sites across different species and tissues. Thus, it is necessary to design a more efficient computational method to predict m6A sites across multiple species and tissues. In this paper, we proposed ELMo4m6A, a contextual language embedding-based method for predicting m6A sites from RNA sequences without any prior knowledge. ELMo4m6A first learns embeddings of RNA sequences using a language model ELMo, then uses a hybrid convolutional neural network (CNN) and long short-term memory (LSTM) to identify m6A sites. The results of 5-fold cross-validation and independent testing demonstrate that ELMo4m6A is superior to state-of-the-art methods. Moreover, we applied integrated gradients to find potential sequence patterns contributing to m6A sites.
Yongxian Fan, Guicong Sun, Xiaoyong Pan
IEEE ACM Trans. Comput. Biol. Bioinform.1
2022 An Improved Tensor Network for Image Classification in Histopathology
Yongxian Fan, Hao Gong 0006
PRCV (2)1
2022 GCRFLDA: scoring lncRNA-disease associations using graph convolution matrix completion with conditional random field
abstract
Long noncoding RNAs (lncRNAs) play important roles in various biological regulatory processes, and are closely related to the occurrence and development of diseases. Identifying lncRNA-disease associations is valuable for revealing the molecular mechanism of diseases and exploring treatment strategies. Thus, it is necessary to computationally predict lncRNA-disease associations as a complementary method for biological experiments. In this study, we proposed a novel prediction method GCRFLDA based on the graph convolutional matrix completion. GCRFLDA first constructed a graph using the available lncRNA-disease association information. Then, it constructed an encoder consisting of conditional random field and attention mechanism to learn efficient embeddings of nodes, and a decoder layer to score lncRNA-disease associations. In GCRFLDA, the Gaussian interaction profile kernels similarity and cosine similarity were fused as side information of lncRNA and disease nodes. Experimental results on four benchmark datasets show that GCRFLDA is superior to other existing methods. Moreover, we conducted case studies on four diseases and observed that 70 of 80 predicted associated lncRNAs were confirmed by the literature.
Yongxian Fan, Meijun Chen, Xiaoyong Pan
Briefings Bioinform.1
2022 StackEPI: identification of cell line-specific enhancer-promoter interactions based on stacking ensemble learning
abstract
BACKGROUND: Understanding the regulatory role of enhancer-promoter interactions (EPIs) on specific gene expression in cells contributes to the understanding of gene regulation, cell differentiation, etc., and its identification has been a challenging task. On the one hand, using traditional wet experimental methods to identify EPIs often means a lot of human labor and time costs. On the other hand, although the currently proposed computational methods have good recognition effects, they generally require a long training time. RESULTS: In this study, we studied the EPIs of six human cell lines and designed a cell line-specific EPIs prediction method based on a stacking ensemble learning strategy, which has better prediction performance and faster training speed, called StackEPI. Specifically, by combining different encoding schemes and machine learning methods, our prediction method can extract the cell line-specific effective information of enhancer and promoter gene sequences comprehensively and in many directions, and make accurate recognition of cell line-specific EPIs. Ultimately, the source code to implement StackEPI and experimental data involved in the experiment are available at https://github.com/20032303092/StackEPI.git . CONCLUSIONS: The comparison results show that our model can deliver better performance on the problem of identifying cell line-specific EPIs and outperform other state-of-the-art models. In addition, our model also has a more efficient computation speed.
Yongxian Fan, Binchao Peng
BMC Bioinform.1
2021 Using multi-layer perceptron to identify origins of replication in eukaryotes via informative features
abstract
BACKGROUND: The origin is the starting site of DNA replication, an extremely vital part of the informational inheritance between parents and children. More importantly, accurately identifying the origin of replication has great application value in the diagnosis and treatment of diseases related to genetic information errors, while the traditional biological experimental methods are time-consuming and laborious. RESULTS: We carried out research on the origin of replication in a variety of eukaryotes and proposed a unique prediction method for each species. Throughout the experiment, we collected data from 7 species, including Homo sapiens, Mus musculus, Drosophila melanogaster, Arabidopsis thaliana, Kluyveromyces lactis, Pichia pastoris and Schizosaccharomyces pombe. In addition to the commonly used sequence feature extraction methods PseKNC-II and Base-content, we designed a feature extraction method based on TF-IDF. Then the two-step method was utilized for feature selection. After comparing a variety of traditional machine learning classification models, the multi-layer perceptron was employed as the classification algorithm. Ultimately, the data and codes involved in the experiment are available at https://github.com/Sarahyouzi/EukOriginPredict . CONCLUSIONS: The prediction accuracy of the training set of the above-mentioned seven species after 100 times fivefold cross validation reach 92.60%, 90.80%, 91.22%, 96.15%, 96.72%, 99.86%, 96.72%, respectively. It denotes that compared with other methods, the methods we designed could accomplish superior performance. In addition, our experiments reveals that the models of multiple species could predict each other with high accuracy, and the results of STREME shows that they have a certain common motif.
Yongxian Fan
BMC Bioinform.1
2021 PyConvU-Net: a lightweight and multiscale network for biomedical image segmentation
abstract
BACKGROUND: With the development of deep learning (DL), more and more methods based on deep learning are proposed and achieve state-of-the-art performance in biomedical image segmentation. However, these methods are usually complex and require the support of powerful computing resources. According to the actual situation, it is impractical that we use huge computing resources in clinical situations. Thus, it is significant to develop accurate DL based biomedical image segmentation methods which depend on resources-constraint computing. RESULTS: A lightweight and multiscale network called PyConvU-Net is proposed to potentially work with low-resources computing. Through strictly controlled experiments, PyConvU-Net predictions have a good performance on three biomedical image segmentation tasks with the fewest parameters. CONCLUSIONS: Our experimental results preliminarily demonstrate the potential of proposed PyConvU-Net in biomedical image segmentation with resources-constraint computing.
Changyong Li, Yongxian Fan, Xiaodong Cai
BMC Bioinform.2
2021 I2DS: Interpretable Intrusion Detection System Using Autoencoder and Additive Tree
abstract
Intrusion detection system (IDS), the second security gate behind the firewall, can monitor the network without affecting the network performance and ensure the system security from the internal maximum. Many researches have applied traditional machine learning models, deep learning models, or hybrid models to IDS to improve detection effect. However, according to Predicted accuracy, Descriptive accuracy, and Relevancy (PDR) framework, most of detection models based on model-based interpretability lack good detection performance. To solve the problem, in this paper, we have proposed a novel intrusion detection system model based on model-based interpretability, called Interpretable Intrusion Detection System (I2DS). We firstly combine normal and attack samples reconstructed by AutoEncoder (AE) with training samples to highlight the normal and attack features, so that the classifier has a gorgeous effect. Then, Additive Tree (AddTree) is used as a binary classifier, which can provide excellent predictive performance in the combined dataset while maintaining good model-based interpretability. In the experiment, UNSW-NB15 dataset is used to evaluate our proposed model. For detection performance, I2DS achieves a detection accuracy of 99.95%, which is better than most of state-of-the-art intrusion detection methods. Moreover, I2DS maintains higher simulatability and captures the decision rules easily.
Yongxian Fan, Changyong Li
Secur. Commun. Networks2