VLDB 2026 Research / reviewers in the wild / expert
Xiujuan Lei
dblp:85/3128
· DBLP profile ↗
86ranked-venue papers
23as first author
47since 2021 · last 2026
0000-0002-9901-1732ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 58 · 13 first-author · 39 since 2021Artificial intelligence and machine learning · 22 · 6 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HybridSeqNet: A Deep Learning Framework for Blood Pressure Estimation
Fei Wang 0095, Feiyu Yu, Xiujuan Lei, Fang-Xiang Wu, Yansen Su, Chun-Hou Zheng 0001, Junfeng Xia |
ICIC (29) | 4 |
| 2026 | circ-EGAT: Enhanced Graph Attention Network with Multi-information Integration for circRNA-RBP Binding Sites Prediction
Yajing Guo, Xiujuan Lei |
ISBRA (1) | 2 |
| 2026 | SpatialPEFT: a parameter-efficient fine-tuning framework for spatial transcriptomics foundation modelsabstractSUMMARY: SpatialPEFT is a unified parameter-efficient fine-tuning framework that enables the robust adaptation of large spatial transcriptomics foundation models (up to 1.4 billion parameters) on a single 16 GB consumer-grade GPU. By integrating Low-Rank Adaptation (LoRA), gradient checkpointing, and a spatial-aware adapter, it reduces peak VRAM by over 87% while substantially improving downstream spatial annotation accuracy. AVAILABILITY AND IMPLEMENTATION: SpatialPEFT is implemented in Python and released under the MIT license. The source code, documentation, and tutorials are freely available at https://github.com/applerplay/SpatialPEFT, with an archival snapshot deposited at Zenodo (DOI: 10.5281/zenodo.20725321). Xiujuan Lei |
Bioinform. | 2 |
| 2026 | DTIBFAI: drug-target interaction prediction based on BERT and feature augment of Informer
Naichao Wang, Yihe Diwu, Mingchen Feng, Yuchen Zhang 0003, Xiujuan Lei |
Frontiers Comput. Sci. | 5 |
| 2026 | Unseen TCR-epitope interaction prediction with self-supervised contrastive learning
Rawshon Raha, Weilai Chi, Xiujuan Lei, Fang-Xiang Wu |
Neurocomputing | 3 |
| 2026 | Dual-Channel Learning Framework for miRNA-Drug Interaction Prediction Based on Structural Features and Signed Bipartite Graph Neural NetworkabstractMicroRNAs (miRNAs) play a vital role in regulating a wide range of biological functions and are key players in the development of many complex human diseases, making them novel therapeutic targets for drug development. Given the high expenses and time demands of traditional experimental methods, it is essential to develop efficient computational approaches for predicting miRNA-drug interactions (MDIs). This article presents a dual-channel learning framework, SSMDI, based on structural features and Signed Bipartite Graph Neural Network (SBGNN) for predicting MDIs. Firstly, Graph Isomorphism Networks (GIN) is employed to extract molecular graph features of drugs. Meanwhile, a combined framework of Convolutional Neural Network (CNN), Bidirectional Long Short-Term Memory (BiLSTM) network and Self-attention Mechanism is utilized to capture sequence features of miRNAs. Compared with traditional networks, signed networks can deliver richer semantic information in drugs and miRNAs. Therefore, SBGNN is then used to aggregate and update the signed topological features of miRNAs and drugs. Finally, structural and signed topological features are integrated to predict MDIs. The predictive performance of the model is evaluated using 5-fold cross-validation (CV), achieving AUC of 0.9447 and AUPR of 0.9238. The case study further demonstrates the effectiveness of SSMDI in predicting MDIs. In summary, the SSMDI model proves to be an accurate tool for predicting MDIs, which holds significant implications for drug development and miRNA-based therapeutic research. Xiujuan Lei, Fang-Xiang Wu, Yi Pan 0001 |
IEEE Trans. Big Data | 2 |
| 2026 | Prediction of circRNA-Drug Associations Based on Bipartite Graph TransformerabstractCircular RNAs (circRNAs) represent a distinctive class of non-coding RNAs with covalently closed loop structures that play crucial regulatory roles in drug response. While existing computational methods have achieved certain progress in prediction tasks, they primarily relied on circRNA genotypes and traditional molecular fingerprints, with limited utilization of multi-omics data and inadequate consideration of heterogeneous network topology. To address these limitations, this study proposed the CircRNA-Drug Bipartite Graph Transformer (CDBGT) framework to predict associations. Rather than limiting to associations between circRNA genotypes and drugs, this study integrated circRNA-drug response and target association information from multiple databases. CDBGT employed pre-trained models RNA-FM and ChemBERTa to extract features of sequence and molecular fingerprint and utilized multi-omics data to construct similarity matrices. The framework incorporated a bipartite graph transformer with topological positional encoding, comprehensively considering degree encoding, degree ranking encoding and spectral encoding to extract topological information from heterogeneous networks. Experimental results showed that CDBGT performed stably in 5-fold cross-validation. On the Response dataset, it achieved ROC-AUC of 0.9674 and PR-AUC of 0.9540, while on the Target dataset it reached ROC-AUC of 0.8621. Compared with existing methods, it showed an improvement of 3.20 to 26.87 percentage points in ROC-AUC. Ablation experiments demonstrated the necessity of each module. Through literature-supported case studies, this work suggested potential directions for circRNA-based therapeutic research. Yuchen Zhang 0003, Xiujuan Lei |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Synergistic Drug Combination Prediction via Graphormer and Drug-Cell Line Pair GraphabstractDrug combination synergy is crucial in pharmacology, as it can enhance disease treatment efficacy or reduce drug resistance when administered in combination. Accurate prediction of drug combination synergy is vital for optimizing therapeutic regimens and improving treatment effectiveness. However, existing computational methods primarily rely on drug sequence and structural features, making them difficult to capture complex network relationships and global information-especially lacking the ability to perform cross-modal fusion. In this study, we proposed a method called SDCGDCP for predicting drug combination synergy. It processed drug molecular structures (graph structures and molecular fingerprints), target biological activity information and integrated cell line whole-genome expression profiles to construct multi-level combined node representations. A drug-cell line pair graph was accordingly generated. SDCGDCP updated node and edge representations via a GatedGCN module and derived five types of structural encodings (centrality, spatial, edge, Laplacian positional and node-similarity encodings), after which the graph language model Graphormer is employed to capture long-range node interactions. Finally, drug combination synergy was predicted using an MLP and SoftMax classifier. Extensive evaluations showed that SDCGDCP outperforms other state-of-the-art methods on the DrugCombDB dataset, achieving an AUROC of 0.923 and AUPRC of 0.885. Ablation experiments validated the effectiveness of each feature and encoding module. Meanwhile, we conducted case analyses on the predicted drug combination synergies. The results were supported by evidence from several pharmaceutical studies. This highlights potential of SDCGDCP in enhancing drug synergy prediction and optimizing combination therapies. The source code and data for SDCGDCP are available at https://github.com/Philosopher-Zhao/SDCGDCP. Yuchen Zhang 0003, Bingzhe Zhao, Zhuoqun Fu, Yiming Han, Beidan Liu, Xiujuan Lei |
BIBM | 6 |
| 2025 | An Adaptive Multi-view Feature Fusion Framework Based on Multiple Graphs for Predicting Drug-Drug Interactions
Fei Wang 0095, Zefan Cheng, Xiujuan Lei, Fang-Xiang Wu, Chun-Hou Zheng 0001, Yansen Su |
ICIC (26) | 3 |
| 2025 | Multi-task Learning with Cross-Stitch for Synergistic Effect of Drug Combination Prediction
Anqi Liang, Xiujuan Lei, Yi Pan 0001 |
ISBRA (1) | 2 |
| 2025 | Hyperbolic multivariate feature learning in higher-order heterogeneous networks for drug-disease prediction
Jianrui Chen 0002, Xiujuan Lei |
Artif. Intell. Medicine | 4 |
| 2025 | Predicting miRNA-drug interactions via dual-channel network based on TCN and BiLSTM
Xiujuan Lei |
Frontiers Comput. Sci. | 2 |
| 2025 | Multi-Source Data with Laplacian Eigenmaps and Denoising Autoencoder for Predicting Microbe-Disease Associations via Convolutional Neural Network
Xiujuan Lei, Yi Pan 0001 |
J. Comput. Sci. Technol. | 1 |
| 2025 | MSFusion: A multi-source hybrid feature fusion network for accurate grading of invasive breast cancer using H&E-stained histopathological images
Jiayang Bai, Jinjie Wang, Duanbo Shi, Xiujuan Lei, Cheng Lu 0001 |
Medical Image Anal. | 7 |
| 2025 | Nucleotide-level circRNA-RBP binding sites prediction based on hybrid encoding scheme and enhanced feature extraction
Yajing Guo, Xiujuan Lei, Zhengfeng Wang, Fang-Xiang Wu, Yi Pan 0001 |
Neural Networks | 2 |
| 2025 | MKMGCN-DDI: Predicting Drug-Drug Interactions via Magnetic Graph Convolutional Network With Multiple KernelsabstractPolypharmacy is a common means of clinical treatments, but detecting drug-drug interactions (DDIs) behind unexpected effects can be costly and faces clinical limitations. Recently, graph neural networks (GNNs) have demonstrated encouraging performance in predicting DDIs. However, most studies overlook the comprehensive aspects of DDIs, such as the coexistence of types of pharmacological changes and the asymmetric roles of drugs. In this article, we define new prediction tasks, taking into account both enhancive or depressive changes and the roles of drugs, and then establish spectral GNNs to predict comprehensive information of DDIs. First, we formally define several tasks, including joint prediction tasks designed to leverage both types and directions. These tasks deduce to sub-tasks in previous studies. Then, we propose a unified framework, the MKMGCN-DDI, via introducing two Magnetic Laplacian matrices to encode comprehension information within DDIs, defining multiple graph filters, and designing multiple-kernel based Magnetic graph convolutional networks (MKMGCN). Experiments across three datasets show that it not only has good adaptability to multiple tasks but also significantly improves results on simple tasks. Case studies on breast neoplasms and lung neoplasms verify its feasibility, as over half of top-10 items are supported. Yunhan Pan, Xiujuan Lei, Chunyan Ji, Yinglong Dai, Yi Pan 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | BiAtt-GVAE: Molecular Design for Specific Target via Graph Variational Autoencoder Based on Bi-Channel Interactive Attention NetworkabstractDesigning bioactive molecules with desired properties for specific targets is a longstanding challenge in drug design. We introduce a model called BiAtt-GVAE, which incorporates a conditional to more effectively constrain and enhance the generated molecules. To capture interaction sites between ligand atoms and protein amino acids, as well as the structure-properties relationship of ligands, we designed a bi-channel interactive attention network. The interactive attention block operates on the protein-ligand graph, facilitating the interactive learning of ligand atoms and protein amino acid features. We consider various properties to learn the ligand structure-properties relationship through a multi-head cross-attention block. Testing on EGFR and CDK2 targets, as well as through a dedicated case study for designing SARS-CoV-2 Mpro inhibitors demonstrates the effectiveness of BiAtt-GVAE. It has been verified through subsequent investigation of the molecular structure and molecular docking of Mpro inhibitors that BiAtt-GVAE can produce compounds with high ratings for both novelty and affinity. Mei Ma, Xiujuan Lei |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2025 | Predicting Drug-miRNA Associations Combining SDNE with BiGRUabstractIt is well recognized that abnormal miRNA expression can result in drug resistance and pose a challenge to miRNA-based treatments. However, the drug-miRNA associations (DMA) are still incompletely understood. Conventional biological experiments have a high failure rate, lengthy cycle times, and expensive expenditures. Consequently, deep learning-based techniques for predicting DMA have been developed. In this work, we propose a novel method named SDNEDMA for DMA prediction that combines SDNE with BiGRU. The two-channel approach is used to combine the attribute and topological features of miRNAs and drugs. To be more precise, we first model the associations between drugs and miRNAs through the known bipartite network, and then utilize SDNE to obtain the topological features. Meanwhile, BiGRU is employed to acquire miRNA k-mer sequence features and drug ECFP fingerprints. Subsequently, both the topological and attribute features are fused jointly to form final features which is aimed to predict the association score for both them. Multiple features drugs and miRNAs are used at the same time, more information is fused, and the features are more accurate, so the prediction performance is better. The experiments show that SDNEDMA outperforms other state-of-the-art methods, yielding AUC of 0.9641 when we use 5-fold cross-validation on the ncDR dataset. SDNEDMA is additionally employed in a case study, showing how accurate and dependable it is. To sum up, the SDNEDMA has the ability to predict DMA with high accuracy and effectiveness, which is really important for drug development. Chenyue Lei, Xiujuan Lei |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | Molecular Structure-Driven Multi-Relation DGI Prediction With High-Low-Order Attention DenoiseabstractDrug-Gene Interaction (DGI) is crucial for drug discovery and personalized medicine. The continuous development of genomics and drug repositioning has brought increasing attention to the complex relations between drugs and genes. However, traditional biological experiments are time-consuming and costly, which makes it challenging to efficiently explore the multi-relational interactions between drugs and genes. Therefore, computational approaches aim to develop efficient schemes for predicting drug-gene relations to reduce the search space and experimental costs. Existing computational methods often suffer from data scarcity and poor generalization, which pose significant challenges for practical applications. To address these issues, we propose a novel multi-relation DGI prediction method based on molecular structure-driving and high-low-order attention denoising framework. Our approach captures molecular structural information through both atom and bond channels with a drug feature encoder. For network structure, we enhance both high- and low-order channels: the low-order channel leverages graph convolutional networks, while the high-order channel employs hypergraph-based message propagation. Additionally, we adopt consistency information loss and inter-channel attention mechanism to refine high- and low-order features. Experimental results on three drug-gene datasets demonstrate the superior performance of our model, particularly on sparse datasets DrugBank and DGIdb, with F1 improvements of 4.06% and 5.67%, respectively. Yizhe Shang, Jianrui Chen 0002, Xiujuan Lei, Fang-Xiang Wu |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | GIAE-DTI: Predicting Drug-Target Interactions Based on Heterogeneous Network and GIN-Based Graph AutoencoderabstractAccurate prediction of drug-target interactions (DTIs) is essential for advancing drug discovery and repurposing. However, the sparsity of DTI data limits the effectiveness of existing computational methods, which primarily focus on sparse DTI networks and have poor performance in aggregating information from neighboring nodes and representing isolated nodes within the network. In this study, we propose a novel deep learning framework, named GIAE-DTI, which considers cross-modal similarity of drugs and targets and constructs a heterogeneous network for DTI prediction. Firstly, the model calculates the cross-modal similarity of drugs and proteins from the relationships among drugs, proteins, diseases, and side effects, and performs similarity integration by taking the average. Then, a drug-target heterogeneous network is constructed, including drug-drug interactions, protein-protein interactions, and drug-target interactions processed by weighted K nearest known neighbors. In the heterogeneous network, a graph autoencoder based on a graph isomorphism network is employed for feature extraction, while a dual decoder is utilized to achieve better self-supervised learning, resulting in latent feature representations for drugs and targets. Finally, a deep neural network is employed to predict DTIs. The experimental results indicate that on the benchmark dataset, GIAE-DTI achieves AUC and AUPR scores of 0.9533 and 0.9619, respectively, in DTI prediction, outperforming the current state-of-the-art methods. Additionally, case studies on four 5-hydroxytryptamine receptor-related targets and five drugs related to mental diseases show the great potential of the proposed method in practical applications. Xiujuan Lei, Jianrui Chen 0002, Fang-Xiang Wu |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | HCCL: Hierarchical Channels and Contrastive Learning for Drug-Gene Multi-Relation PredictionabstractDrug-gene interaction plays a crucial role in drug discovery and personalized medicine. Although existing methods have improved the accuracy of exploring multiple relationships between drugs and genes, there are still some limitations, such as susceptibility to data sparsity and poor generalization, which pose some challenges for practical applications. To address these challenges, we propose a novel Hierarchical Channels and Contrastive Learning (HCCL) framework in which drug feature extractor captures structural information of drug molecules from atom and bond channels. After obtaining the initial features of drugs and genes, we employ high-low-order channels to update them, where the low-order channel adopts graph convolutional networks while the high-order channel leverages hypergraph structures for message propagation. Finally, we adopt contrastive learning and inter-channel attention to fuse high-low-order features, which improves the robustness of the model and prevents feature information loss. Experimental results demonstrate the superior performance of HCCL. Yizhe Shang, Jianrui Chen 0002, Xiujuan Lei, Fang-Xiang Wu |
BIBM | 3 |
| 2024 | Prediction of miRNA-Disease Associations Based on Hybrid Gated GNN and Multi-Data IntegrationabstractIt is well-established that miRNAs play a crucial role in the occurrence and development of diseases. Current miRNA-disease associations prediction research faces several challenges, including model bias due to data sparsity, information loss from overlooked complex relationships during feature fusion and insufficient capability of existing methods to capture the intricate relationships (between miRNAs, genes, lncRNAs and diseases), thereby limiting prediction accuracy. Based on hybrid gated GNN and multi-data fusion, a method (PMDGGM) for predicting miRNA-disease associations is proposed in this study. PMDGGM constructed seven similarity networks by comprehensively considering the relationships between miRNA and related genes, miRNA, lncRNA and diseases. It provides a solid foundation for feature fusion and information propagation. Subsequently, the method captures the complex relationships between heterogeneous nodes through a bilinear pooling layer and uses a gating mechanism to fuse multi-source heterogeneous features, thereby predicting miRNA-disease associations more accurately. The experimental results show that the method performs well and has significant advantages in predicting the miRNA-disease associations. Among various evaluation metrics, especially the AUCs of ROC and PR curves, the performance of method is outstanding, reaching a high level of 0.9413 and 0.9362. The study conducted case analyses on two diseases heart failure and acute myeloid leukemia. The predicted associated miRNAs can be validated by existing biomedical research efforts. The source code and data of PMDGGM can be publicly accessed on GitHub for further research and verification: https://github.com/WangYeQianger/PMDGGM Yeqiang Wang, Sharen Yun, Yuchen Zhang 0003, Xiujuan Lei |
BIBM | 4 |
| 2024 | Drug Combination Side Effect Prediction Based on Polypharmacy Network and GraphSAGE AlgorithmabstractDue to the complexity and diversity of modern diseases, the combination of drugs has become the first choice. According to the graph theory in graph theory, we transform the original link prediction problem into the node identification problem by establishing the graph of drug side effect network - polypharmacy network to predict. A combined drug side effect prediction model (CDSG) was constructed on the polypharmacy network graph, based on graph attention mechanism and Graph Sample and Aggregate (GraphSAGE). Firstly, the side-effect network was constructed by using the side-effect relationship between drugs and drugs, and then the polypharmacy network was constructed. Then the target gene of the drug is encoded and the feature vector of the drug is established. Furthermore, the advanced feature representations of drugs are learned by utilizing the graph attention network and the GraphSAGE algorithm. Finally, the advanced drug characteristics were connected to the fully connected layer for classification prediction. On a baseline dataset of 1138 side effect types of 257 drugs, we conducted different methods of feature fusion experiment, ablation experiment, 5-fold crossover experiment, and compared CDSG with several computing models such as traditional GCN model, the GAT model and matrix decomposition models, and the final model achieved good results. Haiqiang Xiao, Xiujuan Lei, Yuchen Zhang 0003, Fang-Xiang Wu |
BIBM | 2 |
| 2024 | Multi-filter Based Signed Graph Convolutional Networks for Predicting Interactions on Drug Networks
Zitao Hu, Xiujuan Lei, Chunyan Ji, Zhao Tong 0001, Yi Pan 0001 |
ISBRA (2) | 3 |
| 2024 | HoRDA: Learning higher-order structure information for predicting RNA-disease associations
Julong Li, Jianrui Chen 0002, Zhihui Wang 0002, Xiujuan Lei |
Artif. Intell. Medicine | 4 |
| 2024 | A systematic review on deep learning based methods for cervical cell image analysisabstractCervical cytology image analysis is indispensable for the detection of abnormal cervical cells. Traditionally, manual screening is time-consuming and labor-intensive. Therefore, a lot of deep learning (DL)-based automatic detection methods have been employed in this field to provide timely, accurate and objective results. In this study, we systematically review the current developments in cervical cell image analysis with DL methods. Specifically, we first present the most popular DL models that are widely applied in cervical cell analysis. Second, we describe the methodology for conducting this review. Third, we provide all publicly available datasets related to cervical cell images to the best of our knowledge. Then, we introduce relevant evaluation metrics and loss functions. Next, we summarize and assort the applications for cervical cell classification and segmentation. Afterwards, we discuss about current challenges and future research directions in this field. Finally, we draw the conclusion of this review. According to the analysis, we conclude that the studies based on DL models have maintained an increasing trend in recent years, which indicates the potential of DL in cervical cell image analysis. In cervical cell image classification, CNN is the most commonly used DL model. Among CNN models, we can find that VGGNet and ResNet are the most popular network architectures for the classification of cervical cells. Transformer is the second commonly used DL model. Moreover, Herlev and SIPaKMeD are the most popular public datasets used for cervical cell classification. In cervical cell segmentation, U-Net and FCN are the two most popular DL architectures. In addition, ISBI2014 and Herlev datasets are the most frequently used among the existing publicly available segmentation datasets. However, there are some issues in this field, such as poor cervical cell classification performance as a result of similar pathological properties between different cell categories. Therefore, it is necessary to develop more effective methods with DL models to improve these issues in the future research. Bo Liao 0001, Xiujuan Lei, Fang-Xiang Wu |
Neurocomputing | 3 |
| 2023 | circ2CBA: prediction of circRNA-RBP binding sites combining deep learning and attention mechanism
Yajing Guo, Xiujuan Lei, Yi Pan 0001 |
Frontiers Comput. Sci. | 2 |
| 2023 | Identify potential circRNA-disease associations through a multi-objective evolutionary algorithm
Yuchen Zhang 0003, Xiujuan Lei, Cai Dai, Yi Pan 0001, Fang-Xiang Wu |
Inf. Sci. | 2 |
| 2023 | RMDGCN: Prediction of RNA methylation and disease associations based on graph convolutional network with attention mechanismabstractRNA modification is a post transcriptional modification that occurs in all organisms and plays a crucial role in the stages of RNA life, closely related to many life processes. As one of the newly discovered modifications, N1-methyladenosine (m1A) plays an important role in gene expression regulation, closely related to the occurrence and development of diseases. However, due to the low abundance of m1A, verifying the associations between m1As and diseases through wet experiments requires a great quantity of manpower and resources. In this study, we proposed a computational method for predicting the associations of RNA methylation and disease based on graph convolutional network (RMDGCN) with attention mechanism. We build an adjacency matrix through the collected m1As and diseases associations, and use positive-unlabeled learning to increase the number of positive samples. By extracting the features of m1As and diseases, a heterogeneous network is constructed, and a GCN with attention mechanism is adopted to predict the associations between m1As and diseases. The experimental results indicate that under a 5-fold cross validation, RMDGCN is superior to other methods (AUC = 0.9892 and AUPR = 0.8682). In addition, case studies indicate that RMDGCN can predict the relationships between unknown m1As and diseases. In summary, RMDGCN is an effective method for predicting the associations between m1As and diseases. Yumeng Zhou, Xiujuan Lei |
PLoS Comput. Biol. | 3 |
| 2023 | A dual graph neural network for drug-drug interactions prediction based on molecular structure and interactionsabstractExpressive molecular representation plays critical roles in researching drug design, while effective methods are beneficial to learning molecular representations and solving related problems in drug discovery, especially for drug-drug interactions (DDIs) prediction. Recently, a lot of work has been put forward using graph neural networks (GNNs) to forecast DDIs and learn molecular representations. However, under the current GNNs structure, the majority of approaches learn drug molecular representation from one-dimensional string or two-dimensional molecular graph structure, while the interaction information between chemical substructure remains rarely explored, and it is neglected to identify key substructures that contribute significantly to the DDIs prediction. Therefore, we proposed a dual graph neural network named DGNN-DDI to learn drug molecular features by using molecular structure and interactions. Specifically, we first designed a directed message passing neural network with substructure attention mechanism (SA-DMPNN) to adaptively extract substructures. Second, in order to improve the final features, we separated the drug-drug interactions into pairwise interactions between each drug's unique substructures. Then, the features are adopted to predict interaction probability of a DDI tuple. We evaluated DGNN-DDI on real-world dataset. Compared to state-of-the-art methods, the model improved DDIs prediction performance. We also conducted case study on existing drugs aiming to predict drug combinations that may be effective for the novel coronavirus disease 2019 (COVID-19). Moreover, the visual interpretation results proved that the DGNN-DDI was sensitive to the structure information of drugs and able to detect the key substructures for DDIs. These advantages demonstrated that the proposed method enhanced the performance and interpretation capability of DDI prediction modeling. Mei Ma, Xiujuan Lei |
PLoS Comput. Biol. | 2 |
| 2023 | Biomarker Identification via a Factorization Machine-Based Neural Network With Binary Pairwise EncodingabstractBiomolecules, microRNAs (miRNAs) and long non-coding RNAs (lncRNAs), play critical roles in diverse fundamental and vital biological processes. They can serve as disease biomarkers as their dysregulations could cause complex human diseases. Identifying those biomarkers is helpful with the diagnosis, treatment, prognosis, and prevention of diseases. In this study, we propose a factorization machine-based deep neural network with binary pairwise encoding, DFMbpe, to identify the disease-related biomarkers. First, to comprehensively consider the interdependence of features, a binary pairwise encoding method is designed to obtain the raw feature representations for each biomarker-disease pair. Second, the raw features are mapped into their corresponding embedding vectors. Then, the factorization machine is conducted to get the wide low-order feature interdependence, while the deep neural network is applied to obtain the deep high-order feature interdependence. Finally, two kinds of features are combined to get the final prediction results. Unlike other biomarker identification models, the binary pairwise encoding considers the interdependence of features even though they never appear in the same sample, and the DFMbpe architecture emphasizes both low-order and high-order feature interactions simultaneously. The experimental results show that DFMbpe greatly outperforms the state-of-the-art identification models on both cross-validation and independent dataset evaluation. Besides, three types of case studies further demonstrate the effectiveness of this model. Yulian Ding, Xiujuan Lei, Bo Liao 0001, Fang-Xiang Wu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | Microbe-Disease Association Prediction Using RGCN Through Microbe-Drug-Disease NetworkabstractAccumulating evidence has shown that microbes play significant roles in human health and diseases. Therefore, identifying microbe-disease associations is conducive to disease prevention. In this article, a predictive method called TNRGCN is designed for microbe-disease associations based on Microbe-Drug-Disease Network and Relation Graph Convolutional Network (RGCN). First, considering that indirect links between microbes and diseases will be increased by introducing drug related associations, we construct a Microbe-Drug-Disease tripartite network through data processing from four databases including Human Microbe-Disease Association Database (HMDAD), Disbiome Database, Microbe-Drug Association Database (MDAD) and Comparative Toxicoge-nomics Database (CTD). Second, we construct similarity networks for microbes, diseases and drugs via microbe function similarity, disease semantic similarity and Gaussian interaction profile kernel similarity, respectively. Based on the similarity networks, Principal Component Analysis (PCA) is utilized to extract main features of nodes. These features will be input into the RGCN as initial features. Finally, based on the tripartite network and initial features, we design two-layer RGCN to predict microbe-disease associations. Experimental results indicate that TNRGCN achieves best performance in cross validation compared with other methods. Meanwhile, case studies for Type 2 diabetes (T2D), Bipolar disorder and Autism demonstrate the favorable effectiveness of TNRGCN in association prediction. Xiujuan Lei, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | MLRDFM: a multi-view Laplacian regularized DeepFM model for predicting miRNA-disease associationsabstractMOTIVATION: MicroRNAs (miRNAs), as critical regulators, are involved in various fundamental and vital biological processes, and their abnormalities are closely related to human diseases. Predicting disease-related miRNAs is beneficial to uncovering new biomarkers for the prevention, detection, prognosis, diagnosis and treatment of complex diseases. RESULTS: In this study, we propose a multi-view Laplacian regularized deep factorization machine (DeepFM) model, MLRDFM, to predict novel miRNA-disease associations while improving the standard DeepFM. Specifically, MLRDFM improves DeepFM from two aspects: first, MLRDFM takes the relationships among items into consideration by regularizing their embedding features via their similarity-based Laplacians. In this study, miRNA Laplacian regularization integrates four types of miRNA similarity, while disease Laplacian regularization integrates two types of disease similarity. Second, to judiciously train our model, Laplacian eigenmaps are utilized to initialize the weights in the dense embedding layer. The experimental results on the latest HMDD v3.2 dataset show that MLRDFM improves the performance and reduces the overfitting phenomenon of DeepFM. Besides, MLRDFM is greatly superior to the state-of-the-art models in miRNA-disease association prediction in terms of different evaluation metrics with the 5-fold cross-validation. Furthermore, case studies further demonstrate the effectiveness of MLRDFM. Yulian Ding, Xiujuan Lei, Bo Liao 0001, Fang-Xiang Wu |
Briefings Bioinform. | 2 |
| 2022 | Predicting drug-drug interactions by graph convolutional network with multi-kernelabstractDrug repositioning is proposed to find novel usages for existing drugs. Among many types of drug repositioning approaches, predicting drug-drug interactions (DDIs) helps explore the pharmacological functions of drugs and achieves potential drugs for novel treatments. A number of models have been applied to predict DDIs. The DDI network, which is constructed from the known DDIs, is a common part in many of the existing methods. However, the functions of DDIs are different, and thus integrating them in a single DDI graph may overlook some useful information. We propose a graph convolutional network with multi-kernel (GCNMK) to predict potential DDIs. GCNMK adopts two DDI graph kernels for the graph convolutional layers, namely, increased DDI graph consisting of 'increase'-related DDIs and decreased DDI graph consisting of 'decrease'-related DDIs. The learned drug features are fed into a block with three fully connected layers for the DDI prediction. We compare various types of drug features, whereas the target feature of drugs outperforms all other types of features and their concatenated features. In comparison with three different DDI prediction methods, our proposed GCNMK achieves the best performance in terms of area under receiver operating characteristic curve and area under precision-recall curve. In case studies, we identify the top 20 potential DDIs from all unknown DDIs, and the top 10 potential DDIs from the unknown DDIs among breast, colorectal and lung neoplasms-related drugs. Most of them have evidence to support the existence of their interactions. [email protected]. Fei Wang 0095, Xiujuan Lei, Bo Liao 0001, Fang-Xiang Wu |
Briefings Bioinform. | 2 |
| 2022 | Inferring Metabolite-Disease Association Using Graph Convolutional NetworksabstractAs is well known, biological experiments are time-consuming and laborious, so there is absolutely no doubt that developing an effective computational model will help solve these problems. Most of computational models rely on the biological similarity and network-based methods that cannot consider the topological structures of metabolite-disease association graphs. We proposed a novel method based on graph convolutional networks to infer potential metabolite-disease association, named MDAGCN. We first calculated three kinds of metabolite similarities and three kinds of disease similarities. The final similarity of disease and metabolite will be obtained by integrating three kinds' similarities of each and filtering out the noise similarity values. Then metabolite similarity network, disease similarity network and known metabolite-disease association network were used to construct a heterogenous network. Finally, heterogeneous network with rich information is fed into the graph convolutional networks to obtain new features of a node through aggregation of node information so as to infer the potential associations between metabolites and diseases. Experimental results show that MDAGCN achieves more reliable results in cross validation and case studies when compared with other existing methods. Xiujuan Lei, Jiaojiao Tie, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2022 | Identifying Gene Signatures for Cancer Drug Repositioning Based on Sample ClusteringabstractDrug repositioning is an important approach for drug discovery. Computational drug repositioning approaches typically use a gene signature to represent a particular disease and connect the gene signature with drug perturbation profiles. Although disease samples, especially from cancer, may be heterogeneous, most existing methods consider them as a homogeneous set to identify differentially expressed genes (DEGs)for further determining a gene signature. As a result, some genes that should be in a gene signature may be averaged off. In this study, we propose a new framework to identify gene signatures for cancer drug repositioning based on sample clustering (GS4CDRSC). GS4CDRSC first groups samples into several clusters based on their gene expression profiles. Second, an existing method is applied to the samples in each cluster for generating a list of DEGs. Then a weighting approach is used to identify an intergrated gene signature from all the lists of DEGs. The integrated gene signature is used to connect with drug perturbation profiles in the Connectivity Map (CMap)database to generate a list of drug candidates. GS4CDRSC has been tested with several cancer datasets and existing methods. The computational results show that GS4CDRSC outperforms those methods without the sample clustering and weighting approaches in terms of both number and rate of predicted known drugs for specific cancers. Fei Wang 0095, Yulian Ding, Xiujuan Lei, Bo Liao 0001, Fang-Xiang Wu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | Predicting Microbe-Disease Association Based on Multiple Similarities and LINE AlgorithmabstractNumerous microbes have been found to have vital impacts on human health through affecting biological processes. Therefore, exploring potential associations between microbes and diseases will promote the understanding and diagnosis of diseases. In this study, we present a novel computational model, named MSLINE, to infer potential microbe-disease associations by integrating Multiple Similarities and Large-scale Information Network Embedding (LINE) based on known associations. Specifically, on the basis of known microbe-disease associations from the Human Microbe-Disease Association Database, we first increase the known associations by collecting proven associations from existing literatures. We then construct a microbe-disease heterogeneous network (MDHN) by integrating known associations and multiple similarities (including Gaussian interaction profile kernel similarity, microbe function similarity, disease semantic similarity and disease-symptom similarity). After that, we implement random walk and LINE algorithm on MDHN to learn its structure information. Finally, we score the microbe-disease associations according to the structure information for every nodes. In the Leave-one-out cross validation and 5-fold cross validation, MSLINE performs better compared to other existing methods. Moreover, case studies of different diseases proved that MSLINE could predict the potential microbe-disease associations efficiently. Xiujuan Lei, Cheng Lu 0001, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | Predicting miRNA-Disease Associations Based On Multi-View Variational Graph Auto-Encoder With Matrix FactorizationabstractMicroRNAs (miRNAs) have been proved to play critical roles in diverse biological processes, including the human disease development process. Exploring the potential associations between miRNAs and diseases can help us better understand complex disease mechanisms. Given that traditional biological experiments are expensive and time-consuming, computational models can serve as efficient means to uncover potential miRNA-disease associations. This study presents a new computational model based on variational graph auto-encoder with matrix factorization (VGAMF) for miRNA-disease association prediction. More specifically, VGAMF first integrates four different types of information about miRNAs into an miRNA comprehensive similarity network and two types of information about diseases into a disease comprehensive similarity network, respectively. Then, VGAMF gets the non-linear representations of miRNAs and diseases, respectively, from those two comprehensive similarity networks with variational graph auto-encoders. Simultaneously, a non-negative matrix factorization is conducted on the miRNA-disease association matrix to get the linear representations of miRNAs and diseases. Finally, a fully connected neural network combines linear and non-linear representations of miRNAs and diseases to get the final predicted association score for all miRNA-disease pairs. In the 10-fold cross-validation experiments, VGAMF achieves an average AUC of 0.9280 on HMDD v2.0 and 0.9470 on HMDD v3.2, which outperforms other competing methods. Besides, the case studies on colon cancer and esophageal cancer further demonstrate the effectiveness of VGAMF in predicting novel miRNA-disease associations. Yulian Ding, Xiujuan Lei, Bo Liao 0001, Fang-Xiang Wu |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | Predicting Microbe-Disease Association via Tripartite Network and Relation Graph Convolutional Network
Xiujuan Lei, Yi Pan 0001 |
ISBRA | 2 |
| 2021 | A comprehensive survey on computational methods of non-coding RNA and disease association predictionabstractThe studies on relationships between non-coding RNAs and diseases are widely carried out in recent years. A large number of experimental methods and technologies of producing biological data have also been developed. However, due to their high labor cost and production time, nowadays, calculation-based methods, especially machine learning and deep learning methods, have received a lot of attention and been used commonly to solve these problems. From a computational point of view, this survey mainly introduces three common non-coding RNAs, i.e. miRNAs, lncRNAs and circRNAs, and the related computational methods for predicting their association with diseases. First, the mainstream databases of above three non-coding RNAs are introduced in detail. Then, we present several methods for RNA similarity and disease similarity calculations. Later, we investigate ncRNA-disease prediction methods in details and classify these methods into five types: network propagating, recommend system, matrix completion, machine learning and deep learning. Furthermore, we provide a summary of the applications of these five types of computational methods in predicting the associations between diseases and miRNAs, lncRNAs and circRNAs, respectively. Finally, the advantages and limitations of various methods are identified, and future researches and challenges are also discussed. Xiujuan Lei, Thosini Bamunu Mudiyanselage, Yuchen Zhang 0003, Chen Bian, Wei Lan 0001, Ning Yu 0004, Yi Pan 0001 |
Briefings Bioinform. | 1 |
| 2021 | Prediction of RBP binding sites on circRNAs using an LSTM-based deep sequence learning architectureabstractCircular RNAs (circRNAs) are widely expressed in highly diverged eukaryotes. Although circRNAs have been known for many years, their function remains unclear. Interaction with RNA-binding protein (RBP) to influence post-transcriptional regulation is considered to be an important pathway for circRNA function, such as acting as an oncogenic RBP sponge to inhibit cancer. In this study, we design a deep learning framework, CRPBsites, to predict the binding sites of RBPs on circRNAs. In this model, the sequences of variable-length binding sites are transformed into embedding vectors by word2vec model. Bidirectional LSTM is used to encode the embedding vectors of binding sites, and then they are fed into another LSTM decoder for decoding and classification tasks. To train and test the model, we construct four datasets that contain sequences of variable-length binding sites on circRNAs, and each set corresponds to an RBP, which is overexpressed in bladder cancer tissues. Experimental results on four datasets and comparison with other existing models show that CRPBsites has superior performance. Afterwards, we found that there were highly similar binding motifs in the four binding site datasets. Finally, we applied well-trained CRPBsites to identify the binding sites of IGF2BP1 on circCDYL, and the results proved the effectiveness of this method. In conclusion, CRPBsites is an effective prediction model for circRNA-RBP interaction site identification. We hope that CRPBsites can provide valuable guidance for experimental studies on the influence of circRNA on post-transcriptional regulation. Zhengfeng Wang, Xiujuan Lei |
Briefings Bioinform. | 2 |
| 2021 | Identifying the sequence specificities of circRNA-binding proteins based on a capsule network architectureabstractBACKGROUND: Circular RNAs (circRNAs) are widely expressed in cells and tissues and are involved in biological processes and human diseases. Recent studies have demonstrated that circRNAs can interact with RNA-binding proteins (RBPs), which is considered an important aspect for investigating the function of circRNAs. RESULTS: In this study, we design a slight variant of the capsule network, called circRB, to identify the sequence specificities of circRNAs binding to RBPs. In this model, the sequence features of circRNAs are extracted by convolution operations, and then, two dynamic routing algorithms in a capsule network are employed to discriminate between different binding sites by analysing the convolution features of binding sites. The experimental results show that the circRB method outperforms the existing computational methods. Afterwards, the trained models are applied to detect the sequence motifs on the seven circRNA-RBP bound sequence datasets and matched to known human RNA motifs. Some motifs on circular RNAs overlap with those on linear RNAs. Finally, we also predict binding sites on the reported full-length sequences of circRNAs interacting with RBPs, attempting to assist current studies. We hope that our model will contribute to better understanding the mechanisms of the interactions between RBPs and circRNAs. CONCLUSION: In view of the poor studies about the sequence specificities of circRNA-binding proteins, we designed a classification framework called circRB based on the capsule network. The results show that the circRB method is an effective method, and it achieves higher prediction accuracy than other methods. Zhengfeng Wang, Xiujuan Lei |
BMC Bioinform. | 2 |
| 2021 | Logistic regression algorithm to identify candidate disease genes based on reliable protein-protein interaction network
Xiujuan Lei, Wenxiang Zhang |
Sci. China Inf. Sci. | 1 |
| 2021 | Predicting circRNA-disease associations based on autoencoder and graph embedding
Jing Yang 0063, Xiujuan Lei |
Inf. Sci. | 2 |
| 2021 | Predicting CircRNA-Disease Associations Based on Improved Weighted Biased Meta-Structure
Xiujuan Lei, Chen Bian, Yi Pan 0001 |
J. Comput. Sci. Technol. | 1 |
| 2021 | Prediction of disease-associated circRNAs via circRNA-disease pair graph and weighted nuclear norm minimization
Yuchen Zhang 0003, Xiujuan Lei, Yi Pan 0001, Witold Pedrycz |
Knowl. Based Syst. | 2 |
| 2021 | Human Protein Complex-Based Drug Signatures for Personalized Cancer MedicineabstractDisease signature-based drug repositioning approaches typically first identify a disease signature from gene expression profiles of disease samples to represent a particular disease. Then such a disease signature is connected with the drug-induced gene expression profiles to find potential drugs for the particular disease. In order to obtain reliable disease signatures, the size of disease samples should be large enough, which is not always a single case in practice, especially for personalized medicine. On the other hand, the sample sizes of drug-induced gene expression profiles are generally large. In this study, we propose a new drug repositioning approach (HDgS), in which the drug signature is first identified from drug-induced gene expression profiles, and then connected to the gene expression profiles of disease samples to find the potential drugs for patients. In order to take the dependencies among genes into account, the human protein complexes (HPC) are used to define the drug signature. The proposed HDgS is applied to the drug-induced gene expression profiles in LINCS and several types of cancer samples. The results indicate that the HPC-based drug signature can effectively find drug candidates for patients and that the proposed HDgS can be applied for personalized medicine with even one patient sample. Fei Wang 0095, Yulian Ding, Xiujuan Lei, Bo Liao 0001, Fang-Xiang Wu |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Graph Convolution Networks Using Message Passing and Multi-Source Similarity Features for Predicting circRNA-Disease AssociationabstractGraphs can be used to effectively represent complex data structures. Learning these irregular data in graphs is challenging and still suffers from shallow learning. Applying deep learning on graphs has demonstrated good performance in many applications including social analysis, bioinformatics etc. Message passing graph convolution network is a powerful method which has expressive power to learn graph structures. Meanwhile, circular ribonucleic acid (circRNA) is a type of non-coding RNA which plays a critical role in human diseases. Identifying the associations between circRNAs and diseases is important for diagnosis and treatment of complex diseases. However, there are limited number of known associations between them and conducting biological experiments to identify new associations is time consuming and expensive. As a result, there is a need of building efficient and feasible computation methods to predict potential circRNA-disease associations. In this paper, we propose a novel graph convolution network framework to learn features from a graph built with multi-source similarity information to predict circRNA-disease associations. First we use multi-source information of circRNA similarity, disease and circRNA Gaussian Interaction Profile (GIP) kernel similarity to extract the features using first graph convolution. Then we predict disease associations for each circRNA with a second graph convolution. Proposed framework with five-fold cross validation on various experiments shows promising results in predicting circRNA-disease association and outperforms other existing methods. Thosini Bamunu Mudiyanselage, Xiujuan Lei, Nipuna Senanayake, Yan-Qing Zhang 0001, Yi Pan 0001 |
BIBM | 2 |
| 2020 | Matrix factorization with neural network for predicting circRNA-RBP interactionsabstractBACKGROUND: Circular RNA (circRNA) has been extensively identified in cells and tissues, and plays crucial roles in human diseases and biological processes. circRNA could act as dynamic scaffolding molecules that modulate protein-protein interactions. The interactions between circRNA and RNA Binding Proteins (RBPs) are also deemed to an essential element underlying the functions of circRNA. Considering cost-heavy and labor-intensive aspects of these biological experimental technologies, instead, the high-throughput experimental data has enabled the large-scale prediction and analysis of circRNA-RBP interactions. RESULTS: A computational framework is constructed by employing Positive Unlabeled learning (P-U learning) to predict unknown circRNA-RBP interaction pairs with kernel model MFNN (Matrix Factorization with Neural Networks). The neural network is employed to extract the latent factors of circRNA and RBP in the interaction matrix, the P-U learning strategy is applied to alleviate the imbalanced characteristics of data samples and predict unknown interaction pairs. For this purpose, the known circRNA-RBP interaction data samples are collected from the circRNAs in cancer cell lines database (CircRic), and the circRNA-RBP interaction matrix is constructed as the input of the model. The experimental results show that kernel MFNN outperforms the other deep kernel models. Interestingly, it is found that the deeper of hidden layers in neural network framework does not mean the better in our model. Finally, the unlabeled interactions are scored using P-U learning with MFNN kernel, and the predicted interaction pairs are matched to the known interactions database. The results indicate that our method is an effective model to analyze the circRNA-RBP interactions. CONCLUSION: For a poorly studied circRNA-RBP interactions, we design a prediction framework only based on interaction matrix by employing matrix factorization and neural network. We demonstrate that MFNN achieves higher prediction accuracy, and it is an effective method. Zhengfeng Wang, Xiujuan Lei |
BMC Bioinform. | 2 |
| 2020 | Relational completion based non-negative matrix factorization for predicting metabolite-disease associations
Xiujuan Lei, Jiaojiao Tie, Hamido Fujita |
Knowl. Based Syst. | 1 |
| 2020 | A decomposition-based evolutionary algorithm with adaptive weight adjustment for many-objective problems
Cai Dai, Xiujuan Lei, Xiaoguang He |
Soft Comput. | 2 |
| 2020 | Artificial Fish Swarm Optimization Based Method to Identify Essential ProteinsabstractIt is well known that essential proteins play an extremely important role in controlling cellular activities in living organisms. Identifying essential proteins from protein protein interaction (PPI) networks is conducive to the understanding of cellular functions and molecular mechanisms. Hitherto, many essential proteins detection methods have been proposed. Nevertheless, those existing identification methods are not satisfactory because of low efficiency and low sensitivity to noisy data. This paper presents a novel computational approach based on artificial fish swarm optimization for essential proteins prediction in PPI networks (called AFSO_EP). In AFSO_EP, first, a part of known essential proteins are randomly chosen as artificial fishes of priori knowledge. Then, detecting essential proteins by imitating four principal biological behaviors of artificial fishes when searching for food or companions, including foraging behavior, following behavior, swarming behavior, and random behavior, in which process, the network topology, gene expression, gene ontology (GO) annotation, and subcellular localization information are utilized. To evaluate the performance of AFSO_EP, we conduct experiments on two species (Saccharomyces cerevisiae and Drosophila melanogaster), the experimental results show that our method AFSO_EP achieves a better performance for identifying essential proteins in comparison with several other well-known identification methods, which confirms the effectiveness of AFSO_EP. Xiujuan Lei, Xiaoqin Yang, Fang-Xiang Wu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | Introducing Heuristic Information Into Ant Colony Optimization Algorithm for Identifying EpistasisabstractEpistasis learning, which is aimed at detecting associations between multiple Single Nucleotide Polymorphisms (SNPs) and complex diseases, has gained increasing attention in genome wide association studies. Although much work has been done on mapping the SNPs underlying complex diseases, there is still difficulty in detecting epistatic interactions due to the lack of heuristic information to expedite the search process. In this study, a method EACO is proposed to detect epistatic interactions based on the ant colony optimization (ACO) algorithm, the highlights of which are the introduced heuristic information, fitness function, and a candidate solutions filtration strategy. The heuristic information multi-SURF* is introduced into EACO for identifying epistasis, which is incorporated into ant-decision rules to guide the search with linear time. Two functionally complementary fitness functions, mutual information and the Gini index, are combined to effectively evaluate the associations between SNP combinations and the phenotype. Furthermore, a strategy for candidate solutions filtration is provided to adaptively retain all optimal solutions which yields a more accurate way for epistasis searching. Experiments of EACO, as well as three ACO based methods (AntEpiSeeker, MACOED, and epiACO) and four commonly used methods (BOOST, SNPRuler, TEAM, and epiMODE) are performed on both simulation data sets and a real data set of age-related macular degeneration. Results indicate that EACO is promising in identifying epistasis. Yingxia Sun, Junliang Shang, Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Xiujuan Lei |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2019 | PDG-PIO: Predicting Disease-genes Based on Pigeon-inspired OptimizationabstractCombining large-scale biological data, using computational methods to mine potential disease-gene associations is a popular strategy. At the same time, bio-inspired intelligent optimization has always been a hot research field of intelligent computing. In this study, we apply the pigeon-inspired optimization (PIO) algorithm to the identification of human disease-genes. The problem of predicting disease-genes is translated into a single-objective optimization problem. A reasonable objective function is designed to measure the association between genes and inquiring diseases in a heterogeneous network, and the corresponding probability matrix is generated. The experimental results show that the proposed method (PDG-PIO) can accurately identify disease-genes. Yuchen Zhang 0003, Xiujuan Lei, Shi Cheng 0002 |
CEC | 2 |
| 2019 | Dynamic Multimodal Optimization: A Preliminary StudyabstractThe benchmark problems have played a fundamental role in verifying the algorithm's search ability. A dynamic multimodal optimization (DMO) problem is defined as an optimization problem with multiple global optima and characteristics of global optima which are changed during the search process. Two cases are used to illustrate the application scenario of DMO. A set of benchmark functions on DMO, which contains eight problems, are proposed to show the difficulty of DMO. The properties of the proposed benchmark problems, such as the distribution of solutions, the scalability, the number of global/local optima, are discussed. Shi Cheng 0002, Hui Lu 0002, Yinan Guo 0001, Xiujuan Lei, Jing J. Liang, Yuhui Shi 0001 |
CEC | 4 |
| 2019 | Protein complex detection based on flower pollination mechanism in multi-relation reconstructed dynamic protein networksabstractBACKGROUND: Detecting protein complex in protein-protein interaction (PPI) networks plays a significant part in bioinformatics field. It enables us to obtain the better understanding for the structures and characteristics of biological systems. METHODS: In this study, we present a novel algorithm, named Improved Flower Pollination Algorithm (IFPA), to identify protein complexes in multi-relation reconstructed dynamic PPI networks. Specifically, we first introduce a concept called co-essentiality, which considers the protein essentiality to search essential interactions, Then, we devise the multi-relation reconstructed dynamic PPI networks (MRDPNs) and discover the potential cores of protein complexes in MRDPNs. Finally, an IFPA algorithm is put forward based on the flower pollination mechanism to generate protein complexes by simulating the process of pollen find the optimal pollination plants, namely, attach the peripheries to the corresponding cores. RESULTS: The experimental results on three different datasets (DIP, MIPS and Krogan) show that our IFPA algorithm is more superior to some representative methods in the prediction of protein complexes. CONCLUSIONS: Our proposed IFPA algorithm is powerful in protein complex detection by building multi-relation reconstructed dynamic protein networks and using improved flower pollination algorithm. The experimental results indicate that our IFPA algorithm can obtain better performance than other methods. Xiujuan Lei, Fang-Xiang Wu |
BMC Bioinform. | 1 |
| 2019 | Identifying Cancer genes by combining two-rounds RWR based on multiple biological dataabstractBACKGROUND: It's a very urgent task to identify cancer genes that enables us to understand the mechanisms of biochemical processes at a biomolecular level and facilitates the development of bioinformatics. Although a large number of methods have been proposed to identify cancer genes at recent times, the biological data utilized by most of these methods is still quite less, which reflects an insufficient consideration of the relationship between genes and diseases from a variety of factors. RESULTS: In this paper, we propose a two-rounds random walk algorithm to identify cancer genes based on multiple biological data (TRWR-MB), including protein-protein interaction (PPI) network, pathway network, microRNA similarity network, lncRNA similarity network, cancer similarity network and protein complexes. In the first-round random walk, all cancer nodes, cancer-related genes, cancer-related microRNAs and cancer-related lncRNAs, being associated with all the cancer, are used as seed nodes, and then a random walker walks on a quadruple layer heterogeneous network constructed by multiple biological data. The first-round random walk aims to select the top score k of potential cancer genes. Then in the second-round random walk, genes, microRNAs and lncRNAs, being associated with a certain special cancer in corresponding cancer class, are regarded as seed nodes, and then the walker walks on a new quadruple layer heterogeneous network constructed by lncRNAs, microRNAs, cancer and selected potential cancer genes. After the above walks finish, we combine the results of two-rounds RWR as ranking score for experimental analysis. As a result, a higher value of area under the receiver operating characteristic curve (AUC) is obtained. Besides, cases studies for identifying new cancer genes are performed in corresponding section. CONCLUSION: In summary, TRWR-MB integrates multiple biological data to identify cancer genes by analyzing the relationship between genes and cancer from a variety of biological molecular perspective. Wenxiang Zhang, Xiujuan Lei, Chen Bian |
BMC Bioinform. | 2 |
| 2019 | Detecting overlapping protein complexes in weighted PPI network based on overlay network chain in quotient spaceabstractBACKGROUND: Protein complexes are the cornerstones of many biological processes and gather them to form various types of molecular machinery that perform a vast array of biological functions. In fact, a protein may belong to multiple protein complexes. Most existing protein complex detection algorithms cannot reflect overlapping protein complexes. To solve this problem, a novel overlapping protein complexes identification algorithm is proposed. RESULTS: In this paper, a new clustering algorithm based on overlay network chain in quotient space, marked as ONCQS, was proposed to detect overlapping protein complexes in weighted PPI networks. In the quotient space, a multilevel overlay network is constructed by using the maximal complete subgraph to mine overlapping protein complexes. The GO annotation data is used to weight the PPI network. According to the compatibility relation, the overlay network chain in quotient space was calculated. The protein complexes are contained in the last level of the overlay network. The experiments were carried out on four PPI databases, and compared ONCQS with five other state-of-the-art methods in the identification of protein complexes. CONCLUSIONS: We have applied ONCQS to four PPI databases DIP, Gavin, Krogan and MIPS, the results show that it is superior to other five existing algorithms MCODE, MCL, CORE, ClusterONE and COACH in detecting overlapping protein complexes. Jie Zhao 0012, Xiujuan Lei |
BMC Bioinform. | 2 |
| 2019 | Generalized pigeon-inspired optimization algorithms
Shi Cheng 0002, Xiujuan Lei, Hui Lu 0002, Yong Zhang 0016, Yuhui Shi 0001 |
Sci. China Inf. Sci. | 2 |
| 2019 | Predicting the associations between microbes and diseases by integrating multiple data sources and path-based HeteSim scores
Chunyan Fan, Xiujuan Lei, Aidong Zhang 0001 |
Neurocomputing | 2 |
| 2019 | Predicting disease-genes based on network information loss and protein complexes in heterogeneous network
Xiujuan Lei, Yuchen Zhang 0003 |
Inf. Sci. | 1 |
| 2019 | Moth-flame optimization-based algorithm with synthetic dynamic PPI networks for discovering protein complexes
Xiujuan Lei, Hamido Fujita |
Knowl. Based Syst. | 1 |
| 2019 | Random walk based method to identify essential proteins by integrating network topology and biological characteristics
Xiujuan Lei, Xiaoqin Yang, Hamido Fujita |
Knowl. Based Syst. | 1 |
| 2018 | Two-step Random Walk Algorithm to Identify Cancer Genes Based on Various Biological Data
Wenxiang Zhang, Xiujuan Lei |
BIBM | 2 |
| 2018 | Mining Overlapping Protein Complexes in PPI Network Based on Granular Computation in Quotient Space
Jie Zhao 0012, Xiujuan Lei |
ICIC (1) | 2 |
| 2018 | Topology potential based seed-growth method to identify protein complexes on dynamic PPI data
Xiujuan Lei, Yuchen Zhang 0003, Shi Cheng 0002, Fang-Xiang Wu, Witold Pedrycz |
Inf. Sci. | 1 |
| 2018 | Predicting essential proteins based on RNA-Seq, subcellular localization and GO annotation datasets
Xiujuan Lei, Jie Zhao 0012, Hamido Fujita, Aidong Zhang 0001 |
Knowl. Based Syst. | 1 |
| 2017 | A comprehensive survey of brain storm optimization algorithmsabstractThe development, implementation, variant, and future directions of a new swarm intelligence algorithm, brain storm optimization (BSO) algorithm, are comprehensively surveyed. Brain storm optimization algorithm is a new and promising swarm intelligence algorithm, which simulates the human brainstorming process. Through the convergent operation and divergent operation, individuals in BSO are grouped and diverged in the search space/objective space. To the best of our knowledge, there are 75 papers, 8 theses, and 5 patents in total on the development and application of the BSO algorithm. Every individual in the BSO algorithm is not only a solution to the problem to be optimized, but also a data point to reveal the landscape of the problem. Based on the developments of brain storm optimization algorithms, different kinds of optimization problems and real-world applications could be solved. Shi Cheng 0002, Yifei Sun 0004, Quande Qin, Xianghua Chu, Xiujuan Lei, Yuhui Shi 0001 |
CEC | 6 |
| 2017 | Genome-Wide Identification of Essential Proteins by Integrating RNA-seq, Subcellular Location and Complexes Information
Chunyan Fan, Xiujuan Lei |
ICIC (2) | 2 |
| 2017 | An improvement decomposition-based multi-objective evolutionary algorithm with uniform design
Cai Dai, Xiujuan Lei |
Knowl. Based Syst. | 2 |
| 2016 | Detecting protein complexes from DPINs by OPTICS based on particle swarm optimizationabstractDetecting protein complexes has become an important area of system biology for revealing cellular organization and function. It has been indicated that dense sub-networks in protein-protein interaction (PPI) network, especially dynamic PPI network (DPIN), usually correspond to protein complexes by plenty evidences. In this study, we develop a new approach, named OPTICS_PSO, which combines two algorithms for clustering DPINs data: the Ordering Points to Identify the Clustering Structure (OPTICS) algorithm, for identifying protein complexes in dynamic PPI network, and the particle swarm optimization (PSO) algorithm used to optimize the parameter ε in OPTICS when clustering sub-networks. In the DPIN, all sub-networks have different scales. Although OPTICS is effective in clustering, its parameters are user-specified and cannot deal with networks with different scales. We adopt PSO algorithm to optimize OPTICS and adjust its parameters. The identified protein complexes on the DIP dataset and Krogan dataset show that the new algorithm outperforms the state-of-the-art approaches in terms of several criteria such as precision and f-measure. Xiujuan Lei, Fang-Xiang Wu |
BIBM | 1 |
| 2016 | Mining protein complexes based on topology potential from weighted dynamic PPI networkabstractIdentification of protein complexes is very important to investigate the characteristics of biological processes. Most of existing protein complex clustering algorithms were often run only on a static protein-protein interaction (PPI) network. The dynamic characteristics of interactions were ignored. In order to solve the problem, a new clustering algorithm (TP-WDPIN) was proposed which is based on the concept of topological potential to measure the importance of proteins in the process of detecting seed proteins and then to mine protein complexes from weighted dynamic PPI network. The algorithm used features of core-attachment of complexes and split low density cores to improve density of cores for achieving better clustering results. Experiment results showed that the proposed TP-WDPIN algorithm has better performance than other algorithms on two PPI databases. Xiujuan Lei, Yuchen Zhang 0003, Fang-Xiang Wu, Aidong Zhang 0001 |
BIBM | 1 |
| 2016 | Identifying protein complexes in dynamic protein-protein interaction networks based on Cuckoo Search algorithmabstractProtein complexes play a critical role in understanding the function of cell machinery. The existing protein complex detection algorithms are mostly cannot reflect the dynamics of protein complexes. In this paper, a novel algorithm named cuckoo search clustering algorithm (CSCA) is proposed to detect protein complexes in dynamic protein-protein interaction networks (DPIN) inspired by cuckoo search (CS) mechanism. First, we constructed dynamic protein networks and detected protein complex cores in every dynamic sub-network. Then, CS was used to cluster the protein attachments to the cores. The experimental results on DIP dataset and Krogan dataset demonstrated that CSCA is more effective to identify protein complexes than other typical methods. Jie Zhao 0012, Xiujuan Lei, Fang-Xiang Wu |
BIBM | 2 |
| 2016 | A decomposition based evolutionary algorithm with uniform design for multi-objective optimizationabstractThe diversity and convergence of obtained solutions are two main goals for multi-objective evolutionary algorithms. In this paper, a new decomposition based evolutionary algorithm with uniform design (MOEA/DU) is designed to achieve these two goals. Firstly, the objective space of a multi-objective problem is decomposed into a set of sub-regions based on a set of direction vectors, and each sub-region is made to have a solution for maintaining the diversity. Secondly, for domination solutions, a selection strategy and a crossover operator based on uniform design are used to make these solutions as soon as possibly become non-domination solutions. The proposed algorithm has been compared with NSGAII, MOEA/D and MOEA/D-M2M on seven test instances. The experimental results illustrate that the proposed algorithm is able to find a set of solutions with better diversity and convergence. Cai Dai, Xiujuan Lei, Yulian Ding |
CEC | 2 |
| 2016 | A Novel Fitness Function Based on Decomposition for Multi-objective Optimization Problems
Cai Dai, Xiujuan Lei, Xiaofang Guo |
ICIC (2) | 2 |
| 2016 | Detecting protein complexes from DPINs by density based clustering with Pigeon-Inspired Optimization Algorithm
Xiujuan Lei, Yulian Ding, Fang-Xiang Wu |
Sci. China Inf. Sci. | 1 |
| 2016 | Protein complex identification through Markov clustering with firefly algorithm on dynamic protein-protein interaction networks
Xiujuan Lei, Fei Wang 0095, Fang-Xiang Wu, Aidong Zhang 0001, Witold Pedrycz |
Inf. Sci. | 1 |
| 2016 | Identification of dynamic protein complexes based on fruit fly optimization algorithm
Xiujuan Lei, Yulian Ding, Hamido Fujita, Aidong Zhang 0001 |
Knowl. Based Syst. | 1 |
| 2014 | Detecting functional modules in dynamic protein-protein interaction networks using Markov Clustering and Firefly AlgorithmabstractMarkov Clustering (MCL) is a popular algorithm for clustering networks in bioinformatics such as Protein-Protein Interaction (PPI) networks and especially, shows excellent performance in clustering Dynamic Proteinprotein Interaction Networks (DPIN). However, a limitation of MCL and its variants (e.g. regularized MCL and soft regularized MCL) is that the clustering results are mostly dependent on the parameters that user-specified. However we know that different networks with various scales need different parameters. In this article, we propose a new MCL method based on the Firefly Algorithm (FA) to optimize its parameters. The results on DIP dataset show that the new algorithm outperforms the state-of-the-art approaches in terms of accuracy of identifying functional modules on a real DPIN. Xiujuan Lei, Fang-Xiang Wu, Fei Wang 0095, Aidong Zhang 0001 |
BIBM | 1 |
| 2014 | A simplified glowworm swarm optimization algorithmabstractAimed at the poor optimizing ability and the low accuracy of the glowworm swarm optimization algorithm (GSO), a simplified glowworm swarm optimization algorithm (SGSO) was put forward in this paper, which omitted the phases of seeking dynamic decision domain and movement probability calculation, and meanwhile simplified the location updating process. Moreover, elitism was introduced to improve the capacity of searching optimal solution. It was applied to the unimodal and multimodal benchmark function optimization problems. The improved SGSO algorithm is compared with the basic GSO and other swarm intelligent optimization algorithms to demonstrate the performance. Experimental results showed that SGSO improves not only the precision but also the efficiency in function optimization. Mingyu Du, Xiujuan Lei, Zhenqiang Wu |
IEEE Congress on Evolutionary Computation | 2 |
| 2013 | PPI modules detection method through ABC-IFC algorithmabstractA novel clustering model is proposed which combines the optimization mechanism of artificial bee colony (ABC) with the fuzzy membership matrix in this paper. The clustering model contains two parts: one is to search optimum cluster centers using ABC mechanism, the other is to implement clustering using intuitionistic fuzzy clustering (IFC) method. Firstly, the cluster centers are set randomly and the initial clustering results are obtained using fuzzy membership matrix. The new cluster centers are updated with the nodes that contain the maximal amount of information in the previous clusters of onlookers by ABC algorithm. If the onlookers are incapable of updating, the scouts will generate new cluster centers via global searching. Then the clustering result is obtained through IFC method based on the new optimized cluster centers. Considering that some protein nodes in PPI networks are unreachable, which leads to the traditional distance based clustering criteria infeasible. Therefore the new objective function is designed. The improved algorithm, named ABC-IFC, is also compared with the traditional fuzzy C-means clustering and IFC method. The experimental results on MIPS dataset show that the new algorithm does not only get improved in terms of several commonly used evaluation criteria such as precision, recall and P-value, but also obtains a better clustering result. Xiujuan Lei, Jianfang Tian, Fang-Xiang Wu |
BIBM | 1 |
| 2013 | The clustering model and algorithm of PPI network based on propagating mechanism of artificial bee colony
Xiujuan Lei, Jianfang Tian, Aidong Zhang 0001 |
Inf. Sci. | 1 |
| 2011 | Clustering PPI Data Based on Bacteria Foraging Optimization AlgorithmabstractThis paper proposed a novel method using Bacteria Foraging Optimization(BFO) algorithm to avoid the influence of cluster number on experimental result of clustering PPI networks. The initial position that the bacterium located in was considered to be the cluster center and the positions that the bacterium moved were regarded as the adjacent nodes of cluster center. The algorithm classified the nodes selected in the chemotactic operation into cluster when executing the reproduction and elimination-dispersal operations. The procedure kept on creating new clusters until all the nodes were grouped into the clusters. The simulation result showed that the algorithm not only effectively improved the accuracy of cluster result, but also automatically determined the cluster number. Xiujuan Lei, Aidong Zhang 0001 |
BIBM | 1 |
| 2008 | The aircraft departure scheduling based on particle swarm optimization combined with simulated annealing algorithmabstractParticle swarm optimization combined with simulated annealing algorithm (PSOCSA) was an improved particle swarm optimization algorithm which introduced the simulated annealing (SA) strategy in particle swarm optimization (PSO). It was proposed to solve a mathematical model which is built for aircraft departure sequencing problem in this paper. The correlative implementation techniques and detailed design process of the algorithm were presented. Then the simulation is performed to solve a representative problem using PSOCSA, PSO, and SA. The comparison showed that the PSOCSA algorithm was rational and feasible and more easily converge to the global optimal solution of aircraft departure sequencing problem. Method described in this paper will curtail the consumption of aircraft departure effectively, so it is worth researching it further in the field of airport operations and air traffic control. Fu Ali, Xiujuan Lei |
IEEE Congress on Evolutionary Computation | 2 |
| 2008 | The aircraft departure scheduling based on second-order oscillating particle swarm optimization algorithmabstractThe second-order oscillating particle swarm optimization(SO-PSO) algorithm, which introduced the second-order oscillating evolutionary equation to the evolutionary equation of PSO, could adjust the particles’ global and local search capability and avoid the local optimization. It was proposed to solve a mathematical model which was built for aircraft departure sequencing problem in this paper. The correlative implementation techniques and detailed design process of the algorithm were presented Then the simulation was performed to solve this sequencing problem using the SO-PSO algorithm. The results showed that the global optimal solution was obtained, so the SO-PSO algorithm was rational and feasible and curtailed the consumption of aircraft departure effectively. Xiujuan Lei, Fu Ali, Zhongke Shi |
IEEE Congress on Evolutionary Computation | 1 |
| 2008 | Air robot path planning based on Intelligent Water Drops optimizationabstractPath planning of air robot is a complicated global optimum problem. Intelligent water drops (IWD) algorithm is newly presented under the inspiration of the dynamic of river systems and the actions that water drops do in the rivers, and it is easy to combine with other methods in optimization. In this paper, we propose an improved IWD optimization algorithm for solving the air robot path planning problems in various environments. The water drops can act as an agent in searching the optimal path. The detailed realization procedure for this novel approach is also presented. Series experimental comparison results show the proposed IWD optimization algorithm is more effective and feasible in the air robot path planning than the basic IWD model. Haibin Duan, Senqi Liu, Xiujuan Lei |
IJCNN | 3 |