Feng Li 0033

dblp:92/2954-33 · DBLP profile ↗
← Back
57ranked-venue papers
1as first author
52since 2021 · last 2026
0000-0002-5556-3789ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 49 · 1 first-author · 44 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Predicting the Drug Side Effect Frequency via a Kolmogorov-Arnold Graph Isomorphism Network and Subspace-Aware Dynamic Feature Fusion
Chenglong Mi, Feng Li 0033, Ling-Yun Dai
ICIC (3)3
2026 A multi-objective multi-stage genetic algorithm for community detection in biological networks
Mingyuan Bi, Junliang Shang, Yahan Li, Feng Li 0033, Jin-Xing Liu 0001
Future Gener. Comput. Syst.6
2026 A Hierarchical Attention-Based Negative Sampling Method for Drug Repositioning Using Neighborhood Interaction Fusion
abstract
Accurate prediction of drug-disease associations (DDAs) is essential for drug repositioning and the development of novel therapeutic strategies. However, existing methods often suffer from limited prior knowledge and the use of oversimplified negative sampling techniques, which hinder their ability to capture the complex relationships between drugs and diseases. To break through these limitations, we propose a new model, Hierarchical Attention Mechanism-Based Negative Sampling (HA-NegS), which aims to enhance the prediction of potential DDAs. In this study, HA-NegS further computes the similarity information between drugs and diseases and constructs heterogeneous and homogeneous networks based on it. For the similarity network, HA-NegS fuses Graph Convolutional Network (GCN) and Graph Attention Network (GAT) to effectively capture the neighborhood features of the target nodes. Subsequently, the model incorporates a hierarchical sampling strategy using the PageRank algorithm to rank nodes in descending order of global importance. The attention mechanism is then used to calculate the attention score and re-rank the nodes accordingly. This approach ensures the reliability of the negative sample selection. In order to obtain optimized representations, we use graph contrastive learning methods to refine drug and disease features with homogeneous and heterogeneous neighborhood information. Experimental results on a benchmark dataset show that HA-NegS outperforms existing baseline methods in predicting DDA. In addition, case studies for Alzheimer's disease and Parkinson's disease highlight the effectiveness of HA-NegS in discovering new therapeutic applications for existing drugs.
Cheng-Long Mi, Ling-Yun Dai, Junliang Shang, Juan Wang 0003, Feng Li 0033
IEEE J. Biomed. Health Informatics6
2026 Epileptic Seizure Prediction Using Multi-Strategy Data Augmentation and Hierarchical Contrastive Learning
abstract
Accurate early prediction of epileptic seizures is crucial for improving patients' quality of life. However, existing seizure prediction methods often rely on large-scale labeled datasets and face challenges in generalization and real-time performance. To address these issues, this study proposes an efficient seizure prediction framework that achieves high performance even with limited labeled data, significantly reducing dependence on extensive annotations. To better distinguish preictal states, contrastive learning is employed to enhance feature separation between interictal and preictal periods, leading to improved sensitivity in detecting early seizure patterns. First, a data augmentation strategy is designed, incorporating wavelet-based frequency mixing, temporal masking, and window-based masking to enhance model robustness and generalization. Second, a hierarchical contrastive loss function is introduced, integrating instance-level and temporal contrastive learning to improve the model's ability to capture preictal patterns. Finally, a lightweight SE-EEGNet is developed and optimized as a feature extractor, strengthening critical feature extraction and enabling real-time seizure prediction. On the CHB-MIT dataset, the proposed method achieves 94.51% accuracy, 95.05% sensitivity, a 0.024/h false positive rate (FPR), and a 20.12-minute prediction time using only 30% labeled data. On the Siena dataset, it achieves 93.14% accuracy, 92.77% sensitivity, and a 0.030/h FPR. Moreover, performance improves further as the amount of labeled data increases, validating the effectiveness and practical applicability of the proposed approach in seizure prediction.
Longfei Qi, Feng Li 0033, Junliang Shang, Shihan Wang 0009, Shasha Yuan
IEEE J. Biomed. Health Informatics2
2025 MGAMDA: Multi Source Similarity Fusion-Based Graph Convolutional Neural Network and Attention Mechanism Network for Predicting MiRNA-Disease Associations
abstract
A mounting body of research indicates that dysregulation of MicroRNAs (miRNAs) causes disease through a variety of underlying mechanisms. Predicting microRNA (miRNA)-disease associations (MDAs) is essential for disease prognosis and therapeutics. Compared to conventional biological experiments, computational models save time and effort. A new method is proposed inspired by the graph convolutional networks. It has been named Multi source similarity fusion-based graph convolutional neural network and attention mechanism network for predicting miRNA-disease associations (MGAMDA). First, the several similarity networks between miRNAs and diseases were built. Then, multi-source information network is fused. And the feature was aggregated by using GCNs. In order to address the different levels of importance of the information, an attention mechanism was used to assign weights. The similar features of the disease side and miRNA side were finally obtained separately. It is combined with the association features that are obtained from the association information, and then it is fed into the multi-layer perceptron (MLP). To obtain prediction scores for unknown associations between miRNAs and diseases, a multilayer perceptron was utilized. To validate the new methodology's effectiveness, we performed a series of experimental studies using the Human MicroRNA Disease Database (HMDD v3.2). The performance of the$\mathbf{5}$-fold cross-validation on the datasets shows that MGAMDA surpasses other methods in the area of AUC, AUPR, ACC, F1-score, Recall, and Precision. Furthermore, case studies have demonstrated that MGAMDA accurately predicts miRNAs associated with colon, breast, and stomach cancer.
Ling-Yun Dai, Cheng-Long Mi, Juan Wang 0003, Feng Li 0033
BIBM6
2025 TAGCL-DDI: Two-Stage Augmentation-Driven Graph Contrastive Learning for Drug-Drug Interaction Prediction
abstract
Drug-drug interactions (DDIs) occur when the simultaneous administration of multiple drugs alters their pharmacological effects or causes adverse reactions. Existing DDIs prediction methods typically extract features of drug pairs from a DDI event graph and utilize a multilayer perceptron classifier to predict interactions. However, relying solely on the single DDI event graph limits the encoder's ability to capture comprehensive drug features. To overcome this limitation, we propose TAGCLDDI, a novel Two-stage Augmentation-driven Graph Contrastive Learning framework enhance drug feature representation. In the first stage, TAGCL-DDI employs the multi-head graph attention network to adaptively learn edge weights within the DDI event graph, and then generates multiple augmented graph views by mapping the graph into diverse subspaces through different encoders, enabling the capture of richer interaction information. In the second stage, node-level contrastive learning further refines feature representations across subspaces. Experimental results on two benchmark datasets show that TAGCL-DDI achieves state-of-the-art performance in DDIs prediction.
Jiquan Zhao, Feng Li 0033, Yuzhuo Yuan, Wentian Xin, Shasha Yuan
BIBM3
2025 Adaptive Weighting Contrastive Learning for Spatial Domain Identification in Spatial Transcriptomics
abstract
The rapid advancement of spatial transcriptomics has enabled the joint analysis of gene expression and spatial location data. This integration opens new avenues for uncovering tissue heterogeneity. Existing methods attempt to combine spatial and expression information to identify spatial domains. However, they often treat all information sources equally and do not account for their varying impact on results. To address this challenge, we propose AWCST, an adaptive weighted contrastive learning framework for spatial domain identification. AWCST first extracts spatial and expression latent representations and then fuses them using a multi-head attention mechanism. It measures the distributional differences between each view and the fused feature using Maximum Mean Discrepancy. These differences are converted into adaptive weights for the contrastive loss, enhancing the influence of high-quality information sources. Finally, we evaluate AWCST on two independent datasets to demonstrate its effectiveness.
Xiyue Li, Ling-Yun Dai, Junliang Shang, Feng Li 0033
BIBM5
2025 An Adaptive Single-Cell Sequencing Data Cluster Method Under Weight Fusion Constraint
abstract
Single-cell RNA sequencing (scRNA-seq) provides the transcriptome of a single cell, allowing researchers to study cellular phenomena at a higher resolution level. Nevertheless, noise generated by technical limitations and other results seriously interferes with the downstream analysis of sequencing data such as clustering. How to minimize the impact of noise on the accuracy of clustering methods has become a focus of current research. In this case, we propose a novel cell clustering algorithm called low-rank representation constrained clustering based on noise weight fusion (LRBNW). First, we mitigate the noise interference by introducing a noise weight matrix and assigning different weights to the noise through a reliability assessment strategy. By assigning larger weights to smaller reconstruction errors, we then highlight useful features with small errors, which clean features more representative in data analysis. Finally, we impose a k-block diagonal constraint on the affinity matrix through a block strategy, grouping related features into the same block, to eliminate redundancy among them and avoid over-reliance on related features. Extensive experiments demonstrate that LRBNW achieves higher accuracy results than existing state-of-the-art clustering methods on 10 real scRNA-seq datasets. In addition, downstream analysis experiments also indicated that LRBNW can identify biologically significant groups and reduce noise interference in scRNA-seq data. The result proves LRBNW is a powerful cell type identification tool, and has potential in predicting new cell types.
Zhenchang Wang, Shasha Yuan, Feng Li 0033, Juan Wang 0003
BIBM3
2025 Spatial Multi-Omics Integration Via Information-Aware Multi-View Contrastive Learning
abstract
The rapid advancement of spatial multi-omics technology enables the simultaneous acquisition of diverse expression data from the same tissue or slice. Different omics offer unique and critical information about the biological system. However, most existing methods are unable to fully utilize this information for downstream tasks such as spatial domain identification. To integrate this information effectively for downstream analysis, we introduce a novel Spatial Multi-omics data integration method based on Information-Aware Multi-view Contrastive Learning (SM-IAMCL). It optimizes the spatial and feature neighborhood graphs for each omics by the specific graph learner and fused graph learner, and learns the fused graph of spatial and feature neighborhood graphs at the same time. Then, to make fused graph of each omics integrate both shared and unique information of spatial and feature neighborhood graphs, we incorporate graph-level contrastive learning between different views in each omics. Finally, the learned fused representation of each omics is then integrated via a weighted fusion strategy to generate an integrated low-dimensional latent representation of spatial multiomics. This integrated representation is used for a variety of downstream analysis tasks. The experimental results show that SM-IAMCL outperforms other seven existing methods in the downstream tasks such as spatial domain identification.
Conghui Zhang, Ling-Yun Dai, Juan Wang 0003, Junliang Shang, Feng Li 0033
BIBM6
2025 Predicting Potential Associations Between Microbes and Diseases Using Graph Attention Auto-encoder and PU Learning
Ling-Yun Dai, Feng Li 0033
ICIC (27)3
2025 Label-Guided Graph Contrastive Learning for Single-Cell Fusion Clustering
Baojuan Qin, Junliang Shang, Yan Zhao 0045, Feng Li 0033, Jin-Xing Liu 0001
ISBRA (1)5
2025 A Neighborhood Selection Learning Artificial Bee Colony Algorithm Based on Population Backtracking for Detecting Epistatic Interactions
Xiaoqi Tang, Linqian Zhao, Junliang Shang, Feng Li 0033, Jin-Xing Liu 0001
ISBRA (1)6
2025 SUIFS: A Symmetric Uncertainty Based Interactive Feature Selection Method
Junliang Shang, Qianqian Ren, Feng Li 0033
ISBRA (1)6
2025 PDA-GTGCN: Identification of PiRNA-Disease Associations Based on Group Feature Transformation Graph Convolutional Network
Xiaoqi Tang, Xianghan Meng, Junliang Shang, Baojuan Qin, Feng Li 0033
ISBRA (1)8
2025 STDDAE: Identifying spatial domains in spatial transcriptomics by dual denoising autoencoder with attention mechanism
abstract
Spatial transcriptomics provides a novel perspective for comprehending the intricate relationship between tissue structure and function, as well as for discovering new cell types and subtypes. However, it remains a significant challenge to accurately identify spatial domains with similar gene expression, which requires efficient combination of gene expression data , histology image information, and spatial location. To address this challenge, a novel dual denoising autoencoder with attention mechanism (STDDAE) is proposed. STDDAE integrates gene expression data, histology image information and spatial location, and the decoder consists of a master decoder and a follower decoder, which are jointly optimized to generate low-dimensional latent embeddings for precise spatial domain identification. The performance of STDDAE was evaluated across four datasets with varying resolutions and platforms. The experimental findings validated that STDDAE outperformed other cutting-edg methods in spatial domain identification, trajectory inference, and data denoising. Additionally, STDDAE successfully detected differentially expressed genes within identified spatial domains, which may be valuable in disease diagnosis, prognostic assessment, and treatment selection.
Ying-Lian Gao, Cui-Na Jiao, Xu-Ran Dou, Feng Li 0033, Jin-Xing Liu 0001
Eng. Appl. Artif. Intell.5
2025 A Modified Transformer Network for Seizure Detection Using EEG Signals
abstract
Seizures have a serious impact on the physical function and daily life of epileptic patients. The automated detection of seizures can assist clinicians in taking preventive measures for patients during the diagnosis process. The combination of deep learning (DL) model with convolutional neural network (CNN) and transformer network can effectively extract both local and global features, resulting in improved seizure detection performance. In this study, an enhanced transformer network named Inresformer is proposed for seizure detection, which is combined with Inception and Residual network extracting different scale features of electroencephalography (EEG) signals to enrich the feature representation. In addition, the improved transformer network replaces the existing Feedforward layers with two half-step Feedforward layers to enhance the nonlinear representation of the model. The proposed architecture utilizes discrete wavelet transform (DWT) to decompose the original EEG signals, and the three sub-bands are selected for signal reconstruction. Then, the Co-MixUp method is adopted to solve the problem of data imbalance, and the processed signals are sent to the Inresformer network for seizure information capture and recognition. Finally, discriminant fusion is performed on the results of three-scale EEG sub-signals to achieve final seizure recognition. The proposed network achieves the best accuracy of 100% on Bonn dataset and the average accuracy of 98.03%, sensitivity of 95.65%, and specificity of 98.57% on the long-term CHB-MIT dataset. Compared to the existing DL networks, the proposed method holds significant potential for clinical research and diagnosis applications with competitive performance.
Wenrong Hu, Juan Wang 0003, Feng Li 0033, Qingwei Jia, Shasha Yuan
Int. J. Neural Syst.3
2025 A Contrastive Learning-Enhanced Residual Network for Predicting Epileptic Seizures Using EEG Signals
abstract
The models used to predict epileptic seizures based on electroencephalogram (EEG) signals often encounter substantial challenges due to the requirement for large, labeled datasets and the inherent complexity of EEG data, which hinders their robustness and generalization capability. This study proposes CLResNet, a framework for predicting epileptic seizures, which combines contrastive self-supervised learning with a modified deep residual neural network to address the above challenges. In contrast to traditional models, CLResNet uses unlabeled EEG data for pre-training to extract robust feature representations. It is then fine-tuned on a smaller labeled dataset to significantly reduce its reliance on labeled data while improving its efficiency and predictive accuracy. The contrastive learning (CL) framework enhances the ability of the model to distinguish between preictal and interictal states, thus improving its robustness and generalizability. The architecture of CLResNet contains residual connections that enable it to learn deep features of the data and ensure an efficient gradient flow. The results of the evaluation of the model on the CHB-MIT dataset showed that it outperformed prevalent methods in the field, with an accuracy of 92.97%, sensitivity of 94.18%, and false-positive rate of 0.043/h. On the Siena dataset, the model also achieved competitive performance, with an accuracy of 92.79%, a sensitivity of 91.47%, and a false-positive rate of 0.041/h. These results confirm the effectiveness of CLResNet in addressing variations in EEG data, and show that contrastive self-supervised learning is a robust and accurate approach for predicting seizures.
Longfei Qi, Shasha Yuan, Feng Li 0033, Junliang Shang, Juan Wang 0003, Shihan Wang 0009
Int. J. Neural Syst.3
2025 stMHCG: High-confidence multi-view clustering for identification of spatial domains from spatially resolved transcriptomics
Junliang Shang, Yan Zhao 0045, Baojuan Qin, Qianqian Ren, Feng Li 0033, Jin-Xing Liu 0001
Neurocomputing6
2025 RPMVCDA: Random Perturbation and Multi-View Graph Convolutional Networks for CircRNA-Disease Association Prediction
abstract
Numerous studies have demonstrated the regulatory role of circular RNA (circRNA) in various diseases, emphasizing the importance of identifying disease-related circRNAs. Although several computational models have been developed to predict circRNA-disease associations, the limited number of experimentally validated associations has resulted in the sparse association network. Therefore, there is a need for continuously improving circRNA-disease prediction models. In this study, we propose RPMVCDA, a computational model based on random perturbation and multi-view graph convolutional networks (GCNs), to predict circRNA-disease associations. Specifically, RPMVCDA first constructs multiple similarity networks of circRNAs and diseases, applying multi-view GCNs to obtain embedding representations. Second, to enable message passing between circRNA-disease samples, RPMVCDA constructs the feature similarity association network. Third, RPMVCDA introduces a random perturbation association network to further explore the potential associations, which is the highlight of the RPMVCDA. Finally, based on these three association networks, RPMVCDA utilizes the self-attention mechanism to generate high-quality features for circRNAs and diseases, which are used to calculate association scores. To evaluate the performance of RPMVCDA, five-fold cross-validation and case studies on the CircR2Disease dataset are performed, results of which shows that RPMVCDA outperforms the compared models, implying that it might be an alternative for predicting circRNA-disease associations.
Xin He 0008, Junliang Shang, Feng Li 0033, Jin-Xing Liu 0001
IEEE Trans. Comput. Biol. Bioinform.4
2025 MPSO-CD: A Multi-Objective Particle Swarm Optimization Community Detection Method for Identifying Disease Modules
abstract
The dysfunction of biological systems caused by disease-related genes is one of the inducements of complex diseases. To understand molecular mechanisms of complex diseases, the identification of disease-related gene modules in biological networks through community detection is emerging as a promising approach. However, most community detection methods are not suitable for biological networks because their topological structures are complex and the scale of biologically relevant modules are small. In this paper, a novel community detection method called MPSO-CD was proposed based on multi-objective particle swarm optimization, in which negative ratio association and ratio cut were employed as objective functions. Highlights of MPSO-CD are a mutation strategy based on clustering coefficient and the procedure of disease module screening referring to the internal connection density and functional similarity. Experimental results of social and synthetic complex networks indicate that MPSO-CD is comparable and often superior to four compared methods. Eventually, MPSO-CD is applied to the asthma gene co-expression network for identifying potential disease modules that provide the molecular mechanism information about asthma. Most of the captured modules have been proven to be associated with asthma through Gene Ontology and pathway enrichment analysis.
Xuhui Zhu, Mingyuan Bi, Junliang Shang, Feng Li 0033, Yuanyuan Zhang 0008, Ling-Yun Dai, Shengjun Li, Jin-Xing Liu 0001
IEEE Trans. Comput. Biol. Bioinform.5
2024 A multi-objective genetic algorithm based on neighborhood coevolution for community detection
abstract
Community detection has attracted growing interest, with multi-objective evolutionary algorithms proving to be highly competitive in this area. In this paper, a community detection method based on a multi-objective neighborhood coevolution genetic algorithm, NCMOGA, is proposed. To improve the computational efficiency in large-scale networks, NCMOGA introduces a network processing strategy to simplify the network before and during evolution. A neighborhood coevolution strategy is proposed, in which the corresponding subpopulation is formed according to the neighborhood of each individual. A series of operations such as crossover, mutation and update are performed in the subpopulation, emphasizing the synergy between individuals and their neighbors. Mating selection and crossover operations are performed based on the center selection idea of density peak clustering, and the most important nodes are selected to generate offspring. The effectiveness of NCMOGA is verified on synthetic networks and real-world networks. In addition, the results in guiding the classification of disease and healthy samples demonstrate the high quality of the modules detected by NCMOGA.
Mingyuan Bi, Junliang Shang, Xiaotong Kong, Feng Li 0033, Yuanyuan Zhang 0008, Jin-Xing Liu 0001
BIBM4
2024 Spatial domains identification based on multi-view contrastive learning in spatial transcriptomics
abstract
Spatial transcriptomic techniques can be used to obtain transcriptome data from different locations in tissues. The identification of spatial domains is a key task in the analysis of spatial transcriptomic data. Therefore, we propose a multi-view contrastive learning framework named MVCLST for identifying spatial domains. First, MVCLST introduces pathway information on the basis of spatial transcriptomic data to consider the functional correlation between spots. In order to better mine the underlying biological features, MVCLST constructs three biological networks from three biological perspectives: spatial distribution, similarity of gene expression and functional correlation. Secondly, MVCLST introduces a multi-view contrastive learning method, which fully considers the contrastive relationship between multiple views to improve the accuracy and reliability of feature extraction. Then, in order to obtain a more abundant feature representation, the adaptive attention mechanism is used to integrate the common features and specific features. Finally, we compare MVCLST with five spatial transcriptomic methods to verify the accuracy of MVCLST for spatial domains identification. The experimental results show that MVCLST is better than the other five methods.
Yanru Gao, Feng Li 0033, Fanhao Meng, Qianqian Ren, Junliang Shang
BIBM2
2024 Integrating Autoencoder and Multi-View Graph Convolutional Networks for Spatial Domains Identification
abstract
Spatial transcriptomics technology provides high-resolution gene expression profiles and spatial location information, offering a revolutionary approach to identify tissue regions and cell types. However, due to the characteristics of gene expression data such as high dimension and high noise, some information may be lost when performing feature extraction. Therefore, we propose a framework that integrates an autoencoder and a multi-view graph convolutional network, IAMGCN, to reduce information loss and capture more comprehensive information. Specifically, we use the autoencoder to extract the deep representation of the gene expression matrix. The middle layer of the autoencoder is integrated into the middle layer of the graph convolutional network. This integration aims to reduce the information loss during the information aggregation process of the graph convolutional network. In order to learn the specificity and common information of multiple views, we introduce a contrastive learning strategy, which takes the output of the autoencoder as a positive sample, and reduce the distance between the two views and the autoencoder output, respectively. We test IAMGCN on two datasets and compare it with five other methods, and the experimental results show that IAMGCN achieves the highest accuracy on all datasets.
Fanhao Meng, Feng Li 0033, Yanru Gao, Junliang Shang
BIBM2
2024 Multi-Population Ant Colony Optimization With Knowledge-Based Local Searches for Epistasis Detection
abstract
Analysis of epistatic interaction is an important means to study the pathogenesis of complex diseases in genome-wide association studies (GWAS). Epistatic interaction detection aims to identify the ideal combination among single nucleotide polymorphisms (SNPs) and determine whether this combination is significantly associated with complex diseases. However, they suffer from certain limitations, such as low detection power and long execution times. Therefore, this paper proposes a multi-population ant colony optimization algorithm with an adaptive heuristic strategy (MPACO-AHS). MPACO-AHS is a framework based on multi-population approaches, where multiple populations are employed to detect epistatic interactions, helping to avoid the data bias inherent in a single population. Moreover, to guide the search direction of each population, an adaptive heuristic strategy is introduced, allowing the algorithm to focus on areas more likely to contain epistatic interactions, thereby improving the accuracy of the results. Comprehensive experiments are conducted on simulated datasets. The results demonstrate that MPACO-AHS outperforms existing algorithms by overcoming the challenges in detecting epistatic interactions in GWAS.
Qianqian Ren, Shaoyi Liu, Lianlian Zhang, Junliang Shang, Feng Li 0033
BIBM5
2024 CPSORCL: A Cooperative Particle Swarm Optimization Method with Random Contrastive Learning for Interactive Feature Selection
Junliang Shang, Yahan Li, Feng Li 0033, Yuanyuan Zhang 0008, Jin-Xing Liu 0001
ISBRA (2)4
2024 A review of recent advances in spatially resolved transcriptomics data analysis
Ying-Lian Gao, Jing Jing 0001, Feng Li 0033, Chun-Hou Zheng 0001, Jin-Xing Liu 0001
Neurocomputing4
2024 SLGCN: Structure-enhanced line graph convolutional network for predicting drug-disease associations
Bao-Min Liu, Ying-Lian Gao, Feng Li 0033, Chun-Hou Zheng 0001, Jin-Xing Liu 0001
Knowl. Based Syst.3
2024 Diagnosis-Guided Deep Subspace Clustering Association Study for Pathogenetic Markers Identification of Alzheimer's Disease Based on Comparative Atlases
abstract
The roles of brain region activities and genotypic functions in the pathogenesis of Alzheimer's disease (AD) remain unclear. Meanwhile, current imaging genetics methods are difficult to identify potential pathogenetic markers by correlation analysis between brain network and genetic variation. To discover disease-related brain connectome from the specific brain structure and the fine-grained level, based on the Automated Anatomical Labeling (AAL) and human Brainnetome atlases, the functional brain network is first constructed for each subject. Specifically, the upper triangle elements of the functional connectivity matrix are extracted as connectivity features. The clustering coefficient and the average weighted node degree are developed to assess the significance of every brain area. Since the constructed brain network and genetic data are characterized by non-linearity, high-dimensionality, and few subjects, the deep subspace clustering algorithm is proposed to reconstruct the original data. Our multilayer neural network helps capture the non-linear manifolds, and subspace clustering learns pairwise affinities between samples. Moreover, most approaches in neuroimaging genetics are unsupervised learning, neglecting the diagnostic information related to diseases. We presented a label constraint with diagnostic status to instruct the imaging genetics correlation analysis. To this end, a diagnosis-guided deep subspace clustering association (DDSCA) method is developed to discover brain connectome and risk genetic factors by integrating genotypes with functional network phenotypes. Extensive experiments prove that DDSCA achieves superior performance to most association methods and effectively selects disease-relevant genetic markers and brain connectome at the coarse-grained and fine-grained levels.
Cui-Na Jiao, Junliang Shang, Feng Li 0033, Xinchun Cui, Yan-Li Wang, Ying-Lian Gao, Jin-Xing Liu 0001
IEEE J. Biomed. Health Informatics3
2024 KFDAE: CircRNA-Disease Associations Prediction Based on Kernel Fusion and Deep Auto-Encoder
abstract
CircRNA has been proved to play an important role in the diseases diagnosis and treatment. Considering that the wet-lab is time-consuming and expensive, computational methods are viable alternative in these years. However, the number of circRNA-disease associations (CDAs) that can be verified is relatively few, and some methods do not take full advantage of dependencies between attributes. To solve these problems, this paper proposes a novel method based on Kernel Fusion and Deep Auto-encoder (KFDAE) to predict the potential associations between circRNAs and diseases. Firstly, KFDAE uses a non-linear method to fuse the circRNA similarity kernels and disease similarity kernels. Then the vectors are connected to make the positive and negative sample sets, and these data are send to deep auto-encoder to reduce dimension and extract features. Finally, three-layer deep feedforward neural network is used to learn features and gain the prediction score. The experimental results show that compared with existing methods, KFDAE achieves the best performance. In addition, the results of case studies prove the effectiveness and practical significance of KFDAE, which means KFDAE is able to capture more comprehensive information and generate credible candidate for subsequent wet-lab.
Wen-Yue Kang, Ying-Lian Gao, Ying Wang 0143, Feng Li 0033, Jin-Xing Liu 0001
IEEE J. Biomed. Health Informatics4
2024 SGFCCDA: Scale Graph Convolutional Networks and Feature Convolution for circRNA-Disease Association Prediction
abstract
Circular RNAs (circRNAs) have emerged as a novel class of non-coding RNAs with regulatory roles in disease pathogenesis. Computational models aimed at predicting circRNA-disease associations offer valuable insights into disease mechanisms, thereby enabling the development of innovative diagnostic and therapeutic approaches while reducing the reliance on costly wet experiments. In this study, SGFCCDA is proposed for predicting potential circRNA-disease associations based on scale graph convolutional networks and feature convolution. Specifically, SGFCCDA integrates multiple measures of circRNA and disease similarity and combines known association information to construct a heterogeneous network. This network is then explored by scale graph convolutional networks to capture both topological and attribute information. Additionally, convolutional neural networks are employed to further learn the features and obtain higher-order feature representations containing richer information about nodes. The Hadamard product is utilized to effectively combine circRNA features with disease features, and a multilayer perceptron is applied to predict the association between each pair of circRNA and disease. Five-fold cross validation experiments conducted on the CircR2Disease dataset demonstrate the accurate prediction capabilities of SGFCCDA in identifying potential circRNA-disease associations. Furthermore, case studies provide further confirmation of SGFCCDA's ability to identify disease-associated circRNAs.
Junliang Shang, Linqian Zhao, Xin He 0008, Xianghan Meng, Feng Li 0033, Jin-Xing Liu 0001
IEEE J. Biomed. Health Informatics7
2024 M3HOGAT: A Multi-View Multi-Modal Multi-Scale High-Order Graph Attention Network for Microbe-Disease Association Prediction
abstract
Numerous scientific studies have found a link between diverse microorganisms in the human body and complex human diseases. Because traditional experimental approaches are time-consuming and expensive, using computational methods to identify microbes correlated with diseases is critical. In this paper, a new microbe-disease association prediction model is proposed that combines a multi-view multi-modal network and a multi-scale feature fusion mechanism, called M3HOGAT. Firstly, a microbe-disease association network and multiple similarity views are constructed based on multi-source information. Then, consider that neighbor information from disparate orders might be more adept at learning node representations. Consequently, the higher-order graph attention network (HOGAT) is devised to aggregate neighbor information from disparate orders to extract microbe and disease features from different networks and views. Given that the embedding features of microbe and disease from different views possess varying importance, a multi-scale feature fusion mechanism is employed to learn their interaction information, thereby generating the final feature of microbes and diseases. Finally, an inner product decoder is used to reconstruct the microbe-disease association matrix. Compared with five state-of-the-art methods on the HMDAD and Disbiome datasets, the results of 5-fold cross-validations show that M3HOGAT achieves the best performance. Furthermore, case studies on asthma and obesity confirm the effectiveness of M3HOGAT in identifying potential disease-related microbes.
Jin-Xing Liu 0001, Feng Li 0033, Juan Wang 0003, Ying-Lian Gao
IEEE J. Biomed. Health Informatics3
2024 FSCME: A Feature Selection Method Combining Copula Correlation and Maximal Information Coefficient by Entropy Weights
abstract
Feature selection is a critical component of data mining and has garnered significant attention in recent years. However, feature selection methods based on information entropy often introduce complex mutual information forms to measure features, leading to increased redundancy and potential errors. To address this issue, we propose FSCME, a feature selection method combining Copula correlation (Ccor) and the maximum information coefficient (MIC) by entropy weights. The FSCME takes into consideration the relevance between features and labels, as well as the redundancy among candidate features and selected features. Therefore, the FSCME utilizes Ccor to measure the redundancy between features, while also estimating the relevance between features and labels. Meanwhile, the FSCME employs MIC to enhance the credibility of the correlation between features and labels. Moreover, this study employs the Entropy Weight Method (EWM) to evaluate and assign weights to the Ccor and MIC. The experimental results demonstrate that FSCME yields a more effective feature subset for subsequent clustering processes, significantly improving the classification performance compared to the other six feature selection methods.
Junliang Shang, Qianqian Ren, Feng Li 0033, Cui-Na Jiao, Jin-Xing Liu 0001
IEEE J. Biomed. Health Informatics4
2024 MGCNRF: Prediction of Disease-Related miRNAs Based on Multiple Graph Convolutional Networks and Random Forest
abstract
Increasing microRNAs (miRNAs) have been confirmed to be inextricably linked to various diseases, and the discovery of their associations has become a routine way of treating diseases. To overcome the time-consuming and laborious shortcoming of traditional experiments in verifying the associations of miRNAs and diseases (MDAs), a variety of computational methods have emerged. However, these methods still have many shortcomings in terms of predictive performance and accuracy. In this study, a model based on multiple graph convolutional networks and random forest (MGCNRF) was proposed for the prediction MDAs. Specifically, MGCNRF first mapped miRNA functional similarity and sequence similarity, disease semantic similarity and target similarity, and the known MDAs into four different two-layer heterogeneous networks. Second, MGCNRF applied four heterogeneous networks into four different layered attention graph convolutional networks (GCNs), respectively, to extract MDA embeddings. Finally, MGCNRF integrated the embeddings of every MDA into the features of the miRNA-disease pair and predicted potential MDAs through the random forest (RF). Fivefold cross-validation was applied to verify the prediction performance of MGCNRF, which outperforms the other seven state-of-the-art methods by area under curve. Furthermore, the accuracy and the case studies of different diseases further demonstrate the scientific rationale of MGCNRF. In conclusion, MGCNRF can serve as a scientific tool for predicting potential MDAs.
Feng Li 0033, Boxin Guan, Jin-Xing Liu 0001, Junliang Shang
IEEE Trans. Neural Networks Learn. Syst.3
2023 idenLD-AREL: identifying lncRNA-disease associations by random forests based on an ensemble learning framework
abstract
Identification of disease-associated long non-coding RNAs (lncRNAs) facilitates the understanding of the pathogenesis of complex diseases. Many different types of computational models have been proposed. Although some of them have achieved encouraging results in predicting disease-associated lncRNAs, how to obtain stable results is still a challenge. In this paper, we propose a computational model based on an ensemble learning framework via the adaptive random forests, in short, idenLD-AREL. The idenLD-AREL integrates multiple random forest predictors and adaptive strategies to predict the scores of potential lncRNA-disease associations (LDAs), which ensure the stability and accuracy of the prediction results. In addition, there are a large number of false negative samples in the association datasets. For this reason, the resampling strategy is applied to idenLD-AREL to balance the samples. The idenLD-AREL is assessed by five-fold cross-validation in both the benchmark dataset and independent test set, showing excellent performance. Besides, the experimental results of the case study further demonstrate the effectiveness of the idenLD-AREL in predicting potential LDAs. The demo codes of the iLncDA-RSN are available online at https://github.com/CDMBlab/idenLD-AREL.
Yahan Li, Junliang Shang, Feng Li 0033, Jin-Xing Liu 0001
BIBM5
2023 MKGSAGE: A Computational Framework via Multiple Kernel Fusion on GraphSAGE for Inferring Potential Disease-Related Microbes
abstract
Microbes play a crucial role within the human body and are closely associated with the occurrence and development of numerous diseases. Studies have shown that disruptions in the composition and functionality of microbes can lead to immune system imbalances, inflammatory responses, and subsequently impact human health. Therefore, developing computational models to discover the potential connections between microbes and diseases is currently a hot topic. In this paper, a computational framework based on the multiple kernel fusion of graph embedding with sampling and aggregation (GraphSAGE) and dual Laplace regularized least squares called MKGSAGE is proposed for predicting potential links between microbe and disease. First, multiple layers embedding features of microbe and disease are learned from the initial input features by GraphSAGE. The kernel matrices are then calculated separately for each layer based on the Gaussian interaction profile (GIP). Furthermore, the multiple kernel fusion method is proposed for fusing kernel matrices of each layer and the initial similarity matrix. Dual Laplacian regularized least squares are finally applied for potential microbe-disease association prediction. Compared with six state-of-the-art methods on the HMDAD dataset, 5-fold cross-validations show that MKGSAGE performs best. In addition, case studies on asthma and inflammatory bowel disease further validate the effectiveness of MKGSAGE on discovering novel microbe-disease associations.
Jin-Xing Liu 0001, Bao-Min Liu, Ling-Yun Dai, Feng Li 0033, Ying-Lian Gao
BIBM5
2023 Spectral clustering based on multi-similarity learning method for single-cell RNA-seq data
abstract
The inherent complexities of single-cell RNA-seq data (scRNA-seq), such as high dimensionality, low signal-to-noise ratio, cellular heterogeneity, and imbalanced distribution of subcellular types, pose significant challenges when conducting cell type analysis. To address these obstacles, employing appropriate data preprocessing techniques in the single-cell clustering process is crucial, and spectral clustering is particularly effective due to its robustness to noise and outliers. Therefore, this paper presents a novel spectral clustering algorithm based on multi-similarity learning method (MSSC). However, utilizing a pairwise strategy to assess the similarity between two data points in conventional spectral clustering results in an insufficient representation of the intricate relationships within the dataset. In light of this issue, the proposed algorithm employs two similarity measurement methods, namely Euclidean distance and Spearman rank correlation coefficient, to obtain similarity matrices. These matrices are then fused for use in spectral clustering. Additionally, prior to performing spectral clustering, the scRNA-seq data is preprocessed using the Sigmoid kernel similarity method and normalization techniques. As a consequence, our method yields a more extensive and intricate dataset similarity information, thereby enhancing the performance of spectral clustering. Finally, the Gaussian mixture model (GMM) is used for clustering. In most cases, experiments validated that the MSSC method outperforms the other four clustering methods on seven benchmark scRNA-seq datasets.
Lianlian Zhang, Shaoyi Liu, Qianqian Ren, Junliang Shang, Feng Li 0033
BIBM6
2023 Spectral Clustering of Single-Cell RNA-Sequencing Data by Multiple Feature Sets Affinity
Feng Li 0033, Junliang Shang, Qianqian Ren, Shengjun Li
ICIC (3)2
2023 scGASI: A Graph Autoencoder-Based Single-Cell Integration Clustering Method
Tian-Jing Qiao, Feng Li 0033, Shasha Yuan, Ling-Yun Dai, Juan Wang 0003
ISBRA2
2023 ABCAE: Artificial Bee Colony Algorithm with Adaptive Exploitation for Epistatic Interaction Detection
Qianqian Ren, Yahan Li, Feng Li 0033, Jin-Xing Liu 0001, Junliang Shang
ISBRA3
2023 DM-MOGA: a multi-objective optimization genetic algorithm for identifying disease modules of non-small cell lung cancer
abstract
BACKGROUND: Constructing molecular interaction networks from microarray data and then identifying disease module biomarkers can provide insight into the underlying pathogenic mechanisms of non-small cell lung cancer. A promising approach for identifying disease modules in the network is community detection. RESULTS: In order to identify disease modules from gene co-expression networks, a community detection method is proposed based on multi-objective optimization genetic algorithm with decomposition. The method is named DM-MOGA and possesses two highlights. First, the boundary correction strategy is designed for the modules obtained in the process of local module detection and pre-simplification. Second, during the evolution, we introduce Davies-Bouldin index and clustering coefficient as fitness functions which are improved and migrated to weighted networks. In order to identify modules that are more relevant to diseases, the above strategies are designed to consider the network topology of genes and the strength of connections with other genes at the same time. Experimental results of different gene expression datasets of non-small cell lung cancer demonstrate that the core modules obtained by DM-MOGA are more effective than those obtained by several other advanced module identification methods. CONCLUSIONS: The proposed method identifies disease-relevant modules by optimizing two novel fitness functions to simultaneously consider the local topology of each gene and its connection strength with other genes. The association of the identified core modules with lung cancer has been confirmed by pathway and gene ontology enrichment analysis.
Junliang Shang, Xuhui Zhu, Feng Li 0033, Jin-Xing Liu 0001
BMC Bioinform.4
2023 GCCN: Graph Capsule Convolutional Network for Progressive Mild Cognitive Impairment Prediction and Pathogenesis Identification Based on Imaging Genetic Data
abstract
In this study, we proposed a novel method called the graph capsule convolutional network (GCCN) to predict the progression from mild cognitive impairment to dementia and identify its pathogenesis. First, we proposed a novel risk gene discovery component to indirectly target genes with higher interactions with others. These risk genes and brain regions were collected as nodes to construct heterogeneous pathogenic information association graphs. Second, the graph capsules were established by projecting heterogeneous pathogenic information into a set of disentangled latent components. The orientation and length of capsules are representations of the format and intensity of pathogenic information. Third, graph capsule convolution network was used to model the information flows among pathogenic factors and elaborates the convergence of primary capsules to advanced capsules. The advanced capsule is a concept that organizes pathogenic information based on its consistency, and the synergistic effects of advanced capsules directed the development of the disease. Finally, discriminative pathogenic information flows were captured by a straightforward built-in interpretation mechanism, i.e., the dynamic routing mechanism, and applied to the identification of pathogenesis. GCCN has been experimentally shown to be significantly advanced on public datasets. Further experiments have shown that the pathogenic factors identified by GCCN are evidential and closely related to progressive mild cognitive impairment.
Junliang Shang, Qi Zou 0003, Qianqian Ren, Boxin Guan, Feng Li 0033, Jin-Xing Liu 0001
IEEE J. Biomed. Health Informatics5
2023 MSGCA: Drug-Disease Associations Prediction Based on Multi-Similarities Graph Convolutional Autoencoder
abstract
Identifying drug-disease associations (DDAs) is critical to the development of drugs. Traditional methods to determine DDAs are expensive and inefficient. Therefore, it is imperative to develop more accurate and effective methods for DDAs prediction. Most current DDAs prediction methods utilize original DDAs matrix directly. However, the original DDAs matrix is sparse, which greatly affects the prediction consequences. Hence, a prediction method based on multi-similarities graph convolutional autoencoder (MSGCA) is proposed for DDAs prediction. First, MSGCA integrates multiple drug similarities and disease similarities using centered kernel alignment-based multiple kernel learning (CKA-MKL) algorithm to form new drug similarity and disease similarity, respectively. Second, the new drug and disease similarities are improved by linear neighborhood, and the DDAs matrix is reconstructed by weighted K nearest neighbor profiles. Next, the reconstructed DDAs and the improved drug and disease similarities are integrated into a heterogeneous network. Finally, the graph convolutional autoencoder with attention mechanism is utilized to predict DDAs. Compared with extant methods, MSGCA shows superior results on three datasets. Furthermore, case studies further demonstrate the reliability of MSGCA.
Ying Wang 0143, Ying-Lian Gao, Juan Wang 0003, Feng Li 0033, Jin-Xing Liu 0001
IEEE J. Biomed. Health Informatics4
2023 NLRRC: A Novel Clustering Method of Jointing Non-Negative LRR and Random Walk Graph Regularized NMF for Single-Cell Type Identification
abstract
The development of single-cell RNA sequencing (scRNA-seq) technology has opened up a new perspective for us to study disease mechanisms at the single cell level. Cell clustering reveals the natural grouping of cells, which is a vital step in scRNA-seq data analysis. However, the high noise and dropout of single-cell data pose numerous challenges to cell clustering. In this study, we propose a novel matrix factorization method named NLRRC for single-cell type identification. NLRRC joins non-negative low-rank representation (LRR) and random walk graph regularized NMF (RWNMFC) to accurately reveal the natural grouping of cells. Specifically, we find the lowest rank representation of single-cell samples by non-negative LRR to reduce the difficulty of analyzing high-dimensional samples and capture the global information of the samples. Meanwhile, by using random walk graph regularization (RWGR) and NMF, RWNMFC captures manifold structure and cluster information before generating a cluster allocation matrix. The cluster assignment matrix contains cluster labels, which can be used directly to get the clustering results. The performance of NLRRC is validated on simulated and real single-cell datasets. The results of the experiments illustrate that NLRRC has a significant advantage in single-cell type identification.
Juan Wang 0003, Linping Wang, Shasha Yuan, Feng Li 0033, Jin-Xing Liu 0001, Junliang Shang
IEEE J. Biomed. Health Informatics4
2022 scSASSL: Self-attention semi-supervised learning with deep generative models to automatically identify cell types
abstract
High-throughput single-cell sequencing has distinct advantages over previous bulk sequencing technologies. It provides an opportunity for researchers to study cell heterogeneity from the level of individual cells and explain biological relationships between individual cells from a higher resolution perspective. However, single-cell data are characterized by a large number of samples, high dimensionality, and sparseness, which pose a challenge to traditional methods. Therefore, we develop a semi-supervised deep generative model with a self-attention mechanism. The use of deep learning methods allows the denoising and dimensionality reduction of high-dimensional single-cell data nonlinear. We use a neural network with a self-attention mechanism for cell type prediction. This approach facilitates the neural network to extract cell-to-cell relationship features and enhances the model’s ability to extract features. The model can generate data. We apply this ability of imputation to single-cell datasets, thus solving the sparsity problem of single-cell datasets. We have conducted experiments on several simulated and real datasets, and the experimental results show that our proposed method largely outperforms other existing methods both in terms of identifying cell types and imputation on single-cell data. Our method is scalable because it can handle large-scale single-cell datasets of more than a million quantities. This method is promising in other fields as well. All source codes used in our experiments have been deposited at https://github.com/FengLi12/scSASSL.
Hongyu Duan, Feng Li 0033, Xin Chu, Zhensheng Sun, Junliang Shang, Xikui Liu 0001, Yan Li 0041
BIBM2
2022 Probability Connectivity-Based Multimodality Regression Analysis for Associating Disease-Specific Multimodal Brain Imaging Phenotypes with Genetic Risk Factors
abstract
Neuroimaging genetics is a powerful technique for discovering the relationships between genotype and imaging phenotype. However, many univariate or multivariate regression approaches have only focused on imaging quantitative traits (QTs) that are relevant to some genetic markers on distinct pathways and might not be disease specific. In addition, there are complex relations between subjects of distinct modalities and diagnosis labels, which contain useful information for the treatment of Alzheimer’s disease (AD). Here, a novel probability connectivity-based penalty is developed for incorporating the prior information to explore relations among different subjects with disease status. Specifically, the Pearson’s correlation coefficient (PCC) is used to construct a similarity matrix in a probability graph, first to express the connectivity weights between subjects, which can reflect the different correlations among subjects within the same class. Second, a diagnosis-aligned probability connectivity-based multimodal regression (DPCMR) method is employed to find the relations among modalities of distinct subjects. It also mines associations between genetic markers and imaging phenotypes. The AD risk single nucleotide polymorphism (SNP) APOE rs429358 and three modalities of neuroimaging data are used to verify the performance of all of the methods. The experimental results reveal that DPCMR has better performance and identifies some brain regions across multiple modalities related to diseases.
Cui-Na Jiao, Chun-Hou Zheng 0001, Jin-Xing Liu 0001, Feng Li 0033
BIBM4
2022 Artificial bee colony algorithm based on self-adjusting random grouping for high-order epistasis detection
abstract
In the genome-wide association studies (GWAS), epistasis detection is of great significance to study the pathogenesis of complex diseases. Epistasis refers to the effect of interactions between multiple single nucleotide polymorphisms (SNPs) on complex diseases. In this paper, an artificial bee colony algorithm based on self-adjusting random grouping (ABC-SRG) is proposed for high-order epistasis detection. ABC-SRG adopts a new self-adjusting random grouping strategy, which realizes the division of the original data according to the fitness value of each grouping. In addition, a variance-based adaptive iteration strategy is proposed, which implements the adaptive iteration through the variance of the fitness value of each iteration of the algorithm. To demonstrate the effectiveness of the algorithm, the experiments on simulated data and real data were conducted. In the simulation experiments, ABC-SRG was compared with the other five methods for second-order and third-order SNP interaction detection. Age-related macular degeneration (AMD) data were selected for the real data experiment, and most of the SNP interactions detected in the experiment have been confirmed to be related to the AMD disease. Therefore, ABC-SRG is an effective method to detect high-order epistasis.
Junliang Shang, Yijun Gu, Feng Li 0033, Jin-Xing Liu 0001, Boxin Guan
BIBM4
2022 Identification of cancer driver modules by combining network functional and topology information
abstract
Accurate identification of cancer driver modules or pathways is important for controlling disease progression and timely treatment. In recent years, most approaches have been based on mutation data combined with gene interaction networks to identify cancer driver modules, but cancer-related genes tend to interact with each other, and the mutations they experience disruption their neighbors. Therefore, we propose a framework that combines network function and topological information to quantify the extent to which mutated genes disrupt their neighbors. Firstly, similarity in protein-protein interaction networks binds to high coverage and high mutual exclusivity of mutant genes, which are used to obtain the impact of the interaction between two mutant genes on biological function. Secondly, we quantified the degree of gene disruption by mutant genes in their neighborhood using an adaptive spread strength measure to obtain the gene spread strength network (GSSN). Finally, the module is extended using CFinder strategy to obtain the optimal driving module. We apply our method to 12 cancer datasets, and the experimental results show that our method outperforms the other three methods on most datasets. At the same time, we also analyze common and low-frequency driver modules in cancer.
Xin Chu, Feng Li 0033, Hongyu Duan, Junliang Shang, Juan Wang 0003, Jin-Xing Liu 0001
BIBM2
2022 Construction of Gene Network Based on Inter-tumor Heterogeneity for Tumor Type Identification
Zhensheng Sun, Junliang Shang, Hongyu Duan, Jin-Xing Liu 0001, Xikui Liu 0001, Yan Li 0041, Feng Li 0033
ICIC (2)7
2022 A Network-Based Voting Method for Identification and Prioritization of Personalized Cancer Driver Genes
Feng Li 0033, Junliang Shang, Xikui Liu 0001, Yan Li 0041
ISBRA2
2022 TDCOSR: A Multimodality Fusion Framework for Association Analysis Between Genes and ROIs of Alzheimer's Disease
Qi Zou 0003, Feng Li 0033, Juan Wang 0003, Jin-Xing Liu 0001, Junliang Shang
ISBRA3
2021 Multiscale part mutual information for quantifying nonlinear direct associations in networks
abstract
MOTIVATION: For network-assisted analysis, which has become a popular method of data mining, network construction is a crucial task. Network construction relies on the accurate quantification of direct associations among variables. The existence of multiscale associations among variables presents several quantification challenges, especially when quantifying nonlinear direct interactions. RESULTS: In this study, the multiscale part mutual information (MPMI), based on part mutual information (PMI) and nonlinear partial association (NPA), was developed for effectively quantifying nonlinear direct associations among variables in networks with multiscale associations. First, we defined the MPMI in theory and derived its five important properties. Second, an experiment in a three-node network was carried out to numerically estimate its quantification ability under two cases of strong associations. Third, experiments of the MPMI and comparisons with the PMI, NPA and conditional mutual information were performed on simulated datasets and on datasets from DREAM challenge project. Finally, the MPMI was applied to real datasets of glioblastoma and lung adenocarcinoma to validate its effectiveness. Results showed that the MPMI is an effective alternative measure for quantifying nonlinear direct associations in networks, especially those with multiscale associations. AVAILABILITY AND IMPLEMENTATION: The source code of MPMI is available online at https://github.com/CDMB-lab/MPMI. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Junliang Shang, Feng Li 0033, Jin-Xing Liu 0001, Honghai Zhang
Bioinform.4
2021 DSCMF: prediction of LncRNA-disease associations based on dual sparse collaborative matrix factorization
abstract
BACKGROUND: In the development of science and technology, there are increasing evidences that there are some associations between lncRNAs and human diseases. Therefore, finding these associations between them will have a huge impact on our treatment and prevention of some diseases. However, the process of finding the associations between them is very difficult and requires a lot of time and effort. Therefore, it is particularly important to find some good methods for predicting lncRNA-disease associations (LDAs). RESULTS: -norm is added in our method. At the same time, Gaussian interaction profile kernel is added to our method, which increase the network similarity between lncRNA and disease. Finally, the AUC value obtained by the experiment is used to evaluate the quality of our method, and the AUC value is obtained by the ten-fold cross-validation method. CONCLUSIONS: The AUC value obtained by the DSCMF method is 0.8523. At the end of the paper, simulation experiment is carried out, and the experimental results of prostate cancer, breast cancer, ovarian cancer and colorectal cancer are analyzed in detail. The DSCMF method is expected to bring some help to lncRNA-disease associations research. The code can access the https://github.com/Ming-0113/DSCMF website.
Jin-Xing Liu 0001, Ming-Ming Gao, Ying-Lian Gao, Feng Li 0033
BMC Bioinform.5
2020 IDSSIM: an lncRNA functional similarity calculation model based on an improved disease semantic similarity method
abstract
BACKGROUND: It has been widely accepted that long non-coding RNAs (lncRNAs) play important roles in the development and progression of human diseases. Many association prediction models have been proposed for predicting lncRNA functions and identifying potential lncRNA-disease associations. Nevertheless, among them, little effort has been attempted to measure lncRNA functional similarity, which is an essential part of association prediction models. RESULTS: In this study, we presented an lncRNA functional similarity calculation model, IDSSIM for short, based on an improved disease semantic similarity method, highlight of which is the introduction of information content contribution factor into the semantic value calculation to take into account both the hierarchical structures of disease directed acyclic graphs and the disease specificities. IDSSIM and three state-of-the-art models, i.e., LNCSIM1, LNCSIM2, and ILNCSIM, were evaluated by applying their disease semantic similarity matrices and the lncRNA functional similarity matrices, as well as corresponding matrices of human lncRNA-disease associations coming from either lncRNADisease database or MNDR database, into an association prediction method WKNKN for lncRNA-disease association prediction. In addition, case studies of breast cancer and adenocarcinoma were also performed to validate the effectiveness of IDSSIM. CONCLUSIONS: Results demonstrated that in terms of ROC curves and AUC values, IDSSIM is superior to compared models, and can improve accuracy of disease semantic similarity effectively, leading to increase the association prediction ability of the IDSSIM-WKNKN model; in terms of case studies, most of potential disease-associated lncRNAs predicted by IDSSIM can be confirmed by databases and literatures, implying that IDSSIM can serve as a promising tool for predicting lncRNA functions, identifying potential lncRNA-disease associations, and pre-screening candidate lncRNAs to perform biological experiments. The IDSSIM code, all experimental data and prediction results are available online at https://github.com/CDMB-lab/IDSSIM .
Wenwen Fan, Junliang Shang, Feng Li 0033, Shasha Yuan, Jin-Xing Liu 0001
BMC Bioinform.3
2020 Detection of Driver Modules with Rarely Mutated Genes in Cancers
abstract
Identifying driver modules or pathways is a key challenge to interpret the molecular mechanisms and pathogenesis underlying cancer. An increasing number of studies suggest that rarely mutated genes are important for the development of cancer. However, the driver modules consisting of mutated genes with low-frequency driver mutations are not well characterized. To identify driver modules with rarely mutated genes, we propose a functional similarity index to quantify the functional relationship between rarely mutated genes and other ones in the same module. Then, we develop a method to detect Driver Modules with Rarely mutated Genes (DMRG) by incorporating the functional similarity, coverage and mutual exclusivity. By applying DMRG on TCGA cancer dataset on three networks: HINT+HI2012, iRefIndex and MultiNet, we detect driver modules intersecting with the well-known signalling pathways and protein complexes, such as the cell cycle pathway and the mediator complex. DMRG can also detect driver modules effectively with 20, 40, 60 and 80 percent of samples by random selection. When compared with HotNet2, DMRG detects more rarely mutated cancer genes and has higher pathway enrichment. Overall, DMRG provides an effective method for the identification of driver modules with rarely mutated genes.
Feng Li 0033, Lin Gao 0006, Bingbo Wang
IEEE ACM Trans. Comput. Biol. Bioinform.1
2019 Dual Sparse Collaborative Matrix Factorization Method Based on Gaussian Kernel Function for Predicting LncRNA-Disease Associations
Ming-Ming Gao, Ying-Lian Gao, Feng Li 0033, Jin-Xing Liu 0001
ICIC (3)4
2018 Feature related multi-view nonnegative matrix factorization for identifying conserved functional modules in multiple biological networks
abstract
BACKGROUND: Comprehensive analyzing multi-omics biological data in different conditions is important for understanding biological mechanism in system level. Multiple or multi-layer network model gives us a new insight into simultaneously analyzing these data, for instance, to identify conserved functional modules in multiple biological networks. However, because of the larger scale and more complicated structure of multiple networks than single network, how to accurate and efficient detect conserved functional biological modules remains a significant challenge. RESULTS: Here, we propose an efficient method, named ConMod, to discover conserved functional modules in multiple biological networks. We introduce two features to characterize multiple networks, thus all networks are compressed into two feature matrices. The module detection is only performed in the feature matrices by using multi-view non-negative matrix factorization (NMF), which is independent of the number of input networks. Experimental results on both synthetic and real biological networks demonstrate that our method is promising in identifying conserved modules in multiple networks since it improves the accuracy and efficiency comparing with state-of-the-art methods. Furthermore, applying ConMod to co-expression networks of different cancers, we find cancer shared gene modules, the majority of which have significantly functional implications, such as ribosome biogenesis and immune response. In addition, analyzing on brain tissue-specific protein interaction networks, we detect conserved modules related to nervous system development, mRNA processing, etc. CONCLUSIONS: ConMod facilitates finding conserved modules in any number of networks with a low time and space complexity, thereby serve as a valuable tool for inference shared traits and biological functions of multiple biological system.
Peizhuo Wang, Lin Gao 0006, Yuxuan Hu 0004, Feng Li 0033
BMC Bioinform.4
2015 Identifying overlapping mutated driver pathways by constructing gene networks in cancer
abstract
BACKGROUND: Large-scale cancer genomic projects are providing lots of data on genomic, epigenomic and gene expression aberrations in many cancer types. One key challenge is to detect functional driver pathways and to filter out nonfunctional passenger genes in cancer genomics. Vandin et al. introduced the Maximum Weight Sub-matrix Problem to find driver pathways and showed that it is an NP-hard problem. METHODS: To find a better solution and solve the problem more efficiently, we present a network-based method (NBM) to detect overlapping driver pathways automatically. This algorithm can directly find driver pathways or gene sets de novo from somatic mutation data utilizing two combinatorial properties, high coverage and high exclusivity, without any prior information. We firstly construct gene networks based on the approximate exclusivity between each pair of genes using somatic mutation data from many cancer patients. Secondly, we present a new greedy strategy to add or remove genes for obtaining overlapping gene sets with driver mutations according to the properties of high exclusivity and high coverage. RESULTS: To assess the efficiency of the proposed NBM, we apply the method on simulated data and compare results obtained from the NBM, RME, Dendrix and Multi-Dendrix. NBM obtains optimal results in less than nine seconds on a conventional computer and the time complexity is much less than the three other methods. To further verify the performance of NBM, we apply the method to analyze somatic mutation data from five real biological data sets such as the mutation profiles of 90 glioblastoma tumor samples and 163 lung carcinoma samples. NBM detects groups of genes which overlap with known pathways, including P53, RB and RTK/RAS/PI(3)K signaling pathways. New gene sets with p-value less than 1e-3 are found from the somatic mutation data. CONCLUSIONS: NBM can detect more biologically relevant gene sets. Results show that NBM outperforms other algorithms for detecting driver pathways or gene sets. Further research will be conducted with the use of novel machine learning techniques.
Hao Wu 0062, Lin Gao 0006, Feng Li 0033, Xiaofei Yang 0003, Nikola K. Kasabov
BMC Bioinform.3