EDBT 2026 Demo / reviewers in the wild / expert
Ying-Lian Gao
dblp:166/6139
· DBLP profile ↗
85ranked-venue papers
6as first author
53since 2021 · last 2026
0000-0003-0483-5622ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 70 · 6 first-author · 40 since 2021Artificial intelligence and machine learning · 14 · 12 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SSBiA: A Framework for Predicting Microbe-Disease Associations Based on Signed Subgraphs and Bi-Feature Aggregation
Ying-Lian Gao, Ming-Li Cui, Junliang Shang, Chun-Hou Zheng 0001 |
ICIC (27) | 1 |
| 2026 | Predicting microbe-disease associations based on multi-modal using reliable negative sample and cross attention network
Cui-Na Jiao, Xinchun Cui, Ying-Lian Gao, Jin-Xing Liu 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Deep association analysis framework with multi-modal attention fusion for brain imaging genetics
Shuang-Qing Wang, Cui-Na Jiao, Ying-Lian Gao, Xinchun Cui, Yan-Li Wang |
Medical Image Anal. | 3 |
| 2026 | Bi-directional attention network with drop aggregation for microRNA-disease association prediction
Yan-Fang Yang, Ming-Li Cui, Ying-Lian Gao, Yan-Li Wang |
Neural Networks | 4 |
| 2026 | Multi-type Transformer encoding-based graph contrastive learning for drug repositioning
Ming-Li Cui, Ying-Lian Gao, Chun-Hou Zheng 0001, Yan-Li Wang |
Pattern Recognit. | 2 |
| 2026 | Multi-Grained Line Graph Neural Network With Hierarchical Contrastive Learning for Predicting Drug-Disease AssociationsabstractPredicting drug-disease associations is a crucial step in drug repositioning, especially with computational methods that quickly locate potential drug-disease pairs. Heterogenous network is a common tool for introducing multiple type relation information about drugs and diseases. However, the diversity of relations is ignored in most of existing methods, which makes them difficult to explore type semantic information with structure properties. Therefore, we propose a relation-centric GNN framework to encode critical association patterns. Firstly, we utilize a relation-centric graph, line graph, to represent the context of a drug-disease pair identified as the center node. The prediction problem is modeled to learn the embedding vector of the center node. Secondly, a multi-grained line graph neural network (MGLGNN) is designed to excavate fine-grained features that encapsulate local graph structures. We theoretically define a handful of typical nodes that can be regarded as high-order abstractions of relations in each type. Then, MGLGNN distills the local information and passes it to typical nodes from a global perspective. With learned multi-grained features, the center node automatically captures heterogenous relation semantics and structure patterns. Thirdly, a hierarchical contrastive learning (HCL) mechanism is proposed to ensure the quality of multi-grained features in an unsupervised way. Extensive experiments show the great potential of our model in mining drug-disease associations. Bao-Min Liu, Ling-Yun Dai, Junliang Shang, Chun-Hou Zheng 0001, Ying-Lian Gao, Rui Gao 0006, Jin-Xing Liu 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | STDDAE: Identifying spatial domains in spatial transcriptomics by dual denoising autoencoder with attention mechanismabstractSpatial transcriptomics provides a novel perspective for comprehending the intricate relationship between tissue structure and function, as well as for discovering new cell types and subtypes. However, it remains a significant challenge to accurately identify spatial domains with similar gene expression, which requires efficient combination of gene expression data , histology image information, and spatial location. To address this challenge, a novel dual denoising autoencoder with attention mechanism (STDDAE) is proposed. STDDAE integrates gene expression data, histology image information and spatial location, and the decoder consists of a master decoder and a follower decoder, which are jointly optimized to generate low-dimensional latent embeddings for precise spatial domain identification. The performance of STDDAE was evaluated across four datasets with varying resolutions and platforms. The experimental findings validated that STDDAE outperformed other cutting-edg methods in spatial domain identification, trajectory inference, and data denoising. Additionally, STDDAE successfully detected differentially expressed genes within identified spatial domains, which may be valuable in disease diagnosis, prognostic assessment, and treatment selection. Ying-Lian Gao, Cui-Na Jiao, Xu-Ran Dou, Feng Li 0033, Jin-Xing Liu 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | SAMGCN: A spatially-augmented multi-view graph convolutional network for identifying spatial domains
Hao Liu 0075, Ying-Lian Gao, Cui-Na Jiao, Junliang Shang |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | SpaMGAN: Multi-view graph augmentation network for spatial domain identification in spatial transcriptomics
Hao Liu 0075, Cui-Na Jiao, Chun-Hou Zheng 0001, Ying-Lian Gao, Jin-Xing Liu 0001, Yan-Li Wang |
Knowl. Based Syst. | 5 |
| 2025 | TEMCL: Prediction of Drug-Disease Associations Based on Transformer and Enhanced Multi-View Contrastive LearningabstractDrug repositioning (DR) has emerged as an effective method of identifying new indications for existing drugs. Many DR methods have demonstrated superior performance. However, most of them utilize a limited number of biological entities, ignoring the critical role of other entities in addressing data sparsity as well as improving model generalization capabilities. In addition, fully capturing high-order information of biological data still needs to be fully explored. To address above issues, a model based on transformer and enhanced multi-view contrastive learning (TEMCL) is proposed for predicting drug-disease associations (DDAs). Firstly, transformer is employed to obtain high-order features of nodes from similarity information. Secondly, based on similarity matrices and association matrices of nodes, two different types of views are constructed, i.e., homogeneous hypergraphs and heterogeneous association graphs. Among them, to alleviate sparsity problem existing in heterogeneous graphs, protein nodes as well as meta-path enhancement strategy are introduced. Thirdly, hypergraph convolutional network and heterogeneous graph transformer are used to extract node features on above two types of views, respectively. Contrastive learning is applied to obtain more representative features. Finally, multilayer perceptron (MLP) is used for predicting DDAs. Experiments show that TEMCL outperforms existing methods on DR task, exhibiting superior performance. In addition, case studies further demonstrate the effectiveness of this model. TEMCL provides new insights for identifying novel DDAs. Ming-Li Cui, Cui-Na Jiao, Ying-Lian Gao, Junliang Shang, Chun-Hou Zheng 0001, Jin-Xing Liu 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Integrating Local and Global Information to Decipher Spatial Domains of Spatial Transcriptomics by Attention-based Graph Convolutional NetworkabstractRecent developments in spatial transcriptomics technologies have made it possible to obtain gene expression profiles while maintaining spatial context. Precisely identifying spatial domains is essential for downstream analysis, requiring the effective integration of gene expression profiles with spatial information. To overcome the challenge of low accuracy in spatial domain identification, this paper proposed a deep learning model called LGAGCN based on local and global information. It used graph convolutional network to learn the features of local and global views and employed an attention mechanism to integrate embeddings from different views. Moreover, experiments were conducted on the human dorsolateral prefrontal cortex (DLPFC) dataset and the human breast cancer (HBC) dataset to evaluate the effectiveness of the model. The experimental results showed that LGAGCN outperformed state-of-the-art methods in spatial clustering task. Xu-Ran Dou, Junliang Shang, Chun-Hou Zheng 0001, Ying-Lian Gao, Jin-Xing Liu 0001 |
BIBM | 5 |
| 2024 | Multi-modal imaging genetics data fusion by deep auto-encoder and self-representation network for Alzheimer's disease diagnosis and biomarkers extraction
Cui-Na Jiao, Ying-Lian Gao, Junliang Shang, Jin-Xing Liu 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | A review of recent advances in spatially resolved transcriptomics data analysis
Ying-Lian Gao, Jing Jing 0001, Feng Li 0033, Chun-Hou Zheng 0001, Jin-Xing Liu 0001 |
Neurocomputing | 2 |
| 2024 | SLGCN: Structure-enhanced line graph convolutional network for predicting drug-disease associations
Bao-Min Liu, Ying-Lian Gao, Feng Li 0033, Chun-Hou Zheng 0001, Jin-Xing Liu 0001 |
Knowl. Based Syst. | 2 |
| 2024 | Diagnosis-Guided Deep Subspace Clustering Association Study for Pathogenetic Markers Identification of Alzheimer's Disease Based on Comparative AtlasesabstractThe roles of brain region activities and genotypic functions in the pathogenesis of Alzheimer's disease (AD) remain unclear. Meanwhile, current imaging genetics methods are difficult to identify potential pathogenetic markers by correlation analysis between brain network and genetic variation. To discover disease-related brain connectome from the specific brain structure and the fine-grained level, based on the Automated Anatomical Labeling (AAL) and human Brainnetome atlases, the functional brain network is first constructed for each subject. Specifically, the upper triangle elements of the functional connectivity matrix are extracted as connectivity features. The clustering coefficient and the average weighted node degree are developed to assess the significance of every brain area. Since the constructed brain network and genetic data are characterized by non-linearity, high-dimensionality, and few subjects, the deep subspace clustering algorithm is proposed to reconstruct the original data. Our multilayer neural network helps capture the non-linear manifolds, and subspace clustering learns pairwise affinities between samples. Moreover, most approaches in neuroimaging genetics are unsupervised learning, neglecting the diagnostic information related to diseases. We presented a label constraint with diagnostic status to instruct the imaging genetics correlation analysis. To this end, a diagnosis-guided deep subspace clustering association (DDSCA) method is developed to discover brain connectome and risk genetic factors by integrating genotypes with functional network phenotypes. Extensive experiments prove that DDSCA achieves superior performance to most association methods and effectively selects disease-relevant genetic markers and brain connectome at the coarse-grained and fine-grained levels. Cui-Na Jiao, Junliang Shang, Feng Li 0033, Xinchun Cui, Yan-Li Wang, Ying-Lian Gao, Jin-Xing Liu 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | Multi-Kernel Graph Attention Deep Autoencoder for MiRNA-Disease Association PredictionabstractAccumulating evidence indicates that microRNAs (miRNAs) can control and coordinate various biological processes. Consequently, abnormal expressions of miRNAs have been linked to various complex diseases. Recognizable proof of miRNA-disease associations (MDAs) will contribute to the diagnosis and treatment of human diseases. Nevertheless, traditional experimental verification of MDAs is laborious and limited to small-scale. Therefore, it is necessary to develop reliable and effective computational methods to predict novel MDAs. In this work, a multi-kernel graph attention deep autoencoder (MGADAE) method is proposed to predict potential MDAs. In detail, MGADAE first employs the multiple kernel learning (MKL) algorithm to construct an integrated miRNA similarity and disease similarity, providing more biological information for further feature learning. Second, MGADAE combines the known MDAs, disease similarity, and miRNA similarity into a heterogeneous network, then learns the representations of miRNAs and diseases through graph convolution operation. After that, an attention mechanism is introduced into MGADAE to integrate the representations from multiple graph convolutional network (GCN) layers. Lastly, the integrated representations of miRNAs and diseases are input into the bilinear decoder to obtain the final predicted association scores. Corresponding experiments prove that the proposed method outperforms existing advanced approaches in MDA prediction. Furthermore, case studies related to two human cancers provide further confirmation of the reliability of MGADAE in practice. Cui-Na Jiao, Feng Zhou 0021, Bao-Min Liu, Chun-Hou Zheng 0001, Jin-Xing Liu 0001, Ying-Lian Gao |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | KFDAE: CircRNA-Disease Associations Prediction Based on Kernel Fusion and Deep Auto-EncoderabstractCircRNA has been proved to play an important role in the diseases diagnosis and treatment. Considering that the wet-lab is time-consuming and expensive, computational methods are viable alternative in these years. However, the number of circRNA-disease associations (CDAs) that can be verified is relatively few, and some methods do not take full advantage of dependencies between attributes. To solve these problems, this paper proposes a novel method based on Kernel Fusion and Deep Auto-encoder (KFDAE) to predict the potential associations between circRNAs and diseases. Firstly, KFDAE uses a non-linear method to fuse the circRNA similarity kernels and disease similarity kernels. Then the vectors are connected to make the positive and negative sample sets, and these data are send to deep auto-encoder to reduce dimension and extract features. Finally, three-layer deep feedforward neural network is used to learn features and gain the prediction score. The experimental results show that compared with existing methods, KFDAE achieves the best performance. In addition, the results of case studies prove the effectiveness and practical significance of KFDAE, which means KFDAE is able to capture more comprehensive information and generate credible candidate for subsequent wet-lab. Wen-Yue Kang, Ying-Lian Gao, Ying Wang 0143, Feng Li 0033, Jin-Xing Liu 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | M3HOGAT: A Multi-View Multi-Modal Multi-Scale High-Order Graph Attention Network for Microbe-Disease Association PredictionabstractNumerous scientific studies have found a link between diverse microorganisms in the human body and complex human diseases. Because traditional experimental approaches are time-consuming and expensive, using computational methods to identify microbes correlated with diseases is critical. In this paper, a new microbe-disease association prediction model is proposed that combines a multi-view multi-modal network and a multi-scale feature fusion mechanism, called M3HOGAT. Firstly, a microbe-disease association network and multiple similarity views are constructed based on multi-source information. Then, consider that neighbor information from disparate orders might be more adept at learning node representations. Consequently, the higher-order graph attention network (HOGAT) is devised to aggregate neighbor information from disparate orders to extract microbe and disease features from different networks and views. Given that the embedding features of microbe and disease from different views possess varying importance, a multi-scale feature fusion mechanism is employed to learn their interaction information, thereby generating the final feature of microbes and diseases. Finally, an inner product decoder is used to reconstruct the microbe-disease association matrix. Compared with five state-of-the-art methods on the HMDAD and Disbiome datasets, the results of 5-fold cross-validations show that M3HOGAT achieves the best performance. Furthermore, case studies on asthma and obesity confirm the effectiveness of M3HOGAT in identifying potential disease-related microbes. Jin-Xing Liu 0001, Feng Li 0033, Juan Wang 0003, Ying-Lian Gao |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | A New Graph Autoencoder-Based Consensus-Guided Model for scRNA-seq Cell Type DetectionabstractSingle-cell RNA sequencing (scRNA-seq) technology is famous for providing a microscopic view to help capture cellular heterogeneity. This characteristic has advanced the field of genomics by enabling the delicate differentiation of cell types. However, the properties of single-cell datasets, such as high dropout events, noise, and high dimensionality, are still a research challenge in the single-cell field. To utilize single-cell data more efficiently and to better explore the heterogeneity among cells, a new graph autoencoder (GAE)-based consensus-guided model (scGAC) is proposed in this article. The data are preprocessed into multiple top-level feature datasets. Then, feature learning is performed by using GAEs to generate new feature matrices, followed by similarity learning based on distance fusion methods. The learned similarity matrices are fed back to the GAEs to guide their feature learning process. Finally, the abovementioned steps are iterated continuously to integrate the final consistent similarity matrix and perform other related downstream analyses. The scGAC model can accurately identify critical features and effectively preserve the internal structure of the data. This can further improve the accuracy of cell type identification. Dai-Jun Zhang, Ying-Lian Gao, Jing-Xiu Zhao, Chun-Hou Zheng 0001, Jin-Xing Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | GRPGAT: Predicting CircRNA-disease Associations Based on Graph Random Propagation Network and Graph Attention NetworkabstractCircRNA as a biomarker has been shown to have an essential effect on the occurrence and prognosis of a wide range of human diseases. Because of the high cost of wet experiments, computational methods are widely used to explore circRNA. However, the performance and robustness of the computational models still need to be further improved. To solve these problems, this paper proposes a novel method based on graph random propagation network and multi-head dynamic graph attention network (GRPGAT) to predict the potential associations between circRNAs and diseases. Firstly, GRPGAT uses centered kernel alignment method to fuse the circRNA similarity kernels and disease similarity kernels. Then the integrated vectors build a heterogeneous graph and are sent to a graph random propagation network. The remaining nodes are fed into a multi-head dynamic attention network for feature extraction. Finally, a four-layer Multilayer Perceptron is used to learn features and gain the prediction scores. Experiments are supported by cirR2Disease, and achieve Area Under Curve (AUC) scores of 0.9636 in 5-fold cross validation. In comparison with the state-of-the-art models, GRPGAT also shows superior performance. Wen-Yue Kang, Chun-Hou Zheng 0001, Ying-Lian Gao, Juan Wang 0003, Junliang Shang, Jin-Xing Liu 0001 |
BIBM | 3 |
| 2023 | MKGSAGE: A Computational Framework via Multiple Kernel Fusion on GraphSAGE for Inferring Potential Disease-Related MicrobesabstractMicrobes play a crucial role within the human body and are closely associated with the occurrence and development of numerous diseases. Studies have shown that disruptions in the composition and functionality of microbes can lead to immune system imbalances, inflammatory responses, and subsequently impact human health. Therefore, developing computational models to discover the potential connections between microbes and diseases is currently a hot topic. In this paper, a computational framework based on the multiple kernel fusion of graph embedding with sampling and aggregation (GraphSAGE) and dual Laplace regularized least squares called MKGSAGE is proposed for predicting potential links between microbe and disease. First, multiple layers embedding features of microbe and disease are learned from the initial input features by GraphSAGE. The kernel matrices are then calculated separately for each layer based on the Gaussian interaction profile (GIP). Furthermore, the multiple kernel fusion method is proposed for fusing kernel matrices of each layer and the initial similarity matrix. Dual Laplacian regularized least squares are finally applied for potential microbe-disease association prediction. Compared with six state-of-the-art methods on the HMDAD dataset, 5-fold cross-validations show that MKGSAGE performs best. In addition, case studies on asthma and inflammatory bowel disease further validate the effectiveness of MKGSAGE on discovering novel microbe-disease associations. Jin-Xing Liu 0001, Bao-Min Liu, Ling-Yun Dai, Feng Li 0033, Ying-Lian Gao |
BIBM | 6 |
| 2023 | LANCMDA: Predicting MiRNA-Disease Associations via LightGBM with Attributed Network Construction
Xu-Ran Dou, Wen-Yu Xi, Tian-Ru Wu, Cui-Na Jiao, Jin-Xing Liu 0001, Ying-Lian Gao |
ICIC (3) | 6 |
| 2023 | Spatial Domain Identification Based on Graph Attention Denoising Auto-encoder
Dai-Jun Zhang, Cui-Na Jiao, Ying-Lian Gao, Jin-Xing Liu 0001 |
ICIC (3) | 4 |
| 2023 | Identify Complex Higher-Order Associations Between Alzheimer's Disease Genes and Imaging Markers Through Improved Adaptive Sparse Multi-view Canonical Correlation Analysis
Xiang-Zhen Kong, Boxin Guan, Chun-Hou Zheng 0001, Ying-Lian Gao |
ICIC (3) | 5 |
| 2023 | MSF-LRR: Multi-Similarity Information Fusion Through Low-Rank Representation to Predict Disease-Associated MicrobesabstractAn Increase in microbial activity is shown to be intimately connected with the pathogenesis of diseases. Considering the expense of traditional verification methods, researchers are working to develop high-efficiency methods for detecting potential disease-related microbes. In this article, a new prediction method, MSF-LRR, is established, which uses Low-Rank Representation (LRR) to perform multi-similarity information fusion to predict disease-related microbes. Considering that most existing methods only use one class of similarity, three classes of microbe and disease similarity are added. Then, LRR is used to obtain low-rank structural similarity information. Additionally, the method adaptively extracts the local low-rank structure of the data from a global perspective, to make the information used for the prediction more effective. Finally, a neighbor-based prediction method that utilizes the concept of collaborative filtering is applied to predict unknown microbe-disease pairs. As a result, the AUC value of MSF-LRR is superior to other existing algorithms under 5-fold cross-validation. Furthermore, in case studies, excluding originally known associations, 16 and 19 of the top 20 microbes associated with Bacterial Vaginosis and Irritable Bowel Syndrome, respectively, have been confirmed by the recent literature. In summary, MSF-LRR is a good predictor of potential microbe-disease associations and can contribute to drug discovery and biological research. Jin-Xing Liu 0001, Meng-Meng Yin, Ying-Lian Gao, Junliang Shang, Chun-Hou Zheng 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | Non-Negative Low-Rank Representation With Similarity Correction for Cell Type Identification in scRNA-Seq DataabstractSingle-cell RNA sequencing (scRNA-Seq) technology has emerged as a powerful tool to investigate cellular heterogeneity within tissues, organs, and organisms. One fundamental question pertaining to single-cell gene expression data analysis revolves around the identification of cell types, which constitutes a critical step within the data processing workflow. However, existing methods for cell type identification through learning low-dimensional latent embeddings often overlook the intercellular structural relationships. In this paper, we present a novel non-negative low-rank similarity correction model (NLRSIM) that leverages subspace clustering to preserve the global structure among cells. This model introduces a novel manifold learning process to address the issue of imbalanced neighbourhood spatial density in cells, thereby effectively preserving local geometric structures. This procedure utilizes a position-sensitive hashing algorithm to construct the graph structure of the data. The experimental results demonstrate that the NLRSIM surpasses other advanced models in terms of clustering effects and visualization experiments. The validated effectiveness of gene expression information after calibration by the NLRSIM model has been duly ascertained in the realm of relevant biological studies. The NLRSIM model offers unprecedented insights into gene expression, states, and structures at the individual cellular level, thereby contributing novel perspectives to the field. Jin-Xing Liu 0001, Dai-Jun Zhang, Jing-Xiu Zhao, Chun-Hou Zheng 0001, Ying-Lian Gao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2023 | LDCMFC: Predicting Long Non-Coding RNA and Disease Association Using Collaborative Matrix Factorization Based on CorrentropyabstractWith the development of bioinformatics, the important role played by lncRNAs in various intractable diseases has aroused the interest of many experts. In recent studies, researchers have found that several human diseases are related to lncRANs. Moreover, it is very difficult and expensive to explore the unknown lncRNA-disease associations (LDAs), so only a few associations have been confirmed. It is vital to find a more accurate and effective method to identify potential LDAs. In this study, a method of collaborative matrix factorization based on correntropy (LDCMFC) is proposed for the identification of potential LDAs. To improve the robustness of the algorithm, the traditional minimization of the Euclidean distance is replaced with the maximized correntropy. In addition, the weighted K nearest known neighbor (WKNKN) method is used to rebuild the adjacency matrix. Finally, the performance of LDCMFC is tested by 5-fold cross-validation. Compared with other traditional methods, LDACMFC obtains a higher AUC of 0.8628. In different types of studies of three important cancer cases, most of the potentially relevant lncRNAs derived from the experiments have been validated in the databases. The final result shows that LDCMFC is a feasible method to predict LDAs. Wen-Yu Xi, Feng Zhou 0021, Ying-Lian Gao, Jin-Xing Liu 0001, Chun-Hou Zheng 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | BioSTD: A New Tensor Multi-View Framework via Combining Tensor Decomposition and Strong Complementarity Constraint for Analyzing Cancer Omics DataabstractAdvances in omics technology have enriched the understanding of the biological mechanisms of diseases, which has provided a new approach for cancer research. Multi-omics data contain different levels of cancer information, and comprehensive analysis of them has attracted wide attention. However, limited by the dimensionality of matrix models, traditional methods cannot fully use the key high-dimensional global structure of multi-omics data. Moreover, besides global information, local features within each omics are also critical. It is necessary to consider the potential local information together with the high-dimensional global information, ensuring that the shared and complementary features of the omics data are comprehensively observed. In view of the above, this article proposes a new tensor integrative framework called the strong complementarity tensor decomposition model (BioSTD) for cancer multi-omics data. It is used to identify cancer subtype specific genes and cluster subtype samples. Different from the matrix framework, BioSTD utilizes multi-view tensors to coordinate each omics to maximize high-dimensional spatial relationships, which jointly considers the different characteristics of different omics data. Meanwhile, we propose the concept of strong complementarity constraint applicable to omics data and introduce it into BioSTD. Strong complementarity is used to explore the potential local information, which can enhance the separability of different subtypes, allowing consistency and complementarity in the omics data to be fully represented. Experimental results on real cancer datasets show that our model outperforms other advanced models, which confirms its validity. Ying-Lian Gao, Juan Wang 0003, Shasha Yuan, Jin-Xing Liu 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2023 | MSGCA: Drug-Disease Associations Prediction Based on Multi-Similarities Graph Convolutional AutoencoderabstractIdentifying drug-disease associations (DDAs) is critical to the development of drugs. Traditional methods to determine DDAs are expensive and inefficient. Therefore, it is imperative to develop more accurate and effective methods for DDAs prediction. Most current DDAs prediction methods utilize original DDAs matrix directly. However, the original DDAs matrix is sparse, which greatly affects the prediction consequences. Hence, a prediction method based on multi-similarities graph convolutional autoencoder (MSGCA) is proposed for DDAs prediction. First, MSGCA integrates multiple drug similarities and disease similarities using centered kernel alignment-based multiple kernel learning (CKA-MKL) algorithm to form new drug similarity and disease similarity, respectively. Second, the new drug and disease similarities are improved by linear neighborhood, and the DDAs matrix is reconstructed by weighted K nearest neighbor profiles. Next, the reconstructed DDAs and the improved drug and disease similarities are integrated into a heterogeneous network. Finally, the graph convolutional autoencoder with attention mechanism is utilized to predict DDAs. Compared with extant methods, MSGCA shows superior results on three datasets. Furthermore, case studies further demonstrate the reliability of MSGCA. Ying Wang 0143, Ying-Lian Gao, Juan Wang 0003, Feng Li 0033, Jin-Xing Liu 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | NTBiRW: A Novel Neighbor Model Based on Two-Tier Bi-Random Walk for Predicting Potential Disease-Related MicrobesabstractStudies have revealed that microbes have an important effect on numerous physiological processes, and further research on the links between diseases and microbes is significant. Given that laboratory methods are expensive and not optimized, computational models are increasingly used for discovering disease-related microbes. Here, a new neighbor approach based on two-tier Bi-Random Walk is proposed for potential disease-related microbes, known as NTBiRW. In this method, the first step is to construct multiple microbe similarities and disease similarities. Then, three kinds of microbe/disease similarity are integrated through two-tier Bi-Random Walk to obtain the final integrated microbe/disease similarity network with different weights. Finally, Weighted K Nearest Known Neighbors (WKNKN) is used for prediction based on the final similarity network. In addition, leave-one-out cross-validation (LOOCV) and 5-fold cross-validation (5-fold CV) are applied for evaluating the performance of NTBiRW. Multiple evaluating indicators are taken to show the performance from multiple perspectives. And most of the evaluation index values of NTBiRW are better than those of the compared methods. Moreover, in case studies on atopic dermatitis and psoriasis, most of the first 10 candidates in the final result can be proven. This also demonstrates the capability of NTBiRW for discovering new associations. Therefore, this method can contribute to the discovery of disease-related microbes and thus offer new thoughts for further understanding the pathogenesis of diseases. Meng-Meng Yin, Ying-Lian Gao, Chun-Hou Zheng 0001, Jin-Xing Liu 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | An integrated Extreme learning machine based on kernel risk-sensitive loss of q-Gaussian and voting mechanism for sample classificationabstractEnsemble learning is to train and combine multiple learners to complete the corresponding learning tasks. It can improve the stability of the overall model, and a good ensemble method can further improve the accuracy of the model. At the same time, as one of the outstanding representatives of machine learning, Extreme Learning Machine has attracted the continuous attention of experts and scholars. to get a better representation of the feature space, we extend the Gaussian kernel in the kernel risk-sensitive loss and propose a Kernel Risk-Sensitive Loss of q-Gaussian kernel and Hyper-graph Regularized Extreme Learning Machine method. Since the contingency in the ELM training process cannot be completely avoided, the stability of most ELM methods is affected to some extent. What’s more, we introduce the voting mechanism and a new ELM classification model named Kernel Risk-Sensitive Loss of q-Gaussian kernel and Hyper-graph Regularized Integrated Extreme Learning Machine based on Voting Mechanism is proposed. It improves the stability of the model through the idea of ensemble learning. We apply the new model on six real data sets, and through observation and analysis of experimental results, we find that the new model has certain competitiveness, especially in classification accuracy and stability. Ying-Lian Gao, Zhen-Xin Niu, Shasha Yuan, Chun-Hou Zheng 0001, Jin-Xing Liu 0001 |
BIBM | 2 |
| 2022 | HSAELDA: Predicting lncRNA-disease associations based on heterogeneous networks and Stacked AutoencoderabstractIt is well known that the study of the lncRNA-disease associations (LDAs) is of great value for the diagnosis and cure of many complex diseases. However, exploring unknown LDAs is extremely difficult and expensive. Therefore, it is essential to find a more accurate and effective calculation method to predict the potential LDAs. However, most previous studies focused on designing complex similarity-based methods to predict the potential interaction between lncRNAs and diseases. In this research, combining the three biological networks of lncRNA-disease, miRNA-lncRNA and miRNA-disease, a new computing model based on heterogeneous networks and stacked autoencoder (SAE) is proposed, called HSAELDA. Then, the SAE is used to extract the comprehensive features of the lncRNA-disease pairs, the LightGBM classifier is used for training. At the same time, five-fold cross-validation (CV) is used to compare our model with some existing prediction methods. The final comparison results showed HSAELDA obtained the highest AUC value of 0.978. In conclusion, the overall prediction performance of HSAELDA has been greatly improved compared to the state-of-art models. Experimental results and case study results show that HSAELDA is an effective method for predicting potential LDAs. Wen-Yu Xi, Qianqian Ren, Jin-Xing Liu 0001, Ying-Lian Gao |
BIBM | 4 |
| 2022 | Predicting LncRNA-Disease Associations Based on LncRNA-MiRNA-Disease Multilayer Association Network and Bipartite Network RecommendationabstractThe pathogenesis of many human diseases is unclear, but many studies have shown that lncRNAs are deeply involved in the development of diseases. However, the exploration of lncRNA-disease associations in the laboratory requires a lot of time and financial resources, and computational-based methods have obvious advantages and become a promising research direction. But few experiments consider the relationship between other biological factors and lncRNAs and diseases. In this paper, a novel lncRNA-disease association prediction method, MANBNR, is proposed. MiRNAs are introduced by MANBNR to construct a lncRNA-miRNA-disease multilayer association network (MAN). The main innovation of MANBNR is to mine potential lncRNA-disease association information based on miRNA information. For any lncRNA-disease pair, lncRNA-associated miRNAs and disease-associated miRNAs are sorted into two sets, respectively. The association of this lncRNA-disease pair is judged by comparing the number of miRNAs shared in the two sets. This solves the problem that the known lncRNA-disease association matrices are too sparse. Finally, the bipartite network recommendation (BNR) algorithm was used to accurately predict potential lncRNA-disease association. The performance of MANBNR is better than that of many advanced methods at present. Case studies of breast cancer and lung cancer further demonstrate that MANBNR is an effective and reliable method for LDAs prediction. Guozheng Zhang, Shu-Zhen Li, Xu-Ran Dou, Junliang Shang, Qianqian Ren, Ying-Lian Gao |
BIBM | 6 |
| 2022 | A Locality-Constrained Linear Coding-Based Ensemble Learning Framework for Predicting Potentially Disease-Associated MiRNAs
Ying-Lian Gao, Shu-Zhen Li, Boxin Guan, Jin-Xing Liu 0001 |
ISBRA | 2 |
| 2022 | MLMVFE: A Machine Learning Approach Based on Muli-view Features Extraction for Drug-Disease Associations Prediction
Ying Wang 0143, Ying-Lian Gao, Juan Wang 0003, Junliang Shang, Jin-Xing Liu 0001 |
ISBRA | 2 |
| 2022 | A new framework for drug-disease association prediction combing light-gated message passing neural network and gated fusion mechanismabstractWith the development of research on the complex aetiology of many diseases, computational drug repositioning methodology has proven to be a shortcut to costly and inefficient traditional methods. Therefore, developing more promising computational methods is indispensable for finding new candidate diseases to treat with existing drugs. In this paper, a model integrating a new variant of message passing neural network and a novel-gated fusion mechanism called GLGMPNN is proposed for drug-disease association prediction. First, a light-gated message passing neural network (LGMPNN), including message passing, aggregation and updating, is proposed to separately extract multiple pieces of information from the similarity networks and the association network. Then, a gated fusion mechanism consisting of a forget gate and an output gate is applied to integrate the multiple pieces of information to extent. The forget gate calculated by the multiple embeddings is built to integrate the association information into the similarity information. Furthermore, the final node representations are controlled by the output gate, which fuses the topology information of the networks and the initial similarity information. Finally, a bilinear decoder is adopted to reconstruct an adjacency matrix for drug-disease associations. Evaluated by 10-fold cross-validations, GLGMPNN achieves excellent performance compared with the current models. The following studies show that our model can effectively discover novel drug-disease associations. Bao-Min Liu, Ying-Lian Gao, Dai-Jun Zhang, Feng Zhou 0021, Juan Wang 0003, Chun-Hou Zheng 0001, Jin-Xing Liu 0001 |
Briefings Bioinform. | 2 |
| 2022 | Multi-similarity fusion-based label propagation for predicting microbes potentially associated with diseases
Meng-Meng Yin, Ying-Lian Gao, Junliang Shang, Chun-Hou Zheng 0001, Jin-Xing Liu 0001 |
Future Gener. Comput. Syst. | 2 |
| 2022 | Robust Principal Component Analysis Based On Hypergraph Regularization for Sample Clustering and Co-Characteristic Gene SelectionabstractExtracting genes involved in cancer lesions from gene expression data is critical for cancer research and drug development. The method of feature selection has attracted much attention in the field of bioinformatics. Principal Component Analysis (PCA) is a widely used method for learning low-dimensional representation. Some variants of PCA have been proposed to improve the robustness and sparsity of the algorithm. However, the existing methods ignore the high-order relationships between data. In this paper, a new model named Robust Principal Component Analysis via Hypergraph Regularization (HRPCA) is proposed. In detail, HRPCA utilizes L2,1-norm to reduce the effect of outliers and make data sufficiently row-sparse. And the hypergraph regularization is introduced to consider the complex relationship among data. Important information hidden in the data are mined, and this method ensures the accuracy of the resulting data relationship information. Extensive experiments on multi-view biological data demonstrate that the feasible and effective of the proposed approach. Ying-Lian Gao, Ming-Juan Wu, Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Juan Wang 0003 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2022 | Single-Cell RNA Sequencing Data Clustering by Low-Rank Subspace Ensemble FrameworkabstractThe rapid development of single-cell RNA sequencing (scRNA-seq)technology reveals the gene expression status and gene structure of individual cells, reflecting the heterogeneity and diversity of cells. The traditional methods of scRNA-seq data analysis treat data as the same subspace, and hide structural information in other subspaces. In this paper, we propose a low-rank subspace ensemble clustering framework (LRSEC)to analyze scRNA-seq data. Assuming that the scRNA-seq data exist in multiple subspaces, the low-rank model is used to find the lowest rank representation of the data in the subspace. It is worth noting that the penalty factor of the low-rank kernel function is uncertain, and different penalty factors correspond to different low-rank structures. Moreover, the single cluster model is difficult to find the cellular structure of all datasets. To strengthen the correlation between model solutions, we construct a new ensemble clustering framework LRSEC by using the low-rank model as the basic learner. The LRSEC framework captures the global structure of data through low-rank subspaces, which has better clustering performance than a single clustering model. We validate the performance of the LRSEC framework on seven small datasets and one large dataset and obtain satisfactory results. Chuan-Yuan Wang, Ying-Lian Gao, Jin-Xing Liu 0001, Xiang-Zhen Kong, Chun-Hou Zheng 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | NCPLP: A Novel Approach for Predicting Microbe-Associated Diseases With Network Consistency Projection and Label PropagationabstractA growing number of clinical studies have provided substantial evidence of a close relationship between the microbe and the disease. Thus, it is necessary to infer potential microbe-disease associations. But traditional approaches use experiments to validate these associations that often spend a lot of materials and time. Hence, more reliable computational methods are expected to be applied to predict disease-associated microbes. In this article, an innovative mean for predicting microbe-disease associations is proposed, which is based on network consistency projection and label propagation (NCPLP). Given that most existing algorithms use the Gaussian interaction profile (GIP) kernel similarity as the similarity criterion between microbe pairs and disease pairs, in this model, Medical Subject Headings descriptors are considered to calculate disease semantic similarity. In addition, 16S rRNA gene sequences are borrowed for the calculation of microbe functional similarity. In view of the gene-based sequence information, we use two conventional methods (BLAST+ and MEGA7) to assess the similarity between each pair of microbes from different perspectives. Especially, network consistency projection is added to obtain network projection scores from the microbe space and the disease space. Ultimately, label propagation is utilized to reliably predict microbes related to diseases. NCPLP achieves better performance in various evaluation indicators and discovers a greater number of potential associations between microbes and diseases. Also, case studies further confirm the reliable prediction performance of NCPLP. To conclude, our algorithm NCPLP has the ability to discover these underlying microbe-disease associations and can provide help for biological study. Meng-Meng Yin, Jin-Xing Liu 0001, Ying-Lian Gao, Xiang-Zhen Kong, Chun-Hou Zheng 0001 |
IEEE Trans. Cybern. | 3 |
| 2022 | Unsupervised Cluster Analysis and Gene Marker Extraction of scRNA-seq Data Based On Non-Negative Matrix FactorizationabstractThe development of single-cell RNA sequencing (scRNA-seq) technology has made it possible to measure gene expression levels at the resolution of a single cell, which further reveals the complex growth processes of cells such as mutation and differentiation. Recognizing cell heterogeneity is one of the most critical tasks in scRNA-seq research. To solve it, we propose a non-negative matrix factorization framework based on multi-subspace cell similarity learning for unsupervised scRNA-seq data analysis (MscNMF). MscNMF includes three parts: data decomposition, similarity learning, and similarity fusion. The three work together to complete the data similarity learning task. MscNMF can learn the gene features and cell features of different subspaces, and the correlation and heterogeneity between cells will be more prominent in multi-subspaces. The redundant information and noise in each low-dimensional feature space are eliminated, and its gene weight information can be further analyzed to calculate the optimal number of subpopulations. The final cell similarity learning will be more satisfactory due to the fusion of cell similarity information in different subspaces. The advantage of MscNMF is that it can calculate the number of cell types and the rank of Non-negative matrix factorization (NMF) reasonably. Experiments on eight real scRNA-seq datasets show that MscNMF can effectively perform clustering tasks and extract useful genetic markers. To verify its clustering performance, the framework is compared with other latest clustering algorithms and satisfactory results are obtained. The code of MscNMF is free available for academic (https://github.com/wangchuanyuan1/project-MscNMF). Chuan-Yuan Wang, Ying-Lian Gao, Xiang-Zhen Kong, Jin-Xing Liu 0001, Chun-Hou Zheng 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | Sparse Hyper-graph Non-negative Matrix Factorization by Maximizing CorrentropyabstractNon-negative Matrix Factorization (NMF) as a powerful dimension reduction tool, which is widely used in the bioinformatics field. However, the loss function of conventional NMF is sensitive to non-Gaussian noise and outliers. In addition, NMF-based algorithm overlooks the geometric structure of high dimensional data. To improve the robustness of NMF, we propose a novel method called Sparse Hyper-graph regularized Non-negative Matrix Factorization by Maximizing Correntropy (SHNMF-MCC) in this paper. Specifically, the maximum correntropy criterion replaces the Euclidean distance in the loss term of SHNMF-MCC, which can filter out the noise with large outliers. Moreover, the high-order geometric structure in more sample points is completely preserved in the low-dimensional manifold through the hyper-graph regularization. Meanwhile, the sparse constraint is applied to the loss function to reduce matrix complexity and analysis difficulty. Then, the complex optimization problem can be solved by a half-quadratic (HQ) optimization approach. Before carrying out experiments, we analyze the convergence of SHNMF-MCC. Sample clustering experiments on The Cancer Genome Atlas (TCGA) data and single cell RNA-sequencing (scRNA-seq) data verify that the proposed method is more robust and effective than other similar robust approaches. Cui-Na Jiao, Jin-Xing Liu 0001, Ying-Lian Gao, Xiang-Zhen Kong, Chun-Hou Zheng 0001, Xianzi Yu |
BIBM | 3 |
| 2021 | Robust Tensor Method Based on Correntropy and Tensor Singular Value Decomposition for Cancer Genomics DataabstractThe analysis of biological sequencing data can provide significant support for researchers to unravel the mysteries of life further. This paper proposes a robust tensor data analysis method based on correntropy and tensor singular value decomposition (t-SVD) (CoTD) to analyze high-dimensional and multi-way cancer genomics data. CoTD uses the maximum correntropy criterion to increase the sparsity of the sparse tensor and fully exploits the vital information of the tensor data. It can effectively suppress outliers in the process of recovering low-rank and separating sparse data. In addition, through t-SVD, the internal spatial structure of the original tensor data can be well preserved. In this way, essential information can be retained in the low-rank part, which increases the clustering effect. The CoTD model is optimized by the half-quadratic technique and alternating direction method of multipliers (ADMM). Sample clustering and differentially expressed gene (DEG) extraction experiments are carried out on cancer genomics datasets. CoTD model is compared with four similar methods, which proves that the CoTD model has good performance. Ying-Lian Gao, Shasha Yuan, Jin-Xing Liu 0001 |
BIBM | 2 |
| 2021 | Adaptive total-variation joint learning model for analyzing single cell RNA seq dataabstractAn important purpose of single-cell RNA sequencing (scRNA-seq) data research is to explain the complex and diverse heterogeneity information between cells, which can further deepen human understanding of the mechanisms of life and the organization of organisms. However, the high dimensionality and noise are two major factors that hinder the development of scRNA-seq data mining. Therefore, in this paper, an adaptive total-variant joint learning model (JL-ATV) is proposed to overcome these two drawbacks of scRNA-seq data mining. On the one hand, in this model, dimensionality reduction learning and segmentation reconstruction subspace methods is combined to obtain effective features descriptions of the scRNAseq data and improve the interpretability and accuracy of cell identification. On the other hand, a gradient-based learning approach, namely adaptive total variation (ATV), is applied to scRNA-seq data to preserve the internal structure and overcome the interference of noise. Finally, experiments on multiple datasets show that the JL-ATV model can obtain a set of effective features and further improve the accuracy of identifying cell types. Dai-Jun Zhang, Jing-Xiu Zhao, Jin-Xing Liu 0001, Ying-Lian Gao |
BIBM | 4 |
| 2021 | Extreme Learning Machine Based on Double Kernel Risk-Sensitive Loss for Cancer Samples Classification
Zhen-Xin Niu, Liangrui Ren, Xiang-Zhen Kong, Ying-Lian Gao, Jin-Xing Liu 0001 |
ICIC (2) | 5 |
| 2021 | MKL-LP: Predicting Disease-Associated Microbes with Multiple-Similarity Kernel Learning-Based Label Propagation
Ying-Lian Gao, Meng-Meng Yin, Jin-Xing Liu 0001, Junliang Shang, Chun-Hou Zheng 0001 |
ISBRA | 1 |
| 2021 | DSCMF: prediction of LncRNA-disease associations based on dual sparse collaborative matrix factorizationabstractBACKGROUND: In the development of science and technology, there are increasing evidences that there are some associations between lncRNAs and human diseases. Therefore, finding these associations between them will have a huge impact on our treatment and prevention of some diseases. However, the process of finding the associations between them is very difficult and requires a lot of time and effort. Therefore, it is particularly important to find some good methods for predicting lncRNA-disease associations (LDAs). RESULTS: -norm is added in our method. At the same time, Gaussian interaction profile kernel is added to our method, which increase the network similarity between lncRNA and disease. Finally, the AUC value obtained by the experiment is used to evaluate the quality of our method, and the AUC value is obtained by the ten-fold cross-validation method. CONCLUSIONS: The AUC value obtained by the DSCMF method is 0.8523. At the end of the paper, simulation experiment is carried out, and the experimental results of prostate cancer, breast cancer, ovarian cancer and colorectal cancer are analyzed in detail. The DSCMF method is expected to bring some help to lncRNA-disease associations research. The code can access the https://github.com/Ming-0113/DSCMF website. Jin-Xing Liu 0001, Ming-Ming Gao, Ying-Lian Gao, Feng Li 0033 |
BMC Bioinform. | 4 |
| 2021 | Kernel Risk-Sensitive Loss based Hyper-graph Regularized Robust Extreme Learning Machine and Its Semi-supervised Extension for Classification
Liangrui Ren, Jin-Xing Liu 0001, Ying-Lian Gao, Xiang-Zhen Kong, Chun-Hou Zheng 0001 |
Knowl. Based Syst. | 3 |
| 2021 | DSTPCA: Double-Sparse Constrained Tensor Principal Component Analysis Method for Feature SelectionabstractThe identification of differentially expressed genes plays an increasingly important role biologically. Therefore, the feature selection approach has attracted much attention in the field of bioinformatics. The most popular method of principal component analysis studies two-dimensional data without considering the spatial geometric structure of the data. The recently proposed tensor robust principal component analysis method performs sparse and low-rank decomposition on three-dimensional tensors and effectively preserves the spatial structure. Based on this approach, the$L_{2,1}$- norm regularization term is introduced into the DSTPCA (Double-Sparse Constrained Tensor Principal Component Analysis) method. The DSTPCA method removes the redundant noise by double sparse constraints on the objective function to obtain sufficiently sparse results. After the regularization norm is introduced into the model, the ADMM (alternating direction method of multipliers) algorithm is used to solve the optimal problem. In the experiment of feature selection, while the more redundant genes were filtered out, the more genes closely associated with disease were screened. Experimental results using different datasets indicate that our method outperforms other methods. Yue Hu 0017, Jin-Xing Liu 0001, Ying-Lian Gao, Junliang Shang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | Dual Hyper-Graph Regularized Supervised NMF for Selecting Differentially Expressed Genes and Tumor ClassificationabstractNon-negative matrix factorization (NMF) is a dimensionality reduction technique based on high-dimensional mapping. It can learn part-based representations effectively. In this paper, we propose a method called Dual Hyper-graph Regularized Supervised Non-negative Matrix Factorization (HSNMF). To encode the geometric information of the data, the hyper-graph is introduced into the model as a regularization term. The advantage of hyper-graph learning is to find higher order data relationship to enhance data relevance. This method constructs the data hyper-graph and the feature hyper-graph to find the data manifold and the feature manifold simultaneously. The application of hyper-graph theory in cancer datasets can effectively find pathogenic genes. The discrimination information is further introduced into the objective function to obtain more information about the data. Supervised learning with label information greatly improves the classification effect. Furthermore, the real datasets of cancer usually contain sparse noise, so the$L_{2,1}$-norm is applied to enhance the robustness of HSNMF algorithm. Experiments under The Cancer Genome Atlas (TCGA) datasets verify the feasibility of the HSNMF method. Chuan-Yuan Wang, Na Yu 0004, Ming-Juan Wu, Ying-Lian Gao, Jin-Xing Liu 0001, Juan Wang 0003 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2021 | LWPCMF: Logistic Weighted Profile-Based Collaborative Matrix Factorization for Predicting MiRNA-Disease AssociationsabstractAs is known to all, constructing experiments to predict unknown miRNA-disease association is time-consuming, laborious and costly. Accordingly, new prediction model should be conducted to predict novel miRNA-disease associations. What's more, the performance of this method should be high and reliable. In this paper, a new computation model Logistic Weighted Profile-based Collaborative Matrix Factorization (LWPCMF) is put forward. In this method, weighted profile (WP) is combined with collaborative matrix factorization (CMF) to increase the performance of this model. And, the neighbor information is considered. In addition, logistic function is applied to miRNA functional similarity matrix and disease semantic similarity matrix to extract valuable information. At the same time, by adding WP and logistic function, the known correlation can be protected. And, Gaussian Interaction Profile (GIP) kernels of miRNAs and diseases are added to miRNA functional similarity network and disease semantic similarity network to augment kernel similarities. Then, a five-fold cross validation is implemented to evaluate the predictive ability of this method. Besides, case studies are conducted to view the experimental results. The final result contains not only known associations but also newly predicted ones. And, the result proves that our method is better than other existing methods. This model is able to predict potential miRNA-disease associations. Meng-Meng Yin, Ming-Ming Gao, Jin-Xing Liu 0001, Ying-Lian Gao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2021 | Multi-Label Fusion Collaborative Matrix Factorization for Predicting LncRNA-Disease AssociationsabstractAs we all know, science and technology are developing faster and faster. Many experts and scholars have demonstrated that human diseases are related to lncRNA, but only a few associations have been confirmed, and many unknown associations need to be found. In the process of finding associations, it takes a lot of time, so finding an efficient way to predict the associations between lncRNAs and diseases is particularly important. In this paper, we propose a multi-label fusion collaborative matrix factorization (MLFCMF) approach for predicting lncRNA-disease associations (LDAs). Firstly, the lncRNA space and disease space are optimized by multi-label to enhance the intrinsic link between lncRNA and disease and to tap potential information. Multi-label learning can encode a variety of data information from the sample space. Secondly, to learn multi-label information in the data space, the fusion method is used to handle the relationship between multiple labels. More comprehensive information will be obtained by weighing the effects of different labels. The addition of Gaussian interaction profile (GIP) kernel can increase the network similarity. Finally, the lncRNA-disease associations are predicted by the method of collaborative matrix factorization. The ten-fold cross-validation method is used to evaluate the MLFCMF method, and our method finally obtains an AUC value of 0.8612. Detailed analysis of ovarian cancer, colorectal cancer, and lung cancer in the simulation experiment results. So it can be seen that our method MLFCMF is an effective model for predicting lncRNA-disease associations. Ming-Ming Gao, Ying-Lian Gao, Juan Wang 0003, Jin-Xing Liu 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | WGRCMF: A Weighted Graph Regularized Collaborative Matrix Factorization Method for Predicting Novel LncRNA-Disease AssociationsabstractIn recent years, many human diseases have been determined to be associated with certain lncRNAs. Only a small percentage of all lncRNA-disease associations (LDAs) have been discovered by researchers. Predicting novel LDAs is time-consuming and costly. It is crucial to propose a method that can effectively identify potential LDAs to solve this problem based on the available datasets. Although some current methods can effectively predict potential LDAs, the prediction accuracy needs to be improved, and there are few known associations. Moreover, there are notable errors in the method of constructing the network and the bipartite graph, which interfere with the final results. A weighted graph regularized collaborative matrix factorization (WGRCMF) method is proposed to predict novel LDAs. We introduce the graph regularization terms into the collaborative matrix factorization. Considering that manifold learning can recover low-dimensional manifold structures from high-dimensional sampled data, we can find low-dimensional manifolds in high-dimensional space. In addition, a weight matrix is also introduced into the method, the significance of which is to prevent unknown associations from contributing to the final prediction matrix. Finally, the prediction accuracy of this method is better than those of other methods. In several cancer cases, we implemented the corresponding simulation experiments. According to the experimental results, the proposed method is feasible and effective. Jin-Xing Liu 0001, Ying-Lian Gao, Xiang-Zhen Kong |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Locally Manifold Non-negative Matrix Factorization Based on Centroid for scRNA-seq Data AnalysisabstractThe rapid development of single cell RNA sequencing (scRNA-seq) has made it possible to study the association between cells and genes at molecular resolution. When the follow-up analysis is carried out, it is often difficult to extract the cell information in high-dimensional space because of the high gene dimension in single-cell sequencing, which leads to inaccurate results in the follow-up analysis. To solve the problem, we propose a method called locally manifold non-negative matrix factorization based on centroid for scRNA-seq data analysis (MNMFC). MNMFC is a similarity modeling scheme based on locally manifold, which can map cell association in high dimensional space. Through similarity learning based on locally manifold and non-negative matrix decomposition (NMF) algorithm, the data in high-dimensional space can be mapped to low-dimensional space, which provides help for downstream clustering analysis. The performance of the model was validated experimentally on 10 scRNA-seq datasets. Compared with other nine advanced single-cell clustering methods, whether it is a comprehensive analysis or an individual analysis of the dataset, MNMFC has achieved encouraging results. Chuan-Yuan Wang, Ying-Lian Gao, Cui-Na Jiao, Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Xiang-Zhen Kong |
BIBM | 2 |
| 2020 | Robust Graph Regularized Extreme Learning Machine Auto Encoder and Its Application to Single-Cell Samples Classification
Liangrui Ren, Jin-Xing Liu 0001, Ying-Lian Gao, Xiang-Zhen Kong, Chun-Hou Zheng 0001 |
ICIC (2) | 3 |
| 2020 | Correntropy induced loss based sparse robust graph regularized extreme learning machine for cancer classificationabstractAbstract Background As a machine learning method with high performance and excellent generalization ability, extreme learning machine (ELM) is gaining popularity in various studies. Various ELM-based methods for different fields have been proposed. However, the robustness to noise and outliers is always the main problem affecting the performance of ELM. Results In this paper, an integrated method named correntropy induced loss based sparse robust graph regularized extreme learning machine (CSRGELM) is proposed. The introduction of correntropy induced loss improves the robustness of ELM and weakens the negative effects of noise and outliers. By using the L2,1-norm to constrain the output weight matrix, we tend to obtain a sparse output weight matrix to construct a simpler single hidden layer feedforward neural network model. By introducing the graph regularization to preserve the local structural information of the data, the classification performance of the new method is further improved. Besides, we design an iterative optimization method based on the idea of half quadratic optimization to solve the non-convex problem of CSRGELM. Conclusions The classification results on the benchmark dataset show that CSRGELM can obtain better classification results compared with other methods. More importantly, we also apply the new method to the classification problems of cancer samples and get a good classification effect. Liangrui Ren, Ying-Lian Gao, Jin-Xing Liu 0001, Junliang Shang, Chun-Hou Zheng 0001 |
BMC Bioinform. | 2 |
| 2020 | MCCMF: collaborative matrix factorization based on matrix completion for predicting miRNA-disease associationsabstractBACKGROUND: MicroRNAs (miRNAs) are non-coding RNAs with regulatory functions. Many studies have shown that miRNAs are closely associated with human diseases. Among the methods to explore the relationship between the miRNA and the disease, traditional methods are time-consuming and the accuracy needs to be improved. In view of the shortcoming of previous models, a method, collaborative matrix factorization based on matrix completion (MCCMF) is proposed to predict the unknown miRNA-disease associations. RESULTS: The complete matrix of the miRNA and the disease is obtained by matrix completion. Moreover, Gaussian Interaction Profile kernel is added to the miRNA functional similarity matrix and the disease semantic similarity matrix. Then the Weight K Nearest Known Neighbors method is used to pretreat the association matrix, so the model is close to the reality. Finally, collaborative matrix factorization method is applied to obtain the prediction results. Therefore, the MCCMF obtains a satisfactory result in the fivefold cross-validation, with an AUC of 0.9569 (0.0005). CONCLUSIONS: The AUC value of MCCMF is higher than other advanced methods in the fivefold cross validation experiment. In order to comprehensively evaluate the performance of MCCMF, accuracy, precision, recall and f-measure are also added. The final experimental results demonstrate that MCCMF outperforms other methods in predicting miRNA-disease associations. In the end, the effectiveness and practicability of MCCMF are further verified by researching three specific diseases. Tian-Ru Wu, Meng-Meng Yin, Cui-Na Jiao, Ying-Lian Gao, Xiang-Zhen Kong, Jin-Xing Liu 0001 |
BMC Bioinform. | 4 |
| 2020 | LncRNA-Disease Associations Prediction Using Bipartite Local Model With Nearest Profile-Based Association InferringabstractThere is much evidence that long non-coding RNA (lncRNA) is associated with many diseases. However, it is time-consuming and expensive to identify meaningful lncRNA-disease associations (LDAs) through medical or biological experiments. Therefore, investigating how to identify more meaningful LDAs is necessary, and at the same time it is conducive to the prevention, diagnosis and treatment of complex diseases. Considering the limitations of some current prediction models, a novel model based on bipartite local model with nearest profile-based association inferring, BLM-NPAI, is developed for predicting LDAs. This model predicts novel LDAs from the lncRNA side and the disease side, respectively. More importantly, for some lncRNAs and diseases without any association, the model can also be predicted by their nearest neighbors. Leave-one-out cross validation (LOOCV) and 5-fold cross validation are implemented for BLM-NPAI to evaluate the performance of this model. Our model is superior to current advanced methods in most cases. In addition, to verify the validity and reliability of BLM-NPAI, three disease cases and three lncRNA cases are analyzed to further evaluate BLM-NPAI. Finally, these predicted novel LDAs are confirmed by using the LncRNA-disease database. Jin-Xing Liu 0001, Ying-Lian Gao, Shasha Yuan |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Hyper-Graph Regularized Constrained NMF for Selecting Differentially Expressed Genes and Tumor ClassificationabstractNon-negative Matrix Factorization (NMF) is a dimensionality reduction approach for learning a parts-based and linear representation of non-negative data. It has attracted more attention because of that. In practice, NMF not only neglects the manifold structure of data samples, but also overlooks the priori label information of different classes. In this paper, a novel matrix decomposition method called Hyper-graph regularized Constrained Non-negative Matrix Factorization (HCNMF) is proposed for selecting differentially expressed genes and tumor sample classification. The advantage of hyper-graph learning is to capture local spatial information in high dimensional data. This method incorporates a hyper-graph regularization constraint to consider the higher order data sample relationships. The application of hyper-graph theory can effectively find pathogenic genes in cancer datasets. Besides, the label information is further incorporated in the objective function to improve the discriminative ability of the decomposition matrix. Supervised learning with label information greatly improves the classification effect. We also provide the iterative update rules and convergence proofs for the optimization problems of HCNMF. Experiments under The Cancer Genome Atlas (TCGA) datasets confirm the superiority of HCNMF algorithm compared with other representative algorithms through a set of evaluations. Cui-Na Jiao, Ying-Lian Gao, Na Yu 0004, Jin-Xing Liu 0001, Lianyong Qi |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | Integrative Hypergraph Regularization Principal Component Analysis for Sample Clustering and Co-Expression Genes Network Analysis on Multi-Omics DataabstractIn recent years, with the diversity and variability of cancer information, the multi-omics data have been applied in various fields. Many existing models of principal component analysis can only process single data, which makes limitations on cancer research. Therefore, in this paper, a new model called integrative principal component analysis (IPCA) is proposed to achieve the unification of multi-omics data. In addition, in order to preserve the high-order manifold structure between the data, an integrative hypergraph regularization principal component analysis (IHPCA) is further proposed by applying the hypergraph regularization constraint. The effectiveness of IHPCA method is tested on four multi-omics datasets. Experimental results show that the proposed method has better performance than other representative methods on sample clustering and common expression genes (co-expression genes) network analysis. Ming-Juan Wu, Ying-Lian Gao, Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Juan Wang 0003 |
IEEE J. Biomed. Health Informatics | 2 |
| 2019 | Dual Sparse Collaborative Matrix Factorization Method Based on Gaussian Kernel Function for Predicting LncRNA-Disease Associations
Ming-Ming Gao, Ying-Lian Gao, Feng Li 0033, Jin-Xing Liu 0001 |
ICIC (3) | 3 |
| 2019 | L2, 1-GRMF: an improved graph regularized matrix factorization method to predict drug-target interactionsabstractBACKGROUND: Predicting drug-target interactions is time-consuming and expensive. It is important to present the accuracy of the calculation method. There are many algorithms to predict global interactions, some of which use drug-target networks for prediction (ie, a bipartite graph of bound drug pairs and targets known to interact). Although these algorithms can predict some drug-target interactions to some extent, there is little effect for some new drugs or targets that have no known interaction. RESULTS: Since the datasets are usually located at or near low-dimensional nonlinear manifolds, we propose an improved GRMF (graph regularized matrix factorization) method to learn these flow patterns in combination with the previous matrix-decomposition method. In addition, we use one of the pre-processing steps previously proposed to improve the accuracy of the prediction. CONCLUSIONS: Cross-validation is used to evaluate our method, and simulation experiments are used to predict new interactions. In most cases, our method is superior to other methods. Finally, some examples of new drugs and new targets are predicted by performing simulation experiments. And the improved GRMF method can better predict the remaining drug-target interactions. Ying-Lian Gao, Jin-Xing Liu 0001, Ling-Yun Dai, Shasha Yuan |
BMC Bioinform. | 2 |
| 2019 | The computational prediction of drug-disease interactions using the dual-network L2,1-CMF methodabstractBACKGROUND: Predicting drug-disease interactions (DDIs) is time-consuming and expensive. Improving the accuracy of prediction results is necessary, and it is crucial to develop a novel computing technology to predict new DDIs. The existing methods mostly use the construction of heterogeneous networks to predict new DDIs. However, the number of known interacting drug-disease pairs is small, so there will be many errors in this heterogeneous network that will interfere with the final results. RESULTS: -norm are introduced in our method to achieve better results than other advanced methods. The network similarities of drugs and diseases with their chemical and semantic similarities are combined in this method. CONCLUSIONS: Cross validation is used to evaluate our method, and simulation experiments are used to predict new interactions using two different datasets. Finally, our prediction accuracy is better than other existing methods. This proves that our method is feasible and effective. Ying-Lian Gao, Jin-Xing Liu 0001, Juan Wang 0003, Junliang Shang, Ling-Yun Dai |
BMC Bioinform. | 2 |
| 2019 | RCMF: a robust collaborative matrix factorization method to predict miRNA-disease associationsabstractBACKGROUND: Predicting miRNA-disease associations (MDAs) is time-consuming and expensive. It is imminent to improve the accuracy of prediction results. So it is crucial to develop a novel computing technology to predict new MDAs. Although some existing methods can effectively predict novel MDAs, there are still some shortcomings. Especially when the disease matrix is processed, its sparsity is an important factor affecting the final results. RESULTS: -norm are introduced to our method to achieve the highest AUC value than other advanced methods. CONCLUSIONS: 5-fold cross validation is used to evaluate our method, and simulation experiments are used to predict novel associations on Gold Standard Dataset. Finally, our prediction accuracy is better than other existing advanced methods. Therefore, our approach is effective and feasible in predicting novel MDAs. Jin-Xing Liu 0001, Ying-Lian Gao, Chun-Hou Zheng 0001, Juan Wang 0003 |
BMC Bioinform. | 3 |
| 2019 | NPCMF: Nearest Profile-based Collaborative Matrix Factorization method for predicting miRNA-disease associationsabstractBACKGROUND: Predicting meaningful miRNA-disease associations (MDAs) is costly. Therefore, an increasing number of researchers are beginning to focus on methods to predict potential MDAs. Thus, prediction methods with improved accuracy are under development. An efficient computational method is proposed to be crucial for predicting novel MDAs. For improved experimental productivity, large biological datasets are used by researchers. Although there are many effective and feasible methods to predict potential MDAs, the possibility remains that these methods are flawed. RESULTS: A simple and effective method, known as Nearest Profile-based Collaborative Matrix Factorization (NPCMF), is proposed to identify novel MDAs. The nearest profile is introduced to our method to achieve the highest AUC value compared with other advanced methods. For some miRNAs and diseases without any association, we use the nearest neighbour information to complete the prediction. CONCLUSIONS: To evaluate the performance of our method, five-fold cross-validation is used to calculate the AUC value. At the same time, three disease cases, gastric neoplasms, rectal neoplasms and colonic neoplasms, are used to predict novel MDAs on a gold-standard dataset. We predict the vast majority of known MDAs and some novel MDAs. Finally, the prediction accuracy of our method is determined to be better than that of other existing methods. Thus, the proposed prediction model can obtain reliable experimental results. Ying-Lian Gao, Jin-Xing Liu 0001, Juan Wang 0003, Chun-Hou Zheng 0001 |
BMC Bioinform. | 1 |
| 2019 | Supervised Discriminative Sparse PCA for Com-Characteristic Gene Selection and Tumor Classification on Multiview Biological DataabstractPrincipal component analysis (PCA) has been used to study the pathogenesis of diseases. To enhance the interpretability of classical PCA, various improved PCA methods have been proposed to date. Among these, a typical method is the so-called sparse PCA, which focuses on seeking sparse loadings. However, the performance of these methods is still far from satisfactory due to their limitation of using unsupervised learning methods; moreover, the class ambiguity within the sample is high. To overcome this problem, this paper developed a new PCA method, which is named the supervised discriminative sparse PCA (SDSPCA). The main innovation of this method is the incorporation of discriminative information and sparsity into the PCA model. Specifically, in contrast to the traditional sparse PCA, which imposes sparsity on the loadings, here, sparse components are obtained to represent the data. Furthermore, via the linear transformation, the sparse components approximate the given label information. On the one hand, sparse components improve interpretability over the traditional PCA, while on the other hand, they are have discriminative abilities suitable for classification purposes. A simple algorithm is developed, and its convergence proof is provided. SDSPCA has been applied to the common-characteristic gene selection and tumor classification on multiview biological data. The sparsity and classification performance of SDSPCA are empirically verified via abundant, reasonable, and effective experiments, and the obtained results demonstrate that SDSPCA outperforms other state-of-the-art methods. Chun-Mei Feng 0001, Yong Xu 0001, Jin-Xing Liu 0001, Ying-Lian Gao, Chun-Hou Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | Hypergraph regularized NMF by L2, 1-norm for Clustering and Com-abnormal Expression Genes Selection
Na Yu 0004, Ying-Lian Gao, Jin-Xing Liu 0001, Juan Wang 0003, Junliang Shang |
BIBM | 2 |
| 2018 | Performance Analysis of Non-negative Matrix Factorization Methods on TCGA Data
Mi-Xiao Hou, Jin-Xing Liu 0001, Junliang Shang, Ying-Lian Gao, Ling-Yun Dai |
ICIC (2) | 4 |
| 2018 | Regularized Non-Negative Matrix Factorization for Identifying Differentially Expressed Genes and Clustering Samples: A SurveyabstractNon-negative Matrix Factorization (NMF), a classical method for dimensionality reduction, has been applied in many fields. It is based on the idea that negative numbers are physically meaningless in various data-processing tasks. Apart from its contribution to conventional data analysis, the recent overwhelming interest in NMF is due to its newly discovered ability to solve challenging data mining and machine learning problems, especially in relation to gene expression data. This survey paper mainly focuses on research examining the application of NMF to identify differentially expressed genes and to cluster samples, and the main NMF models, properties, principles, and algorithms with its various generalizations, extensions, and modifications are summarized. The experimental results demonstrate the performance of the various NMF algorithms in identifying differentially expressed genes and clustering samples. Jin-Xing Liu 0001, Dong Wang 0019, Ying-Lian Gao, Chun-Hou Zheng 0001, Yong Xu 0001, Jiguo Yu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2017 | Low-rank representation regularized by L2, 1-norm for identifying differentially expressed genesabstractLow-rank representation (LRR) via rank minimization is a high efficiency method for capturing low-dimensional structure embedded in high-dimensional data. However, minimizing the rank of a matrix is NP-hard. In this paper, robust truncated nuclear norm low-rank representation regularized by L2,1-norm method (RTLRR) is proposed. The truncated nuclear norm is introduced to replace the nuclear norm to approximate the rank function. At the same time, L2,1-norm is used to regularize the sparse matrix to achieve better sparse effect of the algorithm. The proposed method is divided into two steps. Firstly, we do singular value decomposition (SVD) to the original data matrix. Then we apply the truncated nuclear norm and L2,1-norm constraints to subproblems and use inexact augmented Lagrange multiplier method to solve subproblems. Finally, the genes with high scores will be identified as differentially expressed genes according to the sparse matrix. The results on The Cancer Genome Atlas (TCGA) data illustrate that the effectiveness of RTLRR method outperforms many methods. Yaxuan Wang, Jin-Xing Liu 0001, Ying-Lian Gao, Chun-Hou Zheng 0001, Ling-Yun Dai |
BIBM | 3 |
| 2017 | Feature selection and clustering via robust graph-laplacian PCA based on capped L1-normabstractIn molecular biology, the selection of feature genes and tumor clustering are the hotspots and difficulties in bioinformatics research. The traditional PCA method based on the minimization of the squares of the loss function is sensitive to the outliers and noise. Therefore, it is necessary to design a new method to weaken the effects of errors and noise. In this paper, we propose a novel PCA method based graph-Laplacian and capped L1-norm, which is called as CgLPCA. The method can preserve the internal geometry of data by introducing graph-Laplacian. In addition, it uses the capped L1-norm on loss function to improve its robustness. The main contribution of this method is to preserve the nonlinear structure of the data while enhancing the robustness of the PCA-based method. We introduce the Augmented Lagrangian multiplier to solve the optimization problem. The CgLPCA method achieves the advanced level in feature selection and tumor clustering among various PCA-based methods. Ming-Juan Wu, Jin-Xing Liu 0001, Ying-Lian Gao, Chun-Mei Feng 0001 |
BIBM | 3 |
| 2017 | Graph regularized robust non-negative matrix factorization for clustering and selecting differentially expressed genesabstractNon-negative Matrix Factorization (NMF) is widely used as a data dimensionality reduction tool. However, the assumption of most conventional NMF-based methods is that the gene expression data are only destroyed by Gaussian noise. In practice, the gene expression data are unavoidably destroyed by sparse noise. Although Sparsity-Regularized Robust NMF by using L1/2constraint (L1/2-RNMF) can achieve satisfactory results when the sparse noise exists, it does not consider the intrinsic geometric structure in data. Hence, we introduce graph regularization into L1/2-RNMF. In this paper, we developed a novel NMF method named Graph regularized Robust Nonnegative Matrix Factorization (GrRNMF), which mainly consists of two aspects: Firstly, the Gaussian noise and sparse noise are modeled, respectively. Secondly, it can reveal the geometric information in data by adding graph regularization term. Extensive experimental results on The Cancer Genome Atlas (TCGA) data indicate that the GrRNMF method has higher accuracy than other state-of-the-art methods in samples clustering and the selection of differentially expressed genes. Na Yu 0004, Jin-Xing Liu 0001, Ying-Lian Gao, Chun-Hou Zheng 0001, Juan Wang 0003, Ming-Juan Wu |
BIBM | 3 |
| 2017 | A joint-L2, 1-norm-constraint-based semi-supervised feature extraction for RNA-Seq data analysis
Jin-Xing Liu 0001, Dong Wang 0019, Ying-Lian Gao, Chun-Hou Zheng 0001, Junliang Shang, Feng Liu 0013, Yong Xu 0001 |
Neurocomputing | 3 |
| 2016 | A graph-Laplacian PCA based on L1/2-norm constraint for characteristic gene selectionabstractPrincipal Component Analysis (PCA) as a tool for dimensionality reduction is widely used in many areas. In the area of bioinformatics, the first principal component of PCA is used to select characteristic genes. In order to improve the robustness of PCA-based method, this paper proposes a novel graph-Laplacian PCA algorithm by adopting L1/2constraint on error function (L1/2gLPCA) for characteristic gene selection. Augmented Lagrange Multipliers (ALM) method is applied to solve the sub-problem. This method gets better results in characteristic gene selection than traditional PCA approach. Meanwhile, the error function based on the L1/2norm helps to reduce the influence of outliers and noise. Extensive experimental results on gene expression data sets demonstrate that our method can get higher identification accuracies than others. Chun-Mei Feng 0001, Jin-Xing Liu 0001, Ying-Lian Gao, Juan Wang 0003, Dong-Qin Wang |
BIBM | 3 |
| 2016 | Characteristic gene selection via L2, 1-norm Sparse Principal Component AnalysisabstractSparse Principal Component Analysis (SPCA) is a method that can get the sparse loadings of the principal components (PCs), and it may formulate PCA as a regression-type optimization problem by using the elastic net. But the selected features are different with each PC and generally independent. A new method named SPCA has been proposed for removing these detect, which replaces the elastic net with L2,1-norm penalty. The results of the method on gene expression data are still unknown. Therefore, we will take a test to prove this point in this paper. Firstly, this method is applied to the simulated data for obtaining an optimal parameter. Secondly, the L2,1SPCA method is applied to the gene expression data, that is the head and neck squamous carcinoma data (HNSC). Thirdly, the characteristic genes are selected according the PCs. The results consist of very lower P-value and very higher hit count, which shows the method of L2,1SPCA can obtain higher recognition accuracy and higher relevancy to the genes. Finally, the experimental results demonstrate that the L2,1SPCA works well and has good performances in the gene expression data. Yao Lu 0008, Ying-Lian Gao, Jin-Xing Liu 0001, Chang-Gang Wen, Yaxuan Wang, Jiguo Yu |
BIBM | 2 |
| 2016 | Differentially expressed genes selection via Truncated Nuclear Norm RegularizationabstractRobust Principal Component Analysis (RPCA) is an efficient method in the selection of differentially expressed genes. However, nuclear norm minimizes all singular values simultaneously, so it may not be the best solution to replace the low-rank function. In this paper, the truncated nuclear norm is introduced. And a new method named Truncated nuclear norm regularized Robust Principal Component Analysis (TRPCA) is proposed. The method decomposes the observation matrix of genomic data into a low-rank matrix and a sparse matrix. The differentially expressed genes can be selected according to the sparse matrix. The experimental results on the The Cancer Genome Atlas (TCGA) data illustrate that the TRPCA method outperforms other state-of-the-art methods in the selection of differentially expressed genes. Yaxuan Wang, Jin-Xing Liu 0001, Ying-Lian Gao, Chun-Hou Zheng 0001 |
BIBM | 3 |
| 2016 | L21-iPaD: An efficient method for drug-pathway association pairs inferenceabstractPathway-based drug discovery overcomes the disadvantages of the “one drug-one target” method, which aims to find the effective drugs to act on single targets. The current method “iPaD” identities the drug-pathway association pairs by taking the lasso-type penalty on the drug-pathway association matrix. In order to enhance the robustness of the methods and be more effective to find the novel drug-pathway association pairs, we introduce a new method named “L2,1-iPaD”. Compared with the iPaD method, we impose the L2,1-norm constraint on the drug-pathway association coefficient matrix. By applying our method to a real widely datasets (CCLE dataset), we demonstrate that our method is superior to the iPaD method. And our method can obtain the smaller P-values than the iPaD method by performing permutation test to assess the significance of the identified drug-pathway association pairs. More importantly, compared with the iPaD method, our method can identify larger numbers of validated drug-pathway association pairs. The experimental results on the real dataset demonstrate the effectiveness of our method. Dong-Qin Wang, Chun-Hou Zheng 0001, Ying-Lian Gao, Jin-Xing Liu 0001, Shasha Wu, Junliang Shang |
BIBM | 3 |
| 2016 | A Simple Review of Sparse Principal Components Analysis
Chun-Mei Feng 0001, Ying-Lian Gao, Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Shengjun Li, Dong Wang 0019 |
ICIC (2) | 2 |
| 2016 | Comparison of Non-negative Matrix Factorization Methods for Clustering Genomic Data
Mi-Xiao Hou, Ying-Lian Gao, Jin-Xing Liu 0001, Junliang Shang, Chun-Hou Zheng 0001 |
ICIC (2) | 2 |
| 2016 | A Class-Information-Based Sparse Component Analysis Method to Identify Differentially Expressed Genes on RNA-Seq DataabstractWith the development of deep sequencing technologies, many RNA-Seq data have been generated. Researchers have proposed many methods based on the sparse theory to identify the differentially expressed genes from these data. In order to improve the performance of sparse principal component analysis, in this paper, we propose a novel class-information-based sparse component analysis (CISCA) method which introduces the class information via a total scatter matrix. First, CISCA normalizes the RNA-Seq data by using a Poisson model to obtain their differential sections. Second, the total scatter matrix is gotten by combining the between-class and within-class scatter matrices. Third, we decompose the total scatter matrix by using singular value decomposition and construct a new data matrix by using singular values and left singular vectors. Then, aiming at obtaining sparse components, CISCA decomposes the constructed data matrix by solving an optimization problem with sparse constraints on loading vectors. Finally, the differentially expressed genes are identified by using the sparse loading vectors. The results on simulation and real RNA-Seq data demonstrate that our method is effective and suitable for analyzing these data. Jin-Xing Liu 0001, Yong Xu 0001, Ying-Lian Gao, Chun-Hou Zheng 0001, Dong Wang 0019, Qi Zhu 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2016 | Characteristic Gene Selection Based on Robust Graph Regularized Non-Negative Matrix FactorizationabstractMany methods have been considered for gene selection and analysis of gene expression data. Nonetheless, there still exists the considerable space for improving the explicitness and reliability of gene selection. To this end, this paper proposes a novel method named robust graph regularized non-negative matrix factorization for characteristic gene selection using gene expression data, which mainly contains two aspects: Firstly, enforcing L21-norm minimization on error function which is robust to outliers and noises in data points. Secondly, it considers that the samples lie in low-dimensional manifold which embeds in a high-dimensional ambient space, and reveals the data geometric structure embedded in the original data. To demonstrate the validity of the proposed method, we apply it to gene expression data sets involving various human normal and tumor tissue samples and the results demonstrate that the method is effective and feasible. Dong Wang 0019, Jin-Xing Liu 0001, Ying-Lian Gao, Chun-Hou Zheng 0001, Yong Xu 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2015 | A Two-Stage Sparse Selection Method for Extracting Characteristic Genes
Ying-Lian Gao, Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Shengjun Li, Yuxia Lei |
ICIC (2) | 1 |
| 2015 | Semi-supervised Feature Extraction for RNA-Seq Data Analysis
Jin-Xing Liu 0001, Yong Xu 0001, Ying-Lian Gao, Dong Wang 0019, Chun-Hou Zheng 0001, Junliang Shang |
ICIC (3) | 3 |
| 2015 | Graph Regularized Non-negative Matrix with L0-Constraints for Selecting Characteristic Genes
Chun-Xia Ma, Ying-Lian Gao, Dong Wang 0019, Jin-Xing Liu 0001 |
ICIC (2) | 2 |
| 2015 | Application of Graph Regularized Non-negative Matrix Factorization in Characteristic Gene Selection
Dong Wang 0019, Ying-Lian Gao, Jin-Xing Liu 0001, Jiguo Yu, Chang-Gang Wen |
ICIC (2) | 2 |