VLDB 2026 Research / reviewers in the wild / expert
Ling-Yun Dai
dblp:193/7857 · also Lingyun Dai
· DBLP profile ↗
39ranked-venue papers
5as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 37 · 5 first-author · 24 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Predicting the Drug Side Effect Frequency via a Kolmogorov-Arnold Graph Isomorphism Network and Subspace-Aware Dynamic Feature Fusion
Chenglong Mi, Feng Li 0033, Ling-Yun Dai |
ICIC (3) | 5 |
| 2026 | Single-cell multi-view clustering based on dual contrastive learning and cross-attention fusion
Meng-Yao Hu, Jin-Xing Liu 0001, Junliang Shang, Ling-Yun Dai |
Expert Syst. Appl. | 5 |
| 2026 | A Novel Low-Dimensional Sparse and Low-Rank Representation Method for Single-Cell RNA Sequencing Data ClusteringabstractThe advancement of single-cell RNA sequencing (scRNA-seq) technology has enabled researchers to capture cellular heterogeneity at the individual cell level, driving progress in diverse fields such as developmental biology, immunology, and cancer research. Accurate cell clustering is a crucial step for researchers utilizing scRNA-seq data; however, inherent characteristics like high dimensionality and sparsity pose significant challenges to obtaining precise clustering results. To achieve accurate clustering, this paper proposes a novel approach that integrates dimensionality reduction, self-representation matrix construction, and the clustering process into an end-to-end model termed LDSLRR (Low-Dimensional Sparse and Low-Rank Representation). Specifically, the original gene expression matrix first undergoes dimensionality reduction via projection. Subsequently, low-rank representation combined with a sparsity constraint facilitates the learning of the self-representation matrix. Finally, the cluster assignment matrix is acquired using graph-regularized non-negative matrix factorization (NMF). These three modules are simultaneously optimized, enhancing the accuracy of the clustering results. Comparative experiments against multiple state-of-the-art clustering methods on various scRNA-seq datasets demonstrate the superiority of the proposed LDSLRR method. Zhenduo Zhang, Junliang Shang, Ling-Yun Dai, Juan Wang 0003 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2026 | Multi-Grained Line Graph Neural Network With Hierarchical Contrastive Learning for Predicting Drug-Disease AssociationsabstractPredicting drug-disease associations is a crucial step in drug repositioning, especially with computational methods that quickly locate potential drug-disease pairs. Heterogenous network is a common tool for introducing multiple type relation information about drugs and diseases. However, the diversity of relations is ignored in most of existing methods, which makes them difficult to explore type semantic information with structure properties. Therefore, we propose a relation-centric GNN framework to encode critical association patterns. Firstly, we utilize a relation-centric graph, line graph, to represent the context of a drug-disease pair identified as the center node. The prediction problem is modeled to learn the embedding vector of the center node. Secondly, a multi-grained line graph neural network (MGLGNN) is designed to excavate fine-grained features that encapsulate local graph structures. We theoretically define a handful of typical nodes that can be regarded as high-order abstractions of relations in each type. Then, MGLGNN distills the local information and passes it to typical nodes from a global perspective. With learned multi-grained features, the center node automatically captures heterogenous relation semantics and structure patterns. Thirdly, a hierarchical contrastive learning (HCL) mechanism is proposed to ensure the quality of multi-grained features in an unsupervised way. Extensive experiments show the great potential of our model in mining drug-disease associations. Bao-Min Liu, Ling-Yun Dai, Junliang Shang, Chun-Hou Zheng 0001, Ying-Lian Gao, Rui Gao 0006, Jin-Xing Liu 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2026 | A Hierarchical Attention-Based Negative Sampling Method for Drug Repositioning Using Neighborhood Interaction FusionabstractAccurate prediction of drug-disease associations (DDAs) is essential for drug repositioning and the development of novel therapeutic strategies. However, existing methods often suffer from limited prior knowledge and the use of oversimplified negative sampling techniques, which hinder their ability to capture the complex relationships between drugs and diseases. To break through these limitations, we propose a new model, Hierarchical Attention Mechanism-Based Negative Sampling (HA-NegS), which aims to enhance the prediction of potential DDAs. In this study, HA-NegS further computes the similarity information between drugs and diseases and constructs heterogeneous and homogeneous networks based on it. For the similarity network, HA-NegS fuses Graph Convolutional Network (GCN) and Graph Attention Network (GAT) to effectively capture the neighborhood features of the target nodes. Subsequently, the model incorporates a hierarchical sampling strategy using the PageRank algorithm to rank nodes in descending order of global importance. The attention mechanism is then used to calculate the attention score and re-rank the nodes accordingly. This approach ensures the reliability of the negative sample selection. In order to obtain optimized representations, we use graph contrastive learning methods to refine drug and disease features with homogeneous and heterogeneous neighborhood information. Experimental results on a benchmark dataset show that HA-NegS outperforms existing baseline methods in predicting DDA. In addition, case studies for Alzheimer's disease and Parkinson's disease highlight the effectiveness of HA-NegS in discovering new therapeutic applications for existing drugs. Cheng-Long Mi, Ling-Yun Dai, Junliang Shang, Juan Wang 0003, Feng Li 0033 |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | MGAMDA: Multi Source Similarity Fusion-Based Graph Convolutional Neural Network and Attention Mechanism Network for Predicting MiRNA-Disease AssociationsabstractA mounting body of research indicates that dysregulation of MicroRNAs (miRNAs) causes disease through a variety of underlying mechanisms. Predicting microRNA (miRNA)-disease associations (MDAs) is essential for disease prognosis and therapeutics. Compared to conventional biological experiments, computational models save time and effort. A new method is proposed inspired by the graph convolutional networks. It has been named Multi source similarity fusion-based graph convolutional neural network and attention mechanism network for predicting miRNA-disease associations (MGAMDA). First, the several similarity networks between miRNAs and diseases were built. Then, multi-source information network is fused. And the feature was aggregated by using GCNs. In order to address the different levels of importance of the information, an attention mechanism was used to assign weights. The similar features of the disease side and miRNA side were finally obtained separately. It is combined with the association features that are obtained from the association information, and then it is fed into the multi-layer perceptron (MLP). To obtain prediction scores for unknown associations between miRNAs and diseases, a multilayer perceptron was utilized. To validate the new methodology's effectiveness, we performed a series of experimental studies using the Human MicroRNA Disease Database (HMDD v3.2). The performance of the$\mathbf{5}$-fold cross-validation on the datasets shows that MGAMDA surpasses other methods in the area of AUC, AUPR, ACC, F1-score, Recall, and Precision. Furthermore, case studies have demonstrated that MGAMDA accurately predicts miRNAs associated with colon, breast, and stomach cancer. Ling-Yun Dai, Cheng-Long Mi, Juan Wang 0003, Feng Li 0033 |
BIBM | 1 |
| 2025 | Adaptive Weighting Contrastive Learning for Spatial Domain Identification in Spatial TranscriptomicsabstractThe rapid advancement of spatial transcriptomics has enabled the joint analysis of gene expression and spatial location data. This integration opens new avenues for uncovering tissue heterogeneity. Existing methods attempt to combine spatial and expression information to identify spatial domains. However, they often treat all information sources equally and do not account for their varying impact on results. To address this challenge, we propose AWCST, an adaptive weighted contrastive learning framework for spatial domain identification. AWCST first extracts spatial and expression latent representations and then fuses them using a multi-head attention mechanism. It measures the distributional differences between each view and the fused feature using Maximum Mean Discrepancy. These differences are converted into adaptive weights for the contrastive loss, enhancing the influence of high-quality information sources. Finally, we evaluate AWCST on two independent datasets to demonstrate its effectiveness. Xiyue Li, Ling-Yun Dai, Junliang Shang, Feng Li 0033 |
BIBM | 2 |
| 2025 | scGZDC: Graph-Based ZINB Deep Clustering for Single-Cell RNA-Seq DataabstractSingle-cell RNA sequencing (scRNA-seq) is a key technology for studying cellular heterogeneity. However, the high levels of sparsity and noise in scRNA-seq data present challenges for accurate cell clustering. To address this, we propose Graph-based ZINB Deep Clustering for Single-cell RNA-seq Data (scGZDC), a novel framework that operates within a variational autoencoder (VAE). The encoder of scGZDC employs a Graph Convolutional Network (GCN) to learn low-dimensional representations by leveraging both a preprocessed gene expression matrix and the cell-cell similarity graph. The decoder, in turn, employs a Graph Attention Network (GAT) to reconstruct gene expression counts via a Zero-Inflated Negative Binomial (ZINB) distribution, a distribution particularly well-suited for scRNA-seq data. To achieve end-to-end optimization, a Deep Embedding for Clustering (DEC) objective is integrated into the framework. Extensive experiments on public datasets demonstrate that scGZDC consistently outperforms existing methods. Our results show that unifying graph structural information with a suitable probabilistic model in an end-to-end clustering framework is an effective strategy for improving single-cell analysis. Hui-Bo Tian, Jin-Xing Liu 0001, Junliang Shang, Juan Wang 0003, Ling-Yun Dai |
BIBM | 6 |
| 2025 | GAEKLRR: A novel clustering method of the low-rank representation based on graph auto-encoder and relaxed k-means for single-cell type identificationabstractClustering is critical for scRNA-seq because it reveals the similarity of single-cell expression patterns. However, single-cell data contains numerous noises and outliers. Traditional clustering algorithms may fail to capture accurate clustering information. In this study, we propose the GAEKLRR method for single-cell type identification, which is a low-rank representation (LRR) method based on a graph autoencoder (GAE) and relaxed k-means. GAEKLRR consists of gedLRR and relaxed k-means. Among them, gedLRR is a GAEbased LRR algorithm that captures structural information and node features of samples using GAE. Relaxed k-means is a soft clustering method that can better preserve complex relationships between samples through soft partitioning. Specifically, to reduce the impact of noise and outliers on the mapping benchmark, GAEKLRR generates a robust graph embedding dictionary using gedLRR. Due to the reconstruction using inner product distance, the graph embedding dictionary has interpretability. Meanwhile, to capture accurate clustering information, GAEKLRR utilizes gedLRR to seek the LRR matrix of the graph embedding dictionary while using relaxed k -means to update the clustering centroid. It is worth noting that the continuous clustering indication matrix captured by relaxed k-means contains clustering labels, which can be used directly for clustering tasks. Finally, experiments on real singlecell datasets demonstrate that GAEKLRR has significant advantages for clustering. Linping Wang, Junliang Shang, Ling-Yun Dai, Juan Wang 0003 |
BIBM | 3 |
| 2025 | Spatial Multi-Omics Integration Via Information-Aware Multi-View Contrastive LearningabstractThe rapid advancement of spatial multi-omics technology enables the simultaneous acquisition of diverse expression data from the same tissue or slice. Different omics offer unique and critical information about the biological system. However, most existing methods are unable to fully utilize this information for downstream tasks such as spatial domain identification. To integrate this information effectively for downstream analysis, we introduce a novel Spatial Multi-omics data integration method based on Information-Aware Multi-view Contrastive Learning (SM-IAMCL). It optimizes the spatial and feature neighborhood graphs for each omics by the specific graph learner and fused graph learner, and learns the fused graph of spatial and feature neighborhood graphs at the same time. Then, to make fused graph of each omics integrate both shared and unique information of spatial and feature neighborhood graphs, we incorporate graph-level contrastive learning between different views in each omics. Finally, the learned fused representation of each omics is then integrated via a weighted fusion strategy to generate an integrated low-dimensional latent representation of spatial multiomics. This integrated representation is used for a variety of downstream analysis tasks. The experimental results show that SM-IAMCL outperforms other seven existing methods in the downstream tasks such as spatial domain identification. Conghui Zhang, Ling-Yun Dai, Juan Wang 0003, Junliang Shang, Feng Li 0033 |
BIBM | 2 |
| 2025 | Predicting Potential Associations Between Microbes and Diseases Using Graph Attention Auto-encoder and PU Learning
Ling-Yun Dai, Feng Li 0033 |
ICIC (27) | 2 |
| 2025 | MPSO-CD: A Multi-Objective Particle Swarm Optimization Community Detection Method for Identifying Disease ModulesabstractThe dysfunction of biological systems caused by disease-related genes is one of the inducements of complex diseases. To understand molecular mechanisms of complex diseases, the identification of disease-related gene modules in biological networks through community detection is emerging as a promising approach. However, most community detection methods are not suitable for biological networks because their topological structures are complex and the scale of biologically relevant modules are small. In this paper, a novel community detection method called MPSO-CD was proposed based on multi-objective particle swarm optimization, in which negative ratio association and ratio cut were employed as objective functions. Highlights of MPSO-CD are a mutation strategy based on clustering coefficient and the procedure of disease module screening referring to the internal connection density and functional similarity. Experimental results of social and synthetic complex networks indicate that MPSO-CD is comparable and often superior to four compared methods. Eventually, MPSO-CD is applied to the asthma gene co-expression network for identifying potential disease modules that provide the molecular mechanism information about asthma. Most of the captured modules have been proven to be associated with asthma through Gene Ontology and pathway enrichment analysis. Xuhui Zhu, Mingyuan Bi, Junliang Shang, Feng Li 0033, Yuanyuan Zhang 0008, Ling-Yun Dai, Shengjun Li, Jin-Xing Liu 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 7 |
| 2024 | MNGCCL: Multi-neighborhood graph collaborative contrastive learning for drug-disease association predictionabstractExploring new therapeutic applications for existing drugs can effectively reduce drug development costs. However, current drug-disease association (DDA) prediction methods often fail to effectively integrate multi-domain information. The lack of multi-domain information integration causes these methods to heavily rely on prior knowledge, thereby limiting their generalization ability. To address this issue, we developed a Multi-Domain Graph Collaborative Contrastive Learning (MNGCCL) model for DDA prediction. In the MNGCCL framework, a feature extraction module is designed to effectively extract both single-domain and multi-domain features. The single-domain and multi-domain feature extraction components in this module run in parallel, extracting key features of drugs and diseases from different latent spaces (e.g., homogeneous and heterogeneous networks). MNGCCL employs graph collaborative contrastive learning to integrate these features and enhances information interaction by designing new node scoring for negative sample sampling. This significantly enriches the semantic features of drugs and diseases. In DDA prediction, MNGCCL outperforms other state-of-the-art models across various datasets and partitioning methods. Notably, MNGCCL excels in drug repositioning for Parkinson’s disease and Alzheimer’s disease, as well as handling data sparsity. These findings highlight its tremendous potential for drug repositioning and DDA prediction, especially in the context of sparse omics data. Cheng-Long Mi, Jin-Xing Liu 0001, Junliang Shang, Juan Wang 0003, Ling-Yun Dai |
BIBM | 7 |
| 2024 | Seizure Types Classification Based on Multi-branch Hybrid Deep Learning Network
Qingwei Jia, Jin-Xing Liu 0001, Junling Shang, Ling-Yun Dai, Wenrong Hu, Shasha Yuan |
ICIC (4) | 4 |
| 2024 | IMPRINTS.CETSA and IMPRINTS.CETSA.app: an R package and a Shiny application for the analysis and interpretation of IMPRINTS-CETSA dataabstractIMPRINTS-CETSA (Integrated Modulation of Protein Interaction States-Cellular Thermal Shift Assay) provides a highly resolved means to systematically study the interactions of proteins with other cellular components, including metabolites, nucleic acids and other proteins, at the proteome level, but no freely available and user-friendly data analysis software has been reported. Here, we report IMPRINTS.CETSA, an R package that provides the basic data processing framework for robust analysis of the IMPRINTS-CETSA data format, from preprocessing and normalization to visualization. We also report an accompanying R package, IMPRINTS.CETSA.app, which offers a user-friendly Shiny interface for analysis and interpretation of IMPRINTS-CETSA results, with seamless features such as functional enrichment and mapping to other databases at a single site. For the hit generation part, the diverse behaviors of protein modulations have been typically segregated with a two-measure scoring method, i.e. the abundance and thermal stability changes. We present a new algorithm to classify modulated proteins in IMPRINTS-CETSA experiments by a robust single-measure scoring. In this way, both the numerical changes and the statistical significances of the IMPRINTS information can be visualized on a single plot. The IMPRINTS.CETSA and IMPRINTS.CETSA.app R packages are freely available on GitHub at https://github.com/nkdailingyun/IMPRINTS.CETSA and https://github.com/mgerault/IMPRINTS.CETSA.app, respectively. IMPRINTS.CETSA.app is also available as an executable program at https://zenodo.org/records/10636134. Marc-Antoine Gerault, Samuel Granjeaud, Luc Camoin, Pär Nordlund, Ling-Yun Dai |
Briefings Bioinform. | 5 |
| 2023 | MLQP: A Machine Learning Based Quadratic Prediction Method for MiRNA-Disease AssociationsabstractAs science and technology continue to advance, more and more studies show that there are so many miRNA-disease associations (MDAs). Known MDAs will help us prevent and treat certain diseases. However, traditional wet experiments greatly consume manpower and time. Therefore, it is critical that a few reliable methods for predicting MDAs are developed. In this study, we have come up with a machine-learning method for quadratic prediction MDAs (MLQP). MLQP is an improvement method on the traditional least squares method. In this method, MDAs information can be used more by using Weight K Nearest Known Neighbors (WKNKN) method and neighborhood similarity processing (NSP). Gaussian interaction profile (GIP) kernel similarity is applied to improve the accuracy of prediction results. The regularization least squares method is also utilized to generate the prediction scores. AUC value of MLQP under five-fold cross-validation is 0.957. MLQP outperforms other methods in predicting MDAs, as demonstrated by the final experimental results. Finally, the efficacy and practicality of the MLQP will be further verified in the study of four specific diseases. Ling-Yun Dai, Jin-Xing Liu 0001 |
BIBM | 2 |
| 2023 | MKGSAGE: A Computational Framework via Multiple Kernel Fusion on GraphSAGE for Inferring Potential Disease-Related MicrobesabstractMicrobes play a crucial role within the human body and are closely associated with the occurrence and development of numerous diseases. Studies have shown that disruptions in the composition and functionality of microbes can lead to immune system imbalances, inflammatory responses, and subsequently impact human health. Therefore, developing computational models to discover the potential connections between microbes and diseases is currently a hot topic. In this paper, a computational framework based on the multiple kernel fusion of graph embedding with sampling and aggregation (GraphSAGE) and dual Laplace regularized least squares called MKGSAGE is proposed for predicting potential links between microbe and disease. First, multiple layers embedding features of microbe and disease are learned from the initial input features by GraphSAGE. The kernel matrices are then calculated separately for each layer based on the Gaussian interaction profile (GIP). Furthermore, the multiple kernel fusion method is proposed for fusing kernel matrices of each layer and the initial similarity matrix. Dual Laplacian regularized least squares are finally applied for potential microbe-disease association prediction. Compared with six state-of-the-art methods on the HMDAD dataset, 5-fold cross-validations show that MKGSAGE performs best. In addition, case studies on asthma and inflammatory bowel disease further validate the effectiveness of MKGSAGE on discovering novel microbe-disease associations. Jin-Xing Liu 0001, Bao-Min Liu, Ling-Yun Dai, Feng Li 0033, Ying-Lian Gao |
BIBM | 4 |
| 2023 | Epileptic Seizure Detection Based on Feature Extraction and CNN-BiGRU Network with Attention Mechanism
Jie Xu 0059, Juan Wang 0003, Jin-Xing Liu 0001, Junliang Shang, Ling-Yun Dai, Kuiting Yan, Shasha Yuan |
ICIC (2) | 5 |
| 2023 | scGASI: A Graph Autoencoder-Based Single-Cell Integration Clustering Method
Tian-Jing Qiao, Feng Li 0033, Shasha Yuan, Ling-Yun Dai, Juan Wang 0003 |
ISBRA | 4 |
| 2022 | KSMDB: A classification method in imbalanced COVID dataset based on KmeansSMOTE and DeBERTabstract2022 is already the third year of the COVID-19 outbreak, and public opinion information about the outbreak has always been at the forefront of hot searches. The imbalance problem prevalent in many reviews of COVID-19 causes classification models to favor most categories in training and prediction process, resulting in low accuracy of small sample classification data generated by imbalanced data sets. Therefore, it is suggested here that the text classification model is based on the combination of the KMeansSMOTE method combined with DeBERT. First of all, during data processing, the KmeansSMOTE algorithm is utilized to oversample the imbalance of the COVID dataset, which increases the classification accuracy of the model. Besides, we put a stacked denoising bidirectional transformer encoder (DeBERT) to use, a more abstract and richer hidden feature vector is extracted by adding an embedded layer after the input tag, and the noise data is reconstructed to solve the noise problem in the process of raw data existence and oversampling. Furthermore, on the basis of model training, overfitting can be alleviated by adopting an early stopping strategy. A world of experiments using the COVID dataset demonstrates the effectiveness of the proposed method for solving simple imbalance and noise problems. With an overall accuracy of 87%, which improves the classification effect of minority samples and provides a new feasible method for the war of epidemic prevention. Hua-Hui Gao, Junliang Shang, Ling-Yun Dai |
BIBM | 4 |
| 2022 | ARGLRR: An Adjusted Random Walk Graph Regularization Sparse Low-Rank Representation Method for Single-Cell RNA-Sequencing Data Clustering
Zhen-Chang Wang, Jin-Xing Liu 0001, Junliang Shang, Ling-Yun Dai, Chun-Hou Zheng 0001, Juan Wang 0003 |
ISBRA | 4 |
| 2022 | Multi-view manifold regularized compact low-rank representation for cancer samples clustering on multi-omics dataabstractBACKGROUND: The identification of cancer types is of great significance for early diagnosis and clinical treatment of cancer. Clustering cancer samples is an important means to identify cancer types, which has been paid much attention in the field of bioinformatics. The purpose of cancer clustering is to find expression patterns of different cancer types, so that the samples with similar expression patterns can be gathered into the same type. In order to improve the accuracy and reliability of cancer clustering, many clustering methods begin to focus on the integration analysis of cancer multi-omics data. Obviously, the methods based on multi-omics data have more advantages than those using single omics data. However, the high heterogeneity and noise of cancer multi-omics data pose a great challenge to the multi-omics analysis method. RESULTS: In this study, in order to extract more complementary information from cancer multi-omics data for cancer clustering, we propose a low-rank subspace clustering method called multi-view manifold regularized compact low-rank representation (MmCLRR). In MmCLRR, each omics data are regarded as a view, and it learns a consistent subspace representation by imposing a consistence constraint on the low-rank affinity matrix of each view to balance the agreement between different views. Moreover, the manifold regularization and concept factorization are introduced into our method. Relying on the concept factorization, the dictionary can be updated in the learning, which greatly improves the subspace learning ability of low-rank representation. We adopt linearized alternating direction method with adaptive penalty to solve the optimization problem of MmCLRR method. CONCLUSIONS: Finally, we apply MmCLRR into the clustering of cancer samples based on multi-omics data, and the clustering results show that our method outperforms the existing multi-view methods. Juan Wang 0003, Cong-Hai Lu, Ling-Yun Dai, Shasha Yuan |
BMC Bioinform. | 4 |
| 2021 | Joint CC and Bimax: A Biclustering Method for Single-Cell RNA-Seq Data Analysis
He-Ming Chu, Jin-Xing Liu 0001, Juan Wang 0003, Shasha Yuan, Ling-Yun Dai |
ISBRA | 6 |
| 2021 | IPCARF: improving lncRNA-disease association prediction using incremental principal component analysis feature selection and a random forest classifierabstractBACKGROUND: Identifying lncRNA-disease associations not only helps to better comprehend the underlying mechanisms of various human diseases at the lncRNA level but also speeds up the identification of potential biomarkers for disease diagnoses, treatments, prognoses, and drug response predictions. However, as the amount of archived biological data continues to grow, it has become increasingly difficult to detect potential human lncRNA-disease associations from these enormous biological datasets using traditional biological experimental methods. Consequently, developing new and effective computational methods to predict potential human lncRNA diseases is essential. RESULTS: Using a combination of incremental principal component analysis (IPCA) and random forest (RF) algorithms and by integrating multiple similarity matrices, we propose a new algorithm (IPCARF) based on integrated machine learning technology for predicting lncRNA-disease associations. First, we used two different models to compute a semantic similarity matrix of diseases from a directed acyclic graph of diseases. Second, a characteristic vector for each lncRNA-disease pair is obtained by integrating disease similarity, lncRNA similarity, and Gaussian nuclear similarity. Then, the best feature subspace is obtained by applying IPCA to decrease the dimension of the original feature set. Finally, we train an RF model to predict potential lncRNA-disease associations. The experimental results show that the IPCARF algorithm effectively improves the AUC metric when predicting potential lncRNA-disease associations. Before the parameter optimization procedure, the AUC value predicted by the IPCARF algorithm under 10-fold cross-validation reached 0.8529; after selecting the optimal parameters using the grid search algorithm, the predicted AUC of the IPCARF algorithm reached 0.8611. CONCLUSIONS: We compared IPCARF with the existing LRLSLDA, LRLSLDA-LNCSIM, TPGLDA, NPCMF, and ncPred prediction methods, which have shown excellent performance in predicting lncRNA-disease associations. The compared results of 10-fold cross-validation procedures show that the predictions of the IPCARF method are better than those of the other compared methods. Jin-Xing Liu 0001, Ling-Yun Dai |
BMC Bioinform. | 4 |
| 2021 | The Automatic Detection of Seizure Based on Tensor Distance And Bayesian Linear Discriminant AnalysisabstractElectroencephalogram (EEG) plays an important role in recording brain activity to diagnose epilepsy. However, it is not only laborious, but also not very cost effective for medical experts to manually identify the features on EEG. Therefore, automatic seizure detection in accordance with the EEG recordings is significant for the diagnosis and treatment of epilepsy. Here, a new method for detecting seizures using tensor distance (TD) is proposed. First, the time-frequency characteristics of EEG signals are obtained by wavelet transformation, and the tensor representation of EEG signals is then obtained. Tucker decomposition is used to obtain the principal components of the EEG tensor. After, the distances between different categories of EEG tensors are calculated as the EEG features. Finally, the TD features are classified through the Bayesian Linear Discriminant Analysis (Bayesian LDA) classifier. The performance of this method is measured by the sensitivity, specificity, and recognition accuracy. Results indicate 95.12% sensitivity, 97.60% specificity, 97.60% recognition accuracy, and a false detection rate of 0.76 per hour in the invasive EEG dataset, which included 566.57[Formula: see text]h of EEG recording data from 21 patients. Taken together, the results show that TD has a good detection effect for seizure classification and that this method has high computational speed and great potential for real-time diagnosis. Delu Ma, Shasha Yuan, Junliang Shang, Jin-Xing Liu 0001, Ling-Yun Dai, Fangzhou Xu |
Int. J. Neural Syst. | 5 |
| 2021 | Logistic Weighted Profile-Based Bi-Random Walk for Exploring MiRNA-Disease Associations
Ling-Yun Dai, Jin-Xing Liu 0001, Juan Wang 0003, Shasha Yuan |
J. Comput. Sci. Technol. | 1 |
| 2020 | Dual Graph regularized PCA based on Different Norm Constraints for Bi-clustering Analysis on Single-cell RNA-seq DataabstractIn recent years, single-cell RNA sequencing (scRNA-seq) technology has made significant progress in many fields and become an important means to study cell dynamics. How to effectively mine valuable biological information from these sequencing data is a topic worthy of researching. In this paper, two new methods based on traditional principal component analysis (PCA) are proposed and used to scRNA-seq data. The first method named dual graph regularized PCA (DGPPCA) is based on Frobenius-norm and L2,p-norm constraints, and the method named the dual graph-regularization PCA (DG2PPCA) is based on the nonconvex proximal Lp-norm ( 02,p-norm constraints. We apply these two new methods to five scRNA-seq datasets, and perform bi-clustering on genes and samples at the same time. Extensive experiments are conducted to explore the influence of the combination of different norm constraints in the two optimization models. Jin-Xing Liu 0001, Juan Wang 0003, Shasha Yuan, Ling-Yun Dai |
BIBM | 6 |
| 2020 | Automatic Seizure Prediction based on Modified Stockwell Transform and Tensor DecompositionabstractReliable epileptic seizure prediction is significantly important in improving the life of patients and enhancing the therapy effect. In this paper, a novel seizure prediction algorithm is proposed employing the tensor decomposition on long-term intracranial EEG recordings. The modified Stockwell transform (MST) is conducted on the segmented EEG signals to transform into two-dimensional instantaneous power spectra. Then, the third-order tensor representation of the multi-channel EEG signals are structured with the models of time, frequency and space. Tucker decomposition, one valid tensor decomposition method, is applied to obtain the principal components of the EEG tensors and the smaller core tensors after decomposition are extracted as features of interictal EEG and preictal EEG. After that, the classification of preictal and interictal data is achieved by feeding the features into Bayesian Linear Discriminant Analysis (BLDA) classifier. The evaluation of the proposed algorithm is carried out on the Freiburg EEG database and a sensitivity of 88.49% for the seizure occurrence period of 30 min, meanwhile, a sensitivity of 97.62% for the seizure occurrence period of 50 min are yielded with a false alarm rate of 0. 25/h. The results show that this algorithm based on tensor analysis has notable performance for seizure prediction. Shasha Yuan, Jin-Xing Liu 0001, Junliang Shang, Fangzhou Xu, Ling-Yun Dai |
BIBM | 5 |
| 2019 | L2, 1-GRMF: an improved graph regularized matrix factorization method to predict drug-target interactionsabstractBACKGROUND: Predicting drug-target interactions is time-consuming and expensive. It is important to present the accuracy of the calculation method. There are many algorithms to predict global interactions, some of which use drug-target networks for prediction (ie, a bipartite graph of bound drug pairs and targets known to interact). Although these algorithms can predict some drug-target interactions to some extent, there is little effect for some new drugs or targets that have no known interaction. RESULTS: Since the datasets are usually located at or near low-dimensional nonlinear manifolds, we propose an improved GRMF (graph regularized matrix factorization) method to learn these flow patterns in combination with the previous matrix-decomposition method. In addition, we use one of the pre-processing steps previously proposed to improve the accuracy of the prediction. CONCLUSIONS: Cross-validation is used to evaluate our method, and simulation experiments are used to predict new interactions. In most cases, our method is superior to other methods. Finally, some examples of new drugs and new targets are predicted by performing simulation experiments. And the improved GRMF method can better predict the remaining drug-target interactions. Ying-Lian Gao, Jin-Xing Liu 0001, Ling-Yun Dai, Shasha Yuan |
BMC Bioinform. | 4 |
| 2019 | The computational prediction of drug-disease interactions using the dual-network L2,1-CMF methodabstractBACKGROUND: Predicting drug-disease interactions (DDIs) is time-consuming and expensive. Improving the accuracy of prediction results is necessary, and it is crucial to develop a novel computing technology to predict new DDIs. The existing methods mostly use the construction of heterogeneous networks to predict new DDIs. However, the number of known interacting drug-disease pairs is small, so there will be many errors in this heterogeneous network that will interfere with the final results. RESULTS: -norm are introduced in our method to achieve better results than other advanced methods. The network similarities of drugs and diseases with their chemical and semantic similarities are combined in this method. CONCLUSIONS: Cross validation is used to evaluate our method, and simulation experiments are used to predict new interactions using two different datasets. Finally, our prediction accuracy is better than other existing methods. This proves that our method is feasible and effective. Ying-Lian Gao, Jin-Xing Liu 0001, Juan Wang 0003, Junliang Shang, Ling-Yun Dai |
BMC Bioinform. | 6 |
| 2019 | PCA via joint graph Laplacian and sparse constraint: Identification of differentially expressed genes and sample clustering on gene expression dataabstractBACKGROUND: In recent years, identification of differentially expressed genes and sample clustering have become hot topics in bioinformatics. Principal Component Analysis (PCA) is a widely used method in gene expression data. However, it has two limitations: first, the geometric structure hidden in data, e.g., pair-wise distance between data points, have not been explored. This information can facilitate sample clustering; second, the Principal Components (PCs) determined by PCA are dense, leading to hard interpretation. However, only a few of genes are related to the cancer. It is of great significance for the early diagnosis and treatment of cancer to identify a handful of the differentially expressed genes and find new cancer biomarkers. RESULTS: In this study, a new method gLSPCA is proposed to integrate both graph Laplacian and sparse constraint into PCA. gLSPCA on the one hand improves the clustering accuracy by exploring the internal geometric structure of the data, on the other hand identifies differentially expressed genes by imposing a sparsity constraint on the PCs. CONCLUSIONS: Experiments of gLSPCA and its comparison with existing methods, including Z-SPCA, GPower, PathSPCA, SPCArt, gLPCA, are performed on real datasets of both pancreatic cancer (PAAD) and head & neck squamous carcinoma (HNSC). The results demonstrate that gLSPCA is effective in identifying differentially expressed genes and sample clustering. In addition, the applications of gLSPCA on these datasets provide several new clues for the exploration of causative factors of PAAD and HNSC. Chun-Mei Feng 0001, Yong Xu 0001, Mi-Xiao Hou, Ling-Yun Dai, Junliang Shang |
BMC Bioinform. | 4 |
| 2019 | Multi-cancer samples clustering via graph regularized low-rank representation method under sparse and symmetric constraintsabstractBACKGROUND: Identifying different types of cancer based on gene expression data has become hotspot in bioinformatics research. Clustering cancer gene expression data from multiple cancers to their own class is a significance solution. However, the characteristics of high-dimensional and small samples of gene expression data and the noise of the data make data mining and research difficult. Although there are many effective and feasible methods to deal with this problem, the possibility remains that these methods are flawed. RESULTS: In this paper, we propose the graph regularized low-rank representation under symmetric and sparse constraints (sgLRR) method in which we introduce graph regularization based on manifold learning and symmetric sparse constraints into the traditional low-rank representation (LRR). For the sgLRR method, by means of symmetric constraint and sparse constraint, the effect of raw data noise on low-rank representation is alleviated. Further, sgLRR method preserves the important intrinsic local geometrical structures of the raw data by introducing graph regularization. We apply this method to cluster multi-cancer samples based on gene expression data, which improves the clustering quality. First, the gene expression data are decomposed by sgLRR method. And, a lowest rank representation matrix is obtained, which is symmetric and sparse. Then, an affinity matrix is constructed to perform the multi-cancer sample clustering by using a spectral clustering algorithm, i.e., normalized cuts (Ncuts). Finally, the multi-cancer samples clustering is completed. CONCLUSIONS: A series of comparative experiments demonstrate that the sgLRR method based on low rank representation has a great advantage and remarkable performance in the clustering of multi-cancer samples. Juan Wang 0003, Cong-Hai Lu, Jin-Xing Liu 0001, Ling-Yun Dai |
BMC Bioinform. | 4 |
| 2019 | ACCBN: ant-Colony-clustering-based bipartite network method for predicting long non-coding RNA-protein interactionsabstractBACKGROUND: Long non-coding RNA (lncRNA) studies play an important role in the development, invasion, and metastasis of the tumor. The analysis and screening of the differential expression of lncRNAs in cancer and corresponding paracancerous tissues provides new clues for finding new cancer diagnostic indicators and improving the treatment. Predicting lncRNA-protein interactions is very important in the analysis of lncRNAs. This article proposes an Ant-Colony-Clustering-Based Bipartite Network (ACCBN) method and predicts lncRNA-protein interactions. The ACCBN method combines ant colony clustering and bipartite network inference to predict lncRNA-protein interactions. RESULTS: A five-fold cross-validation method was used in the experimental test. The results show that the values of the evaluation indicators of ACCBN on the test set are significantly better after comparing the predictive ability of ACCBN with RWR, ProCF, LPIHN, and LPBNI method. CONCLUSIONS: With the continuous development of biology, besides the research on the cellular process, the research on the interaction function between proteins becomes a new key topic of biology. The studies on protein-protein interactions had important implications for bioinformatics, clinical medicine, and pharmacology. However, there are many kinds of proteins, and their functions of interactions are complicated. Moreover, the experimental methods require time to be confirmed because it is difficult to estimate. Therefore, a viable solution is to predict protein-protein interactions efficiently with computers. The ACCBN method has a good effect on the prediction of protein-protein interactions in terms of sensitivity, precision, accuracy, and F1-score. Guangshun Li, Jin-Xing Liu 0001, Ling-Yun Dai, Ying Guo 0002 |
BMC Bioinform. | 4 |
| 2018 | Sparse Orthogonal Nonnegative Matrix Factorization for Identifying Differentially Expressed Genes and Clustering Tumor Samples
Ling-Yun Dai, Jin-Xing Liu 0001, Mi-Xiao Hou, Shasha Yuan |
BIBM | 1 |
| 2018 | A Fast Quantum Clustering Approach for Cancer Gene Clustering
Guangshun Li, Jin-Xing Liu 0001, Ling-Yun Dai, Shasha Yuan, Ying Guo 0002 |
BIBM | 4 |
| 2018 | Performance Analysis of Non-negative Matrix Factorization Methods on TCGA Data
Mi-Xiao Hou, Jin-Xing Liu 0001, Junliang Shang, Ying-Lian Gao, Ling-Yun Dai |
ICIC (2) | 6 |
| 2017 | Robust graph regularized sparse orthogonal nonnegative matrix factorization for identifying differentially expressed genesabstractWith the advent of sequencing technology, numerous gene expression data are generated. Identifying differentially expressed genes play an important role in the gene therapy of cancer patients. As an useful mathematical tool, nonnegative matrix factorization (NMF) has been successfully used for identifying differentially expressed genes. In this paper, a novel method named robust graph regularized sparse orthogonal nonnegative matrix factorization (RGSON) is proposed and used for identifying differentially expressed genes, which introduces manifold learning, L1and orthogonal constraints into the objective function. In particular, L2,1-norm minimization is enforced on the objective function to improve the robustness of the algorithm. To prove the validity of the algorithm, experiments on the real genomic dataset are conducted. The results show that RGSON performs more effective than many other methods for identifying differentially expressed genes. Ling-Yun Dai, Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Junliang Shang, Chun-Mei Feng 0001, Yaxuan Wang |
BIBM | 1 |
| 2017 | Low-rank representation regularized by L2, 1-norm for identifying differentially expressed genesabstractLow-rank representation (LRR) via rank minimization is a high efficiency method for capturing low-dimensional structure embedded in high-dimensional data. However, minimizing the rank of a matrix is NP-hard. In this paper, robust truncated nuclear norm low-rank representation regularized by L2,1-norm method (RTLRR) is proposed. The truncated nuclear norm is introduced to replace the nuclear norm to approximate the rank function. At the same time, L2,1-norm is used to regularize the sparse matrix to achieve better sparse effect of the algorithm. The proposed method is divided into two steps. Firstly, we do singular value decomposition (SVD) to the original data matrix. Then we apply the truncated nuclear norm and L2,1-norm constraints to subproblems and use inexact augmented Lagrange multiplier method to solve subproblems. Finally, the genes with high scores will be identified as differentially expressed genes according to the sparse matrix. The results on The Cancer Genome Atlas (TCGA) data illustrate that the effectiveness of RTLRR method outperforms many methods. Yaxuan Wang, Jin-Xing Liu 0001, Ying-Lian Gao, Chun-Hou Zheng 0001, Ling-Yun Dai |
BIBM | 5 |
| 2016 | Robust graph regularized discriminative nonnegative matrix factorization for characteristic gene selectionabstractRecent research shows that characteristic gene selection based on gene expression data remains faced with considerable challenges. This is primarily because vast amount of gene expression data have been generated with the development of gene detection technology. Nonetheless, the recognition rate and reliability of gene selection still need to be improved. In this paper, we propose a novel constrained method: robust graph regularized discriminative nonnegative matrix factorization (RGDNMF) for characteristic gene selection. The method mainly includes two aspects: firstly, we incorporate both intrinsic geometrical structure and discriminative label information into the NMF model. Secondly, we adopt L2,1 -norm minimization to both the error function and the regularization term which is robust to noises and outliers in gene data. Furthermore we present the multiplicative update rules and the convergence proof. Our experiments demonstrate that RGDNMF is far more effective than other existing methods. Ling-Yun Dai, Chun-Mei Feng 0001, Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Mi-Xiao Hou, Jiguo Yu |
BIBM | 1 |