EDBT 2026 Demo / reviewers in the wild / expert
Junliang Shang
dblp:123/6047
· DBLP profile ↗
116ranked-venue papers
11as first author
91since 2021 · last 2026
0000-0002-8488-2228ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 101 · 9 first-author · 77 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 11 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SSBiA: A Framework for Predicting Microbe-Disease Associations Based on Signed Subgraphs and Bi-Feature Aggregation
Ying-Lian Gao, Ming-Li Cui, Junliang Shang, Chun-Hou Zheng 0001 |
ICIC (27) | 5 |
| 2026 | GCNFormer: Trustworthy Multi-Omics Integration Method Using Global-Local Graph Transformer for Patient Classification
Yazhuo Han, Junliang Shang, Xiaoqi Tang, Baojuan Qin, Jin-Xing Liu 0001 |
ICIC (30) | 3 |
| 2026 | MRGBMDAT: a multi-relational graph encoder network with bilinear fusion for miRNA-disease association type predictionabstractMOTIVATION: MicroRNAs (miRNAs) are key post-transcriptional regulators involved in diverse biological processes, and their dysregulation is closely associated with the onset and progression of many diseases. Accurate prediction of miRNA-disease association types is therefore essential for understanding disease mechanisms and advancing precision medicine. Although computational methods provide efficient alternatives to wet-lab experiments, existing approaches often focus on binary association prediction, inadequately integrate local semantic dependencies and global topological structures, and suffer from class imbalance. RESULTS: To address these limitations, we propose MRGBMDAT, a multi-relational graph encoder network with bilinear fusion for miRNA-disease association type prediction. Specifically, a multi-relational graph convolution module with bidirectional cross-attention captures global topological structures, while a local subgraph sampling module extracts local semantic dependencies. A bilinear fusion decoder with element-wise attention jointly models their linear and nonlinear interactions. In addition, an iterative feature similarity-based negative sample selection strategy is introduced to alleviate class imbalance. Experimental results on the HMDD v3.2 dataset demonstrate that MRGBMDAT significantly outperforms five state-of-the-art methods across multiple evaluation metrics, exhibiting strong discriminative power and generalization capability. AVAILABILITY AND IMPLEMENTATION: The source code is publicly available at https://github.com/CDMBlab/MRGBMDAT. Siqi Zhu, Shijia Yan, Xuenan Shi, Junliang Shang |
Bioinform. | 6 |
| 2026 | DHGCMDA: a dual-view heterogeneous graph contrastive learning framework for miRNA-disease association type predictionabstractBACKGROUND: Accumulating evidence demonstrates that microRNA (miRNA) dysregulation drives the pathogenesis of diverse human diseases via intricate, context-dependent molecular mechanisms. Hence, prediction of miRNA-disease association types is a critical prerequisite for dissecting functional roles of miRNAs in disease initiation and progression. Although computational methods offer cost-effective, time-efficient alternatives to wet-lab experiments for miRNA-disease association type prediction, most of them are hampered by three key limitations: excessive reliance on association-derived similarity metrics gives rise to quantification bias, traditional pairwise graph architectures inadequately capture high-order biological interactions, and existing representation learning strategies fail to generate consistent embeddings across heterogeneous views and modalities. RESULTS: To address these issues, this study presents DHGCMDA, a dual-view heterogeneous graph contrastive learning framework for miRNA-disease association type prediction. Specifically, dual-view hypergraphs are first constructed based on heterogeneous similarity data to avoid excessive reliance on association-derived similarity metrics. A hypergraph convolutional network is then employed to capture high-order topological relationships between miRNAs and diseases, with its convolution cooperatively integrated with contrastive learning, intra-modality for cross-view consistency and cross-modality for embedding space alignment, to enhance feature representation quality. Finally, an attention-guided adaptive view fusion strategy dynamically weights and integrates distinct view representations, and type-aware message passing via heterogeneous graph Transformer simultaneously enables prediction of association presence and functional types. 5-fold cross-validation on HMDD v2.0 and v3.2 datasets demonstrates that DHGCMDA outperforms several state-of-the-art methods. Furthermore, case studies on breast neoplasms and hepatocellular carcinoma reveal that most predicted association types are corroborated by published literature, thereby validating the efficacy of DHGCMDA in miRNA-disease association type prediction. CONCLUSIONS: DHGCMDA exhibits robust discriminative power and generalization capability, providing a reliable computational alternative for miRNA-disease association type prediction. The source code is publicly available at https://github.com/CDMBlab/DHGCMDA . Fanyu Zhang, Shijia Yan, Xiaotong Kong, Hanxiang Wang, Junliang Shang |
BMC Bioinform. | 6 |
| 2026 | Single-cell distillation discriminative clustering based on asymmetric autoencoder
Junliang Shang, Aitian Fan, Baojuan Qin, Shoujia Jiang |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | Single-cell multi-view clustering based on dual contrastive learning and cross-attention fusion
Meng-Yao Hu, Jin-Xing Liu 0001, Junliang Shang, Ling-Yun Dai |
Expert Syst. Appl. | 4 |
| 2026 | A multi-objective multi-stage genetic algorithm for community detection in biological networks
Mingyuan Bi, Junliang Shang, Yahan Li, Feng Li 0033, Jin-Xing Liu 0001 |
Future Gener. Comput. Syst. | 2 |
| 2026 | Autoencoder-aided graph convolutional networks integrating multi-view and multi-scale for improving spatial domain identification
Juan Wang 0003, Xuena Liang, Shasha Yuan, Jin-Xing Liu 0001, Junliang Shang |
Knowl. Based Syst. | 5 |
| 2026 | A Novel Low-Dimensional Sparse and Low-Rank Representation Method for Single-Cell RNA Sequencing Data ClusteringabstractThe advancement of single-cell RNA sequencing (scRNA-seq) technology has enabled researchers to capture cellular heterogeneity at the individual cell level, driving progress in diverse fields such as developmental biology, immunology, and cancer research. Accurate cell clustering is a crucial step for researchers utilizing scRNA-seq data; however, inherent characteristics like high dimensionality and sparsity pose significant challenges to obtaining precise clustering results. To achieve accurate clustering, this paper proposes a novel approach that integrates dimensionality reduction, self-representation matrix construction, and the clustering process into an end-to-end model termed LDSLRR (Low-Dimensional Sparse and Low-Rank Representation). Specifically, the original gene expression matrix first undergoes dimensionality reduction via projection. Subsequently, low-rank representation combined with a sparsity constraint facilitates the learning of the self-representation matrix. Finally, the cluster assignment matrix is acquired using graph-regularized non-negative matrix factorization (NMF). These three modules are simultaneously optimized, enhancing the accuracy of the clustering results. Comparative experiments against multiple state-of-the-art clustering methods on various scRNA-seq datasets demonstrate the superiority of the proposed LDSLRR method. Zhenduo Zhang, Junliang Shang, Ling-Yun Dai, Juan Wang 0003 |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2026 | CCDM: Continuous-Time Conditional Diffusion Model for Blind CMRI Super-ResolutionabstractDiffusion probabilistic models have effectively addressed the ill-posed nature of cardiac magnetic resonance imaging (CMRI) super-resolution (SR) by learning high-resolution image distributions from low-resolution inputs. However, the iterative sampling process in these models often suffers from slow inference speeds, as well as limitations in the quality and structural consistency of the generated images. To address these challenges, we propose a continuous-time conditional diffusion model (CCDM) for blind CMRI SR. Specifically, we propose a continuous-time conditional diffusion module that reduces the time consumption of the diffusion probability model by maintaining the mean and variance of the data in the forward process. Meanwhile, we design a cascaded residual attention network as a feature extractor to enhance the model’s discriminative power and feature representation capabilities. To further elevate image fidelity, we propose an image quality loss module that integrates a score matching loss, significantly improving detail reconstruction and overall perceptual quality. Furthermore, we develop a hybrid score predictor that approximates the conditional score function via a hybrid parameterized denoising network, facilitating efficient CMRI generation through probability flow sampling. Extensive experimental results demonstrate that compared to existing diffusion model-based SR methods, our CCDM achieves significant improvements in SR quality while substantially reducing time consumption. Defu Qiu, Junliang Shang, Yuanke Zhang, Wenjun Zhang 0005, Kelvin K. L. Wong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Multi-Grained Line Graph Neural Network With Hierarchical Contrastive Learning for Predicting Drug-Disease AssociationsabstractPredicting drug-disease associations is a crucial step in drug repositioning, especially with computational methods that quickly locate potential drug-disease pairs. Heterogenous network is a common tool for introducing multiple type relation information about drugs and diseases. However, the diversity of relations is ignored in most of existing methods, which makes them difficult to explore type semantic information with structure properties. Therefore, we propose a relation-centric GNN framework to encode critical association patterns. Firstly, we utilize a relation-centric graph, line graph, to represent the context of a drug-disease pair identified as the center node. The prediction problem is modeled to learn the embedding vector of the center node. Secondly, a multi-grained line graph neural network (MGLGNN) is designed to excavate fine-grained features that encapsulate local graph structures. We theoretically define a handful of typical nodes that can be regarded as high-order abstractions of relations in each type. Then, MGLGNN distills the local information and passes it to typical nodes from a global perspective. With learned multi-grained features, the center node automatically captures heterogenous relation semantics and structure patterns. Thirdly, a hierarchical contrastive learning (HCL) mechanism is proposed to ensure the quality of multi-grained features in an unsupervised way. Extensive experiments show the great potential of our model in mining drug-disease associations. Bao-Min Liu, Ling-Yun Dai, Junliang Shang, Chun-Hou Zheng 0001, Ying-Lian Gao, Rui Gao 0006, Jin-Xing Liu 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2026 | A Hierarchical Attention-Based Negative Sampling Method for Drug Repositioning Using Neighborhood Interaction FusionabstractAccurate prediction of drug-disease associations (DDAs) is essential for drug repositioning and the development of novel therapeutic strategies. However, existing methods often suffer from limited prior knowledge and the use of oversimplified negative sampling techniques, which hinder their ability to capture the complex relationships between drugs and diseases. To break through these limitations, we propose a new model, Hierarchical Attention Mechanism-Based Negative Sampling (HA-NegS), which aims to enhance the prediction of potential DDAs. In this study, HA-NegS further computes the similarity information between drugs and diseases and constructs heterogeneous and homogeneous networks based on it. For the similarity network, HA-NegS fuses Graph Convolutional Network (GCN) and Graph Attention Network (GAT) to effectively capture the neighborhood features of the target nodes. Subsequently, the model incorporates a hierarchical sampling strategy using the PageRank algorithm to rank nodes in descending order of global importance. The attention mechanism is then used to calculate the attention score and re-rank the nodes accordingly. This approach ensures the reliability of the negative sample selection. In order to obtain optimized representations, we use graph contrastive learning methods to refine drug and disease features with homogeneous and heterogeneous neighborhood information. Experimental results on a benchmark dataset show that HA-NegS outperforms existing baseline methods in predicting DDA. In addition, case studies for Alzheimer's disease and Parkinson's disease highlight the effectiveness of HA-NegS in discovering new therapeutic applications for existing drugs. Cheng-Long Mi, Ling-Yun Dai, Junliang Shang, Juan Wang 0003, Feng Li 0033 |
IEEE J. Biomed. Health Informatics | 3 |
| 2026 | Epileptic Seizure Prediction Using Multi-Strategy Data Augmentation and Hierarchical Contrastive LearningabstractAccurate early prediction of epileptic seizures is crucial for improving patients' quality of life. However, existing seizure prediction methods often rely on large-scale labeled datasets and face challenges in generalization and real-time performance. To address these issues, this study proposes an efficient seizure prediction framework that achieves high performance even with limited labeled data, significantly reducing dependence on extensive annotations. To better distinguish preictal states, contrastive learning is employed to enhance feature separation between interictal and preictal periods, leading to improved sensitivity in detecting early seizure patterns. First, a data augmentation strategy is designed, incorporating wavelet-based frequency mixing, temporal masking, and window-based masking to enhance model robustness and generalization. Second, a hierarchical contrastive loss function is introduced, integrating instance-level and temporal contrastive learning to improve the model's ability to capture preictal patterns. Finally, a lightweight SE-EEGNet is developed and optimized as a feature extractor, strengthening critical feature extraction and enabling real-time seizure prediction. On the CHB-MIT dataset, the proposed method achieves 94.51% accuracy, 95.05% sensitivity, a 0.024/h false positive rate (FPR), and a 20.12-minute prediction time using only 30% labeled data. On the Siena dataset, it achieves 93.14% accuracy, 92.77% sensitivity, and a 0.030/h FPR. Moreover, performance improves further as the amount of labeled data increases, validating the effectiveness and practical applicability of the proposed approach in seizure prediction. Longfei Qi, Feng Li 0033, Junliang Shang, Shihan Wang 0009, Shasha Yuan |
IEEE J. Biomed. Health Informatics | 3 |
| 2026 | Cluster-Guided Contrastive Learning With Masked Autoencoder for Spatial Domain Identification Based on Spatial TranscriptomicsabstractRecent advancements in spatial transcriptomics technology have enabled the capture of gene expression profiles while maintaining spatial information. Accurately identifying spatial clustering plays a pivotal role in analyzing spatial transcriptomics data and understanding tissue microenvironments. However, current spatial domain identification methods cannot explore the complex relationship of gene expression profiles and spatial topology. To alleviate this issue, we propose STMCCL, a novel self-supervised learning framework that jointly trains a masked autoencoder and cluster-guided contrastive learning. This framework extracts informative latent representations from gene expression profiles and spatial information. Specifically, we first use data augmentation strategies to build augmented views and employ a masked encoder to generate a feature view. Then, encoders are applied to learn view-unique embeddings of each view. Furthermore, we introduce a multiple cluster-perspectives module that considers both geometric and structural relationships between clusters to produce more reliable cluster assignments. Finally, to derive more discriminative positives and negatives, the cluster-guided contrastive module calculates the confidence of each sample based on the initial cluster. Comprehensive experiments on 7 public datasets demonstrate that STMCCL outperforms the state-of-the-art baselines with finer-scale spatial domain identification. Juan Wang 0003, Shasha Yuan, Junliang Shang |
IEEE J. Biomed. Health Informatics | 4 |
| 2026 | MiRNA-Disease Association Prediction via Cosine Annealing and Multi-Head Self-Attention in HyperGCNabstractBiological studies have demonstrated that understanding the association between miRNAs and disease is critical for disease prevention, assessment, and therapy. However, traditional experimental methods for inferring these connections are not only costly but also inefficient. Hence, there is a pressing need to develop novel methods to improve the accuracy and efficiency of forecasting. Currently, graph convolutional networks (GCNs) techniques are one of the mainstream methods for predicting disease correlations. Nevertheless, traditional GCNs suffer from gradient vanishing and gradient explosion problems when dealing with long-range dependencies. To overcome these problems, we suggest a new approach called HGCMMDA, which relies on HyperGCN and combines a cosine annealing algorithm and a multi-head self-attention mechanism. In HGCMMDA, similarity networks for miRNAs and diseases are constructed, and GCN is used for feature extraction. A heterogeneity hypergraph is then built via HyperGCN for improved information propagation. Multi-head self-attention captures diverse node relations, while cosine annealing adjusts the learning rate. A combined BCE-Dice loss ensures accurate prediction. To evaluate the effectiveness of the proposed method, a comprehensive set of experiments was conducted using the Human microRNA Disease Database (HMDD v3.2). The method achieved a peak area under the receiver operating characteristic curve (AUC) of 0.9515, along with competitive performance in other evaluation metrics. The experimental findings indicate that HGCMMDA achieves notable enhancements over previously established approaches. These results strongly support the assertion that HGCMMDA serves as a dependable and effective framework for comprehensively exploring the intricate associations between microRNAs and human diseases. Zheng-Hua Chang, Jin-Xing Liu 0001, Junliang Shang |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Adaptive Weighting Contrastive Learning for Spatial Domain Identification in Spatial TranscriptomicsabstractThe rapid advancement of spatial transcriptomics has enabled the joint analysis of gene expression and spatial location data. This integration opens new avenues for uncovering tissue heterogeneity. Existing methods attempt to combine spatial and expression information to identify spatial domains. However, they often treat all information sources equally and do not account for their varying impact on results. To address this challenge, we propose AWCST, an adaptive weighted contrastive learning framework for spatial domain identification. AWCST first extracts spatial and expression latent representations and then fuses them using a multi-head attention mechanism. It measures the distributional differences between each view and the fused feature using Maximum Mean Discrepancy. These differences are converted into adaptive weights for the contrastive loss, enhancing the influence of high-quality information sources. Finally, we evaluate AWCST on two independent datasets to demonstrate its effectiveness. Xiyue Li, Ling-Yun Dai, Junliang Shang, Feng Li 0033 |
BIBM | 3 |
| 2025 | scGZDC: Graph-Based ZINB Deep Clustering for Single-Cell RNA-Seq DataabstractSingle-cell RNA sequencing (scRNA-seq) is a key technology for studying cellular heterogeneity. However, the high levels of sparsity and noise in scRNA-seq data present challenges for accurate cell clustering. To address this, we propose Graph-based ZINB Deep Clustering for Single-cell RNA-seq Data (scGZDC), a novel framework that operates within a variational autoencoder (VAE). The encoder of scGZDC employs a Graph Convolutional Network (GCN) to learn low-dimensional representations by leveraging both a preprocessed gene expression matrix and the cell-cell similarity graph. The decoder, in turn, employs a Graph Attention Network (GAT) to reconstruct gene expression counts via a Zero-Inflated Negative Binomial (ZINB) distribution, a distribution particularly well-suited for scRNA-seq data. To achieve end-to-end optimization, a Deep Embedding for Clustering (DEC) objective is integrated into the framework. Extensive experiments on public datasets demonstrate that scGZDC consistently outperforms existing methods. Our results show that unifying graph structural information with a suitable probabilistic model in an end-to-end clustering framework is an effective strategy for improving single-cell analysis. Hui-Bo Tian, Jin-Xing Liu 0001, Junliang Shang, Juan Wang 0003, Ling-Yun Dai |
BIBM | 4 |
| 2025 | GAEKLRR: A novel clustering method of the low-rank representation based on graph auto-encoder and relaxed k-means for single-cell type identificationabstractClustering is critical for scRNA-seq because it reveals the similarity of single-cell expression patterns. However, single-cell data contains numerous noises and outliers. Traditional clustering algorithms may fail to capture accurate clustering information. In this study, we propose the GAEKLRR method for single-cell type identification, which is a low-rank representation (LRR) method based on a graph autoencoder (GAE) and relaxed k-means. GAEKLRR consists of gedLRR and relaxed k-means. Among them, gedLRR is a GAEbased LRR algorithm that captures structural information and node features of samples using GAE. Relaxed k-means is a soft clustering method that can better preserve complex relationships between samples through soft partitioning. Specifically, to reduce the impact of noise and outliers on the mapping benchmark, GAEKLRR generates a robust graph embedding dictionary using gedLRR. Due to the reconstruction using inner product distance, the graph embedding dictionary has interpretability. Meanwhile, to capture accurate clustering information, GAEKLRR utilizes gedLRR to seek the LRR matrix of the graph embedding dictionary while using relaxed k -means to update the clustering centroid. It is worth noting that the continuous clustering indication matrix captured by relaxed k-means contains clustering labels, which can be used directly for clustering tasks. Finally, experiments on real singlecell datasets demonstrate that GAEKLRR has significant advantages for clustering. Linping Wang, Junliang Shang, Ling-Yun Dai, Juan Wang 0003 |
BIBM | 2 |
| 2025 | Epileptic Seizure Detection Using ECA-EEGNet with Earth Mover's Distance-Based Metric LearningabstractAccurate and efficient detection of epileptic seizures from electroencephalogram (EEG) signals is of great significance for clinical diagnosis and real-time monitoring. However, traditional EEG analysis methods face limitations in feature extraction and classification accuracy, primarily due to their heavy reliance on handcrafted features and rigid decision boundaries. Moreover, in clinical settings, the scarcity of seizure EEG signals and the difficulty in obtaining labeled data often lead to overfitting, especially when training data is insufficient. To address these challenges, this paper proposes a novel end-to-end seizure detection framework that integrates an attention-guided lightweight neural network with an advanced metric learning strategy. Based on the baseline EEGNet architecture, the proposed framework incorporates an Efficient Channel Attention (ECA) module to enhance the extraction of discriminative features from multichannel EEG signals. Furthermore, to improve the separability between seizure and non-seizure interictal states, we introduce a triplet loss-based metric learning method using the Earth Mover's Distance. By employing Earth Mover's Distance as the distance metric in the feature space and introducing a triplet loss function to constrain the relative distance relationships between samples, the proposed method ensures more compact embeddings for intra-class samples and better separation between inter-class feature distributions, thereby effectively mitigating overfitting risks under limited data conditions. Experimental evaluations on the CHB-MIT dataset demonstrate the superior performance of the proposed method, achieving an average accuracy of 97.51 %, sensitivity of 95.56 %, and specificity of 97.93 %. These results indicate that the proposed framework provides a promising and computationally efficient solution for automatic seizure detection in practical EEG analysis. Shihan Wang 0009, Junliang Shang, Juan Wang 0003, Longfei Qi, Shasha Yuan |
BIBM | 2 |
| 2025 | A Scalable and Unified Hierarchical Coarsening Hypergraph Framework for Single-Cell Multi-Omics Data AnalysisabstractNowadays, technological advancements enable the simultaneous profiling of the epigenome, transcriptome, and proteome at the single-cell level, and an urgent need for tools to integrate such complex tri-modal data is demanded. In this paper, we propose a novel method named scHCHF, a scalable and unified Hierarchical Coarsening Hypergraph Framework. scHCHF first employs a hierarchical coarsening process to groups cells into hypervertices. A unified multiomics hypergraph is then constructed upon these hypervertices, which substantially reduces computational complexity and memory footprint. Subsequently, a hypergraph attention network captures high-order relationships and learns discriminative cell representations, guided by a weakly supervised prototypical contrastive loss. Building upon the pre-trained network, scHCHF further employs a fine-tuning paradigm to achieve accurate cell type annotation using a minimal number of labeled cells. Experiments on real tri-modal datasets demonstrate that scHCHF outperforms other state-of-the-art methods in both cell clustering and cell type annotation tasks. Ke-Ze Yu, Xiang-Zhen Kong, Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Junliang Shang |
BIBM | 5 |
| 2025 | Spatial Multi-Omics Integration Via Information-Aware Multi-View Contrastive LearningabstractThe rapid advancement of spatial multi-omics technology enables the simultaneous acquisition of diverse expression data from the same tissue or slice. Different omics offer unique and critical information about the biological system. However, most existing methods are unable to fully utilize this information for downstream tasks such as spatial domain identification. To integrate this information effectively for downstream analysis, we introduce a novel Spatial Multi-omics data integration method based on Information-Aware Multi-view Contrastive Learning (SM-IAMCL). It optimizes the spatial and feature neighborhood graphs for each omics by the specific graph learner and fused graph learner, and learns the fused graph of spatial and feature neighborhood graphs at the same time. Then, to make fused graph of each omics integrate both shared and unique information of spatial and feature neighborhood graphs, we incorporate graph-level contrastive learning between different views in each omics. Finally, the learned fused representation of each omics is then integrated via a weighted fusion strategy to generate an integrated low-dimensional latent representation of spatial multiomics. This integrated representation is used for a variety of downstream analysis tasks. The experimental results show that SM-IAMCL outperforms other seven existing methods in the downstream tasks such as spatial domain identification. Conghui Zhang, Ling-Yun Dai, Juan Wang 0003, Junliang Shang, Feng Li 0033 |
BIBM | 5 |
| 2025 | PDA-PAGCN: Predicting Disease-Related PiRNA Based on Proxy Attention Graph Convolutional Network
Xiaotong Kong, Xianghan Meng, Junliang Shang, Linqian Zhao, Jin-Xing Liu 0001 |
ICIC (26) | 3 |
| 2025 | Label-Guided Graph Contrastive Learning for Single-Cell Fusion Clustering
Baojuan Qin, Junliang Shang, Yan Zhao 0045, Feng Li 0033, Jin-Xing Liu 0001 |
ISBRA (1) | 2 |
| 2025 | A Neighborhood Selection Learning Artificial Bee Colony Algorithm Based on Population Backtracking for Detecting Epistatic Interactions
Xiaoqi Tang, Linqian Zhao, Junliang Shang, Feng Li 0033, Jin-Xing Liu 0001 |
ISBRA (1) | 5 |
| 2025 | SUIFS: A Symmetric Uncertainty Based Interactive Feature Selection Method
Junliang Shang, Qianqian Ren, Feng Li 0033 |
ISBRA (1) | 4 |
| 2025 | PDA-GTGCN: Identification of PiRNA-Disease Associations Based on Group Feature Transformation Graph Convolutional Network
Xiaoqi Tang, Xianghan Meng, Junliang Shang, Baojuan Qin, Feng Li 0033 |
ISBRA (1) | 3 |
| 2025 | scRDAN: a robust domain adaptation network for cell type annotation across single-cell RNA sequencing dataabstractSingle-cell RNA sequencing technology facilitates the recognition of diverse cell types and subgroups, playing a crucial role in investigating cellular heterogeneity. Cell type annotation, a crucial process in single-cell RNA sequencing analysis, is often influenced by noise and batch effects. To address these challenges, we propose scRDAN, which is a robust domain adaptation network comprising three modules: the denoising domain adaptation module, the fine-grained discrimination module, and the robustness enhancement module. The denoising domain adaptation module mitigates noise interference through feature reconstruction in domains, while leveraging adversarial learning to align data distributions, improving annotation accuracy and robustness against batch effects. The fine-grained discrimination module maintains intra-class compactness and enhances inter-class separability, reducing feature overlap and improving cell type distinction. Finally, the robustness enhancement module introduces noise from various perspectives in both domains, enhancing robustness and generalization. We evaluate scRDAN on simulated, cross-platforms, and cross-species datasets, comparing it with advanced methods. Results demonstrate that scRDAN outperforms existing methods in handling batch effects and cell type annotation. Junliang Shang, Baojuan Qin |
Briefings Bioinform. | 3 |
| 2025 | SAMGCN: A spatially-augmented multi-view graph convolutional network for identifying spatial domains
Hao Liu 0075, Ying-Lian Gao, Cui-Na Jiao, Junliang Shang |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | A Contrastive Learning-Enhanced Residual Network for Predicting Epileptic Seizures Using EEG SignalsabstractThe models used to predict epileptic seizures based on electroencephalogram (EEG) signals often encounter substantial challenges due to the requirement for large, labeled datasets and the inherent complexity of EEG data, which hinders their robustness and generalization capability. This study proposes CLResNet, a framework for predicting epileptic seizures, which combines contrastive self-supervised learning with a modified deep residual neural network to address the above challenges. In contrast to traditional models, CLResNet uses unlabeled EEG data for pre-training to extract robust feature representations. It is then fine-tuned on a smaller labeled dataset to significantly reduce its reliance on labeled data while improving its efficiency and predictive accuracy. The contrastive learning (CL) framework enhances the ability of the model to distinguish between preictal and interictal states, thus improving its robustness and generalizability. The architecture of CLResNet contains residual connections that enable it to learn deep features of the data and ensure an efficient gradient flow. The results of the evaluation of the model on the CHB-MIT dataset showed that it outperformed prevalent methods in the field, with an accuracy of 92.97%, sensitivity of 94.18%, and false-positive rate of 0.043/h. On the Siena dataset, the model also achieved competitive performance, with an accuracy of 92.79%, a sensitivity of 91.47%, and a false-positive rate of 0.041/h. These results confirm the effectiveness of CLResNet in addressing variations in EEG data, and show that contrastive self-supervised learning is a robust and accurate approach for predicting seizures. Longfei Qi, Shasha Yuan, Feng Li 0033, Junliang Shang, Juan Wang 0003, Shihan Wang 0009 |
Int. J. Neural Syst. | 4 |
| 2025 | scCDAN: Constraint domain adaptation network for cell type annotation across single cell RNA sequencing data
Junliang Shang, Yan Zhao 0045, Baojuan Qin, Xianghan Meng, Jin-Xing Liu 0001 |
Neurocomputing | 1 |
| 2025 | stMHCG: High-confidence multi-view clustering for identification of spatial domains from spatially resolved transcriptomics
Junliang Shang, Yan Zhao 0045, Baojuan Qin, Qianqian Ren, Feng Li 0033, Jin-Xing Liu 0001 |
Neurocomputing | 2 |
| 2025 | RPMVCDA: Random Perturbation and Multi-View Graph Convolutional Networks for CircRNA-Disease Association PredictionabstractNumerous studies have demonstrated the regulatory role of circular RNA (circRNA) in various diseases, emphasizing the importance of identifying disease-related circRNAs. Although several computational models have been developed to predict circRNA-disease associations, the limited number of experimentally validated associations has resulted in the sparse association network. Therefore, there is a need for continuously improving circRNA-disease prediction models. In this study, we propose RPMVCDA, a computational model based on random perturbation and multi-view graph convolutional networks (GCNs), to predict circRNA-disease associations. Specifically, RPMVCDA first constructs multiple similarity networks of circRNAs and diseases, applying multi-view GCNs to obtain embedding representations. Second, to enable message passing between circRNA-disease samples, RPMVCDA constructs the feature similarity association network. Third, RPMVCDA introduces a random perturbation association network to further explore the potential associations, which is the highlight of the RPMVCDA. Finally, based on these three association networks, RPMVCDA utilizes the self-attention mechanism to generate high-quality features for circRNAs and diseases, which are used to calculate association scores. To evaluate the performance of RPMVCDA, five-fold cross-validation and case studies on the CircR2Disease dataset are performed, results of which shows that RPMVCDA outperforms the compared models, implying that it might be an alternative for predicting circRNA-disease associations. Xin He 0008, Junliang Shang, Feng Li 0033, Jin-Xing Liu 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2025 | MPSO-CD: A Multi-Objective Particle Swarm Optimization Community Detection Method for Identifying Disease ModulesabstractThe dysfunction of biological systems caused by disease-related genes is one of the inducements of complex diseases. To understand molecular mechanisms of complex diseases, the identification of disease-related gene modules in biological networks through community detection is emerging as a promising approach. However, most community detection methods are not suitable for biological networks because their topological structures are complex and the scale of biologically relevant modules are small. In this paper, a novel community detection method called MPSO-CD was proposed based on multi-objective particle swarm optimization, in which negative ratio association and ratio cut were employed as objective functions. Highlights of MPSO-CD are a mutation strategy based on clustering coefficient and the procedure of disease module screening referring to the internal connection density and functional similarity. Experimental results of social and synthetic complex networks indicate that MPSO-CD is comparable and often superior to four compared methods. Eventually, MPSO-CD is applied to the asthma gene co-expression network for identifying potential disease modules that provide the molecular mechanism information about asthma. Most of the captured modules have been proven to be associated with asthma through Gene Ontology and pathway enrichment analysis. Xuhui Zhu, Mingyuan Bi, Junliang Shang, Feng Li 0033, Yuanyuan Zhang 0008, Ling-Yun Dai, Shengjun Li, Jin-Xing Liu 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | TEMCL: Prediction of Drug-Disease Associations Based on Transformer and Enhanced Multi-View Contrastive LearningabstractDrug repositioning (DR) has emerged as an effective method of identifying new indications for existing drugs. Many DR methods have demonstrated superior performance. However, most of them utilize a limited number of biological entities, ignoring the critical role of other entities in addressing data sparsity as well as improving model generalization capabilities. In addition, fully capturing high-order information of biological data still needs to be fully explored. To address above issues, a model based on transformer and enhanced multi-view contrastive learning (TEMCL) is proposed for predicting drug-disease associations (DDAs). Firstly, transformer is employed to obtain high-order features of nodes from similarity information. Secondly, based on similarity matrices and association matrices of nodes, two different types of views are constructed, i.e., homogeneous hypergraphs and heterogeneous association graphs. Among them, to alleviate sparsity problem existing in heterogeneous graphs, protein nodes as well as meta-path enhancement strategy are introduced. Thirdly, hypergraph convolutional network and heterogeneous graph transformer are used to extract node features on above two types of views, respectively. Contrastive learning is applied to obtain more representative features. Finally, multilayer perceptron (MLP) is used for predicting DDAs. Experiments show that TEMCL outperforms existing methods on DR task, exhibiting superior performance. In addition, case studies further demonstrate the effectiveness of this model. TEMCL provides new insights for identifying novel DDAs. Ming-Li Cui, Cui-Na Jiao, Ying-Lian Gao, Junliang Shang, Chun-Hou Zheng 0001, Jin-Xing Liu 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | pscAdapt: Pre-Trained Domain Adaptation Network Based on Structural Similarity for Cell Type Annotation in Single Cell RNA-seq DataabstractCell type annotation refers to the process of categorizing and labeling cells to identify their specific cell types, which is crucial for understanding cell functions and biological processes. Although many methods have been developed for automated cell type annotation, they often encounter challenges such as batch effects due to variations in data distribution across platforms and species, thereby compromising their performance. To address batch effects, in this study, a pre-trained domain adaptation model based on structural similarity, named pscAdapt, is proposed for cell type annotation. Specifically, a pre-trained strategy is employed to initialize model parameters to learn the data distribution of source domain. This strategy is also combined with an adversarial learning strategy to train the domain adaptation network for achieving domain level alignment and reducing domain discrepancy. Furthermore, to better distinguish different types of cells, a structural similarity loss is designed, aiming to shorten distances between cells of the same type and increase distances between cells of different types in feature space, thus achieving cell level alignment and enhancing the discriminability of cell types. Comprehensive experiments were conducted on simulated datasets, cross-platforms datasets and cross-species datasets to validate the effectiveness of pscAdapt, results of which demonstrate that pscAdapt outperforms several popular cell type annotation methods. Yan Zhao 0045, Junliang Shang, Baojuan Qin, Xin He 0008, Qianqian Ren, Jin-Xing Liu 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | BIGFormer: A Graph Transformer With Local Structure Awareness for Diagnosis and Pathogenesis Identification of Alzheimer's Disease Using Imaging Genetic DataabstractAlzheimer's disease (AD) is a highly inheritable neurological disorder, and brain imaging genetics (BIG) has become a rapidly advancing field for comprehensive understanding its pathogenesis. However, most of the existing approaches underestimate the complexity of the interactions among factors that cause AD. To take full appreciate of these complexity interactions, we propose BIGFormer, a graph Transformer with local structural awareness, for AD diagnosis and identification of pathogenic mechanisms. Specifically, the factors interaction graph is constructed with lesion brain regions and risk genes as nodes, where the connection between nodes intuitively represents the interaction between nodes. After that, a perception with local structure awareness is built to extract local structure around nodes, which is then injected into node representation. Then, the global reliance inference component assembles the local structure into higher-order structure, and multi-level interaction structures are jointly aggregated into a classification projection head for disease state prediction. Experimental results show that BIGFormer demonstrated superiority in four classification tasks on the AD neuroimaging initiative dataset and proved to identify biomarkers closely intimately related to AD. Qi Zou 0003, Junliang Shang, Jin-Xing Liu 0001, Rui Gao 0006 |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | A multi-objective genetic algorithm based on neighborhood coevolution for community detectionabstractCommunity detection has attracted growing interest, with multi-objective evolutionary algorithms proving to be highly competitive in this area. In this paper, a community detection method based on a multi-objective neighborhood coevolution genetic algorithm, NCMOGA, is proposed. To improve the computational efficiency in large-scale networks, NCMOGA introduces a network processing strategy to simplify the network before and during evolution. A neighborhood coevolution strategy is proposed, in which the corresponding subpopulation is formed according to the neighborhood of each individual. A series of operations such as crossover, mutation and update are performed in the subpopulation, emphasizing the synergy between individuals and their neighbors. Mating selection and crossover operations are performed based on the center selection idea of density peak clustering, and the most important nodes are selected to generate offspring. The effectiveness of NCMOGA is verified on synthetic networks and real-world networks. In addition, the results in guiding the classification of disease and healthy samples demonstrate the high quality of the modules detected by NCMOGA. Mingyuan Bi, Junliang Shang, Xiaotong Kong, Feng Li 0033, Yuanyuan Zhang 0008, Jin-Xing Liu 0001 |
BIBM | 2 |
| 2024 | Integrating Local and Global Information to Decipher Spatial Domains of Spatial Transcriptomics by Attention-based Graph Convolutional NetworkabstractRecent developments in spatial transcriptomics technologies have made it possible to obtain gene expression profiles while maintaining spatial context. Precisely identifying spatial domains is essential for downstream analysis, requiring the effective integration of gene expression profiles with spatial information. To overcome the challenge of low accuracy in spatial domain identification, this paper proposed a deep learning model called LGAGCN based on local and global information. It used graph convolutional network to learn the features of local and global views and employed an attention mechanism to integrate embeddings from different views. Moreover, experiments were conducted on the human dorsolateral prefrontal cortex (DLPFC) dataset and the human breast cancer (HBC) dataset to evaluate the effectiveness of the model. The experimental results showed that LGAGCN outperformed state-of-the-art methods in spatial clustering task. Xu-Ran Dou, Junliang Shang, Chun-Hou Zheng 0001, Ying-Lian Gao, Jin-Xing Liu 0001 |
BIBM | 3 |
| 2024 | Spatial domains identification based on multi-view contrastive learning in spatial transcriptomicsabstractSpatial transcriptomic techniques can be used to obtain transcriptome data from different locations in tissues. The identification of spatial domains is a key task in the analysis of spatial transcriptomic data. Therefore, we propose a multi-view contrastive learning framework named MVCLST for identifying spatial domains. First, MVCLST introduces pathway information on the basis of spatial transcriptomic data to consider the functional correlation between spots. In order to better mine the underlying biological features, MVCLST constructs three biological networks from three biological perspectives: spatial distribution, similarity of gene expression and functional correlation. Secondly, MVCLST introduces a multi-view contrastive learning method, which fully considers the contrastive relationship between multiple views to improve the accuracy and reliability of feature extraction. Then, in order to obtain a more abundant feature representation, the adaptive attention mechanism is used to integrate the common features and specific features. Finally, we compare MVCLST with five spatial transcriptomic methods to verify the accuracy of MVCLST for spatial domains identification. The experimental results show that MVCLST is better than the other five methods. Yanru Gao, Feng Li 0033, Fanhao Meng, Qianqian Ren, Junliang Shang |
BIBM | 6 |
| 2024 | Integrating Autoencoder and Multi-View Graph Convolutional Networks for Spatial Domains IdentificationabstractSpatial transcriptomics technology provides high-resolution gene expression profiles and spatial location information, offering a revolutionary approach to identify tissue regions and cell types. However, due to the characteristics of gene expression data such as high dimension and high noise, some information may be lost when performing feature extraction. Therefore, we propose a framework that integrates an autoencoder and a multi-view graph convolutional network, IAMGCN, to reduce information loss and capture more comprehensive information. Specifically, we use the autoencoder to extract the deep representation of the gene expression matrix. The middle layer of the autoencoder is integrated into the middle layer of the graph convolutional network. This integration aims to reduce the information loss during the information aggregation process of the graph convolutional network. In order to learn the specificity and common information of multiple views, we introduce a contrastive learning strategy, which takes the output of the autoencoder as a positive sample, and reduce the distance between the two views and the autoencoder output, respectively. We test IAMGCN on two datasets and compare it with five other methods, and the experimental results show that IAMGCN achieves the highest accuracy on all datasets. Fanhao Meng, Feng Li 0033, Yanru Gao, Junliang Shang |
BIBM | 6 |
| 2024 | MNGCCL: Multi-neighborhood graph collaborative contrastive learning for drug-disease association predictionabstractExploring new therapeutic applications for existing drugs can effectively reduce drug development costs. However, current drug-disease association (DDA) prediction methods often fail to effectively integrate multi-domain information. The lack of multi-domain information integration causes these methods to heavily rely on prior knowledge, thereby limiting their generalization ability. To address this issue, we developed a Multi-Domain Graph Collaborative Contrastive Learning (MNGCCL) model for DDA prediction. In the MNGCCL framework, a feature extraction module is designed to effectively extract both single-domain and multi-domain features. The single-domain and multi-domain feature extraction components in this module run in parallel, extracting key features of drugs and diseases from different latent spaces (e.g., homogeneous and heterogeneous networks). MNGCCL employs graph collaborative contrastive learning to integrate these features and enhances information interaction by designing new node scoring for negative sample sampling. This significantly enriches the semantic features of drugs and diseases. In DDA prediction, MNGCCL outperforms other state-of-the-art models across various datasets and partitioning methods. Notably, MNGCCL excels in drug repositioning for Parkinson’s disease and Alzheimer’s disease, as well as handling data sparsity. These findings highlight its tremendous potential for drug repositioning and DDA prediction, especially in the context of sparse omics data. Cheng-Long Mi, Jin-Xing Liu 0001, Junliang Shang, Juan Wang 0003, Ling-Yun Dai |
BIBM | 4 |
| 2024 | SeizureGuard: an adaptive fusion converter model for seizure predictionabstractThe majority of studies on seizure prediction have concentrated on temporal and frequency features, which may result in inconsistencies when evaluating the accuracy of pre-dictions. To address this issue, we propose an innovative seizure prediction model that incorporates channel features as one of the key component. The model effectively integrates time, frequency and channel features by combining three convolutional neural networks (CNNs) and three transformer encoders. We validated our method on the CHB-MIT dataset, obtaining 99.00% accuracy, 99.10% sensitivity and 0.011/h false prediction rate (FPR). The experimental results demonstrate the efficient performance of our method in EEG signal classification and prediction, or comparable to state-of-the-art models. Wen-Xin Pan, Jin-Xing Liu 0001, Junliang Shang |
BIBM | 4 |
| 2024 | Multi-Population Ant Colony Optimization With Knowledge-Based Local Searches for Epistasis DetectionabstractAnalysis of epistatic interaction is an important means to study the pathogenesis of complex diseases in genome-wide association studies (GWAS). Epistatic interaction detection aims to identify the ideal combination among single nucleotide polymorphisms (SNPs) and determine whether this combination is significantly associated with complex diseases. However, they suffer from certain limitations, such as low detection power and long execution times. Therefore, this paper proposes a multi-population ant colony optimization algorithm with an adaptive heuristic strategy (MPACO-AHS). MPACO-AHS is a framework based on multi-population approaches, where multiple populations are employed to detect epistatic interactions, helping to avoid the data bias inherent in a single population. Moreover, to guide the search direction of each population, an adaptive heuristic strategy is introduced, allowing the algorithm to focus on areas more likely to contain epistatic interactions, thereby improving the accuracy of the results. Comprehensive experiments are conducted on simulated datasets. The results demonstrate that MPACO-AHS outperforms existing algorithms by overcoming the challenges in detecting epistatic interactions in GWAS. Qianqian Ren, Shaoyi Liu, Lianlian Zhang, Junliang Shang, Feng Li 0033 |
BIBM | 4 |
| 2024 | A Particle Swarm Optimization Algorithm Based on Multi-Population Mutual Learning for SNP-SNP Interaction DetectionabstractSingle nucleotide polymorphism (SNPs) data have become abundant thanks to the quick advancement of high-throughput sequencing technology, which provides convenience for genome-wide association studies. Single SNPs have been proven to be the cause of some diseases, and the emergence of complex diseases is often thought to be the result of the interaction of multiple SNPs. However, the possible interaction of millions of SNPs imposes a heavy computational burden for uncovering complex disease mechanisms. The existing SNP-SNP interaction detection algorithms frequently have flaws including high computation complexity and poor optimization effectiveness. In this study, a particle swarm optimization algorithm based on multi-population mutual learning (PSOMPML) is proposed to detect SNP-SNP interactions. In this algorithm, the mutual learning strategy is introduced to deal with different particles in different sub-populations to facilitate knowledge exchange. In addition, the elite preservation mechanism is incorporated into PSOMPML, to better preserve the good SNPs in the elite particles. The promising region local search strategy searches the optimal solution along the target solution and its near space to increase the convergence speed of the proposed algorithm. Experiments on simulated data sets and real data also demonstrate the effectiveness of the proposed algorithm. Linqian Zhao, Yahan Li, Junliang Shang, Qianqian Ren, Yuanyuan Zhang 0008, Jin-Xing Liu 0001 |
BIBM | 3 |
| 2024 | CPSORCL: A Cooperative Particle Swarm Optimization Method with Random Contrastive Learning for Interactive Feature Selection
Junliang Shang, Yahan Li, Feng Li 0033, Yuanyuan Zhang 0008, Jin-Xing Liu 0001 |
ISBRA (2) | 1 |
| 2024 | Multi-modal imaging genetics data fusion by deep auto-encoder and self-representation network for Alzheimer's disease diagnosis and biomarkers extraction
Cui-Na Jiao, Ying-Lian Gao, Junliang Shang, Jin-Xing Liu 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | A Delayed Spiking Neural Membrane System for Adaptive Nearest Neighbor-Based Density Peak ClusteringabstractAlthough the density peak clustering (DPC) algorithm can effectively distribute samples and quickly identify noise points, it lacks adaptability and cannot consider the local data structure. In addition, clustering algorithms generally suffer from high time complexity. Prior research suggests that clustering algorithms grounded in P systems can mitigate time complexity concerns. Within the realm of membrane systems (P systems), spiking neural P systems (SN P systems), inspired by biological nervous systems, are third-generation neural networks that possess intricate structures and offer substantial parallelism advantages. Thus, this study first improved the DPC by introducing the maximum nearest neighbor distance and K-nearest neighbors (KNN). Moreover, a method based on delayed spiking neural P systems (DSN P systems) was proposed to improve the performance of the algorithm. Subsequently, the DSNP-ANDPC algorithm was proposed. The effectiveness of DSNP-ANDPC was evaluated through comprehensive evaluations across four synthetic datasets and 10 real-world datasets. The proposed method outperformed the other comparison methods in most cases. Qianqian Ren, Lianlian Zhang, Shaoyi Liu, Jin-Xing Liu 0001, Junliang Shang, Xiyu Liu 0001 |
Int. J. Neural Syst. | 5 |
| 2024 | Enhancing Spatial Domain Identification in Spatially Resolved Transcriptomics Using Graph Convolutional Networks With Adaptively Feature-Spatial Balance and Contrastive LearningabstractRecent advancements in spatially transcriptomics (ST) technologies have enabled the comprehensive measurement of gene expression profiles while preserving the spatial information of cells. Combining gene expression profiles and spatial information has been the most commonly used method to identify spatial functional domains and genes. However, most existing spatial domain decipherer methods are more focused on spatially neighboring structures and fail to take into account balancing the self-characteristics and the spatial structure dependency of spots. Therefore, we propose a novel model called SpaGCAC, which recognizes spatial domains with the help of an adaptive feature-spatial balanced graph convolutional network named AFSBGCN. The AFSBGCN can dynamically learn the relationship between spatial local topology structures and the self-characteristics of spots by adaptively increasing or declining the weight on the self-characteristics during message aggregation. Moreover, to better capture the local structures of spots, SpaGCAC exploits a local topology structure contrastive learning strategy. Meanwhile, SpaGCAC utilizes a probability distribution contrastive learning strategy to increase the similarity of probability distributions for points belonging to the same category. We validate the performance of SpaGCAC for spatial domain identification on four spatial transcriptomic datasets. In comparison with seven spatial domain recognition methods, SpaGCAC achieved the highest NMI median of 0.683 and the second highest ARI median of 0.559 on the multi-slice DLPFC dataset. SpaGCAC achieved the best results on all three other single-slice datasets. The above-mentioned results show that SpaGCAC outperforms most existing methods, providing enhanced insights into tissue heterogeneity. Xuena Liang, Junliang Shang, Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Juan Wang 0003 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2024 | A New Graph Autoencoder-Based Multi-Level Kernel Subspace Fusion Framework for Single-Cell Type IdentificationabstractThe advent of single-cell RNA sequencing (scRNA-seq) technology offers the opportunity to conduct biological research at the cellular level. Single-cell type identification based on unsupervised clustering is one of the fundamental tasks of scRNA-seq data analysis. Although many single-cell clustering methods have been developed recently, few can fully exploit the deep potential relationships between cells, resulting in suboptimal clustering. In this paper, we propose scGAMF, a graph autoencoder-based multi-level kernel subspace fusion framework for scRNA-seq data analysis. Based on multiple top feature sets, scGAMF unifies deep feature embedding and kernel space analysis into a single framework to learn an accurate clustering affinity matrix. First, we construct multiple top feature sets to avoid the high variability caused by single feature set learning. Second, scGAMF uses a graph autoencoder (GAEs) to extract deep information embedded in the data, and learn embeddings including gene expression patterns and cell-cell relationships. Third, to fully explore the deep potential relationships between cells, we design a multi-level kernel space fusion strategy. This strategy uses a kernel expression model with adaptive similarity preservation to learn a self-expression matrix shared by all embedding spaces of a given feature set, and a consensus affinity matrix across multiple top feature sets. Finally, the consensus affinity matrix is used for spectral clustering, visualization, and identification of gene markers. Extensive validation on real datasets shows that scGAMF achieves higher clustering accuracy than many popular single-cell analysis methods. Juan Wang 0003, Tian-Jing Qiao, Chun-Hou Zheng 0001, Jin-Xing Liu 0001, Junliang Shang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2024 | A Clustering Method for Single-Cell RNA-Seq Data Based on Automatic Weighting Penalty and Low-Rank RepresentationabstractAdvances in high-throughput single-cell RNA sequencing (scRNA-seq) technology have provided more comprehensive biological information on cell expression. Clustering analysis is a critical step in scRNA-seq research and provides clear knowledge of the cell identity. Unfortunately, the characteristics of scRNA-seq data and the limitations of existing technologies make clustering encounter a considerable challenge. Meanwhile, some existing methods treat different features equally and ignore differences in feature contributions, which leads to a loss of information. To overcome limitations, we introduce a weighted distance constraint into the construction of the similarity graph and combine the similarity constraint. We propose the Joint Automatic Weighting Similarity Graph and Low-rank Representation (JAGLRR) clustering method. Evaluating the contributions of each feature and assigning various weight values can increase the significance of valuable features while decreasing the interference of redundant features. The similarity constraint allows the model to generate a more symmetric affinity matrix. Benefitting from that affinity matrix, JAGLRR recovers the original linear relationship of the data more accurately and obtains more discriminative information. The results on simulated datasets and 8 real datasets show that JAGLRR outperforms 11 existing comparison methods in clustering experiments, with higher clustering accuracy and stability. Juan Wang 0003, Zhen-Chang Wang, Shasha Yuan, Chun-Hou Zheng 0001, Jin-Xing Liu 0001, Junliang Shang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2024 | Diagnosis-Guided Deep Subspace Clustering Association Study for Pathogenetic Markers Identification of Alzheimer's Disease Based on Comparative AtlasesabstractThe roles of brain region activities and genotypic functions in the pathogenesis of Alzheimer's disease (AD) remain unclear. Meanwhile, current imaging genetics methods are difficult to identify potential pathogenetic markers by correlation analysis between brain network and genetic variation. To discover disease-related brain connectome from the specific brain structure and the fine-grained level, based on the Automated Anatomical Labeling (AAL) and human Brainnetome atlases, the functional brain network is first constructed for each subject. Specifically, the upper triangle elements of the functional connectivity matrix are extracted as connectivity features. The clustering coefficient and the average weighted node degree are developed to assess the significance of every brain area. Since the constructed brain network and genetic data are characterized by non-linearity, high-dimensionality, and few subjects, the deep subspace clustering algorithm is proposed to reconstruct the original data. Our multilayer neural network helps capture the non-linear manifolds, and subspace clustering learns pairwise affinities between samples. Moreover, most approaches in neuroimaging genetics are unsupervised learning, neglecting the diagnostic information related to diseases. We presented a label constraint with diagnostic status to instruct the imaging genetics correlation analysis. To this end, a diagnosis-guided deep subspace clustering association (DDSCA) method is developed to discover brain connectome and risk genetic factors by integrating genotypes with functional network phenotypes. Extensive experiments prove that DDSCA achieves superior performance to most association methods and effectively selects disease-relevant genetic markers and brain connectome at the coarse-grained and fine-grained levels. Cui-Na Jiao, Junliang Shang, Feng Li 0033, Xinchun Cui, Yan-Li Wang, Ying-Lian Gao, Jin-Xing Liu 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | SGFCCDA: Scale Graph Convolutional Networks and Feature Convolution for circRNA-Disease Association PredictionabstractCircular RNAs (circRNAs) have emerged as a novel class of non-coding RNAs with regulatory roles in disease pathogenesis. Computational models aimed at predicting circRNA-disease associations offer valuable insights into disease mechanisms, thereby enabling the development of innovative diagnostic and therapeutic approaches while reducing the reliance on costly wet experiments. In this study, SGFCCDA is proposed for predicting potential circRNA-disease associations based on scale graph convolutional networks and feature convolution. Specifically, SGFCCDA integrates multiple measures of circRNA and disease similarity and combines known association information to construct a heterogeneous network. This network is then explored by scale graph convolutional networks to capture both topological and attribute information. Additionally, convolutional neural networks are employed to further learn the features and obtain higher-order feature representations containing richer information about nodes. The Hadamard product is utilized to effectively combine circRNA features with disease features, and a multilayer perceptron is applied to predict the association between each pair of circRNA and disease. Five-fold cross validation experiments conducted on the CircR2Disease dataset demonstrate the accurate prediction capabilities of SGFCCDA in identifying potential circRNA-disease associations. Furthermore, case studies provide further confirmation of SGFCCDA's ability to identify disease-associated circRNAs. Junliang Shang, Linqian Zhao, Xin He 0008, Xianghan Meng, Feng Li 0033, Jin-Xing Liu 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | Spatiotemporal Network Based on GCN and BiGRU for Seizure DetectionabstractAs an important tool for detecting and diagnosing epilepsy, multi-channel EEG records the neuronal activities of different brain regions. Visual identification of abnormal EEG signals poses challenges, making the use of artificial intelligence techniques for automated seizure detection an inevitable trend. However, existing seizure detection methods often overlook the spatial relationship between EEG channels, which can't take full advantage of brain network structure. In this paper, we design an end-to-end spatiotemporal architecture for seizure detection based on Graph Convolutional Networks (GCN) and Bidirectional Gated Recurrent Units (BiGRU) to efficiently model the spatial dependence and temporal dynamics of EEG. Firstly, the original EEG signals are preprocessed by applying wavelet transform for temporal-frequency analysis. The Pearson correlation matrix is computed for specific frequency bands and GCN is utilized to extract spatial features between EEG channels. Then, these features are sent into the BiGRU network to capture temporal relationships. Finally, the detection decisions are achieved using fully connected layers and the multi-level decision rules are implemented to provide the final results. The proposed method is validated on CHB-MIT EEG dataset, achieving 98.85% sensitivity, 95.83% specificity, 97.35% accuracy, 97.4% F1-score, and 97.33% AUC. This network fusions multiple EEG characteristics in the spatial-temporal-frequency domains to improve the detection performance and the promising result demonstrates that the performance of this model is superior to or on par with existing methods. Jie Xu 0059, Shasha Yuan, Junliang Shang, Juan Wang 0003, Kuiting Yan, Yankai Yang |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | FSCME: A Feature Selection Method Combining Copula Correlation and Maximal Information Coefficient by Entropy WeightsabstractFeature selection is a critical component of data mining and has garnered significant attention in recent years. However, feature selection methods based on information entropy often introduce complex mutual information forms to measure features, leading to increased redundancy and potential errors. To address this issue, we propose FSCME, a feature selection method combining Copula correlation (Ccor) and the maximum information coefficient (MIC) by entropy weights. The FSCME takes into consideration the relevance between features and labels, as well as the redundancy among candidate features and selected features. Therefore, the FSCME utilizes Ccor to measure the redundancy between features, while also estimating the relevance between features and labels. Meanwhile, the FSCME employs MIC to enhance the credibility of the correlation between features and labels. Moreover, this study employs the Entropy Weight Method (EWM) to evaluate and assign weights to the Ccor and MIC. The experimental results demonstrate that FSCME yields a more effective feature subset for subsequent clustering processes, significantly improving the classification performance compared to the other six feature selection methods. Junliang Shang, Qianqian Ren, Feng Li 0033, Cui-Na Jiao, Jin-Xing Liu 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | MGCNRF: Prediction of Disease-Related miRNAs Based on Multiple Graph Convolutional Networks and Random ForestabstractIncreasing microRNAs (miRNAs) have been confirmed to be inextricably linked to various diseases, and the discovery of their associations has become a routine way of treating diseases. To overcome the time-consuming and laborious shortcoming of traditional experiments in verifying the associations of miRNAs and diseases (MDAs), a variety of computational methods have emerged. However, these methods still have many shortcomings in terms of predictive performance and accuracy. In this study, a model based on multiple graph convolutional networks and random forest (MGCNRF) was proposed for the prediction MDAs. Specifically, MGCNRF first mapped miRNA functional similarity and sequence similarity, disease semantic similarity and target similarity, and the known MDAs into four different two-layer heterogeneous networks. Second, MGCNRF applied four heterogeneous networks into four different layered attention graph convolutional networks (GCNs), respectively, to extract MDA embeddings. Finally, MGCNRF integrated the embeddings of every MDA into the features of the miRNA-disease pair and predicted potential MDAs through the random forest (RF). Fivefold cross-validation was applied to verify the prediction performance of MGCNRF, which outperforms the other seven state-of-the-art methods by area under curve. Furthermore, the accuracy and the case studies of different diseases further demonstrate the scientific rationale of MGCNRF. In conclusion, MGCNRF can serve as a scientific tool for predicting potential MDAs. Feng Li 0033, Boxin Guan, Jin-Xing Liu 0001, Junliang Shang |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | GRPGAT: Predicting CircRNA-disease Associations Based on Graph Random Propagation Network and Graph Attention NetworkabstractCircRNA as a biomarker has been shown to have an essential effect on the occurrence and prognosis of a wide range of human diseases. Because of the high cost of wet experiments, computational methods are widely used to explore circRNA. However, the performance and robustness of the computational models still need to be further improved. To solve these problems, this paper proposes a novel method based on graph random propagation network and multi-head dynamic graph attention network (GRPGAT) to predict the potential associations between circRNAs and diseases. Firstly, GRPGAT uses centered kernel alignment method to fuse the circRNA similarity kernels and disease similarity kernels. Then the integrated vectors build a heterogeneous graph and are sent to a graph random propagation network. The remaining nodes are fed into a multi-head dynamic attention network for feature extraction. Finally, a four-layer Multilayer Perceptron is used to learn features and gain the prediction scores. Experiments are supported by cirR2Disease, and achieve Area Under Curve (AUC) scores of 0.9636 in 5-fold cross validation. In comparison with the state-of-the-art models, GRPGAT also shows superior performance. Wen-Yue Kang, Chun-Hou Zheng 0001, Ying-Lian Gao, Juan Wang 0003, Junliang Shang, Jin-Xing Liu 0001 |
BIBM | 5 |
| 2023 | idenLD-AREL: identifying lncRNA-disease associations by random forests based on an ensemble learning frameworkabstractIdentification of disease-associated long non-coding RNAs (lncRNAs) facilitates the understanding of the pathogenesis of complex diseases. Many different types of computational models have been proposed. Although some of them have achieved encouraging results in predicting disease-associated lncRNAs, how to obtain stable results is still a challenge. In this paper, we propose a computational model based on an ensemble learning framework via the adaptive random forests, in short, idenLD-AREL. The idenLD-AREL integrates multiple random forest predictors and adaptive strategies to predict the scores of potential lncRNA-disease associations (LDAs), which ensure the stability and accuracy of the prediction results. In addition, there are a large number of false negative samples in the association datasets. For this reason, the resampling strategy is applied to idenLD-AREL to balance the samples. The idenLD-AREL is assessed by five-fold cross-validation in both the benchmark dataset and independent test set, showing excellent performance. Besides, the experimental results of the case study further demonstrate the effectiveness of the idenLD-AREL in predicting potential LDAs. The demo codes of the iLncDA-RSN are available online at https://github.com/CDMBlab/idenLD-AREL. Yahan Li, Junliang Shang, Feng Li 0033, Jin-Xing Liu 0001 |
BIBM | 3 |
| 2023 | scNMF-Impute: imputation for single-cell RNA-seq data based on nonnegative matrix factorizationabstractSingle-cell RNA sequencing (scRNA-seq) data are collected at an unheard-of rate thanks to the advancement of high-throughput sequencing technologies. However, due to the limitations of current technology, scRNA-seq is sometimes unable to capture the expressed genes, resulting in a large number of zero counts (also known as dropout events) in the data. These dropout events can cause data loss in the gene expression matrix and severely hampers the accuracy of downstream analysis. To address this problem, in this paper, we propose a new imputation method called scNMF-impute. The scNMF-impute method imputes the dropout events and performs dimensionality reduction under the framework of nonnegative matrix factorization (NMF). To effectively identify the location of the dropout and recover the value of the dropout, we explicitly model the dropout events as a matrix. Therefore, the gene expression matrix without dropout is represented as the sum of the original data matrix and the dropout matrix. In addition, to reduce the influence of dropout on factorization, we introduce the similarity information between genes into the NMF model. The introduction of gene similarity information can ensure the accurate recovery of data structures obscured by dropout events in the gene expression matrix. We conducted extensive experiments on simulated datasets and real scRNA-seq datasets to verify the effectiveness of scNMF-impute and other state-of-the-art methods. The results show that scNMF-impute can accurately calculate missing data and restore true gene expression, thus improving the accuracy of existing clustering methods and obtaining more accurate cell clustering results. Juan Wang 0003, Na-Na Zhang, Junliang Shang, Jin-Xing Liu 0001 |
BIBM | 3 |
| 2023 | Spectral clustering based on multi-similarity learning method for single-cell RNA-seq dataabstractThe inherent complexities of single-cell RNA-seq data (scRNA-seq), such as high dimensionality, low signal-to-noise ratio, cellular heterogeneity, and imbalanced distribution of subcellular types, pose significant challenges when conducting cell type analysis. To address these obstacles, employing appropriate data preprocessing techniques in the single-cell clustering process is crucial, and spectral clustering is particularly effective due to its robustness to noise and outliers. Therefore, this paper presents a novel spectral clustering algorithm based on multi-similarity learning method (MSSC). However, utilizing a pairwise strategy to assess the similarity between two data points in conventional spectral clustering results in an insufficient representation of the intricate relationships within the dataset. In light of this issue, the proposed algorithm employs two similarity measurement methods, namely Euclidean distance and Spearman rank correlation coefficient, to obtain similarity matrices. These matrices are then fused for use in spectral clustering. Additionally, prior to performing spectral clustering, the scRNA-seq data is preprocessed using the Sigmoid kernel similarity method and normalization techniques. As a consequence, our method yields a more extensive and intricate dataset similarity information, thereby enhancing the performance of spectral clustering. Finally, the Gaussian mixture model (GMM) is used for clustering. In most cases, experiments validated that the MSSC method outperforms the other four clustering methods on seven benchmark scRNA-seq datasets. Lianlian Zhang, Shaoyi Liu, Qianqian Ren, Junliang Shang, Feng Li 0033 |
BIBM | 4 |
| 2023 | Crow Search Algorithm Based on Information Interaction for Epistasis DetectionabstractIn the genome-wide association study, the interactions of single nucleotide polymorphisms (SNPs) play an important role in revealing the genetic mechanism of complex diseases, and such interaction is called epistasis or epistatic interactions. In recent years, swarm intelligence methods have been widely used to detect epistatic interactions because they can effectively deal with global optimization problems. In this study, we propose a crow search algorithm based on information interaction (FICSA) to detect epistatic interactions. FICSA combines particle swarm optimization (PSO) and crow search algorithm (CSA) to balance the exploration and exploitation in the search process, which can effectively improve the ability of the algorithm to detect epistatic interactions. In addition, opposition-based learning strategy and adaptive parameters are used to further improve the performance of the algorithm. We compare FICSA with seven other epistasis detection algorithms using both simulated datasets and a real-life age-related macular degeneration (AMD) dataset. The results on simulated datasets show that FICSA has better detection power, while the results on the real dataset demonstrate the effectiveness of the proposed algorithm. Junliang Shang, Yijun Gu, Qianqian Ren |
BIBM | 2 |
| 2023 | Spectral Clustering of Single-Cell RNA-Sequencing Data by Multiple Feature Sets Affinity
Feng Li 0033, Junliang Shang, Qianqian Ren, Shengjun Li |
ICIC (3) | 3 |
| 2023 | Epileptic Seizure Detection Based on Feature Extraction and CNN-BiGRU Network with Attention Mechanism
Jie Xu 0059, Juan Wang 0003, Jin-Xing Liu 0001, Junliang Shang, Ling-Yun Dai, Kuiting Yan, Shasha Yuan |
ICIC (2) | 4 |
| 2023 | Seizure Prediction Based on Hybrid Deep Learning Model Using Scalp Electroencephalogram
Kuiting Yan, Junliang Shang, Juan Wang 0003, Jie Xu 0059, Shasha Yuan |
ICIC (2) | 2 |
| 2023 | ABCAE: Artificial Bee Colony Algorithm with Adaptive Exploitation for Epistatic Interaction Detection
Qianqian Ren, Yahan Li, Feng Li 0033, Jin-Xing Liu 0001, Junliang Shang |
ISBRA | 5 |
| 2023 | DM-MOGA: a multi-objective optimization genetic algorithm for identifying disease modules of non-small cell lung cancerabstractBACKGROUND: Constructing molecular interaction networks from microarray data and then identifying disease module biomarkers can provide insight into the underlying pathogenic mechanisms of non-small cell lung cancer. A promising approach for identifying disease modules in the network is community detection. RESULTS: In order to identify disease modules from gene co-expression networks, a community detection method is proposed based on multi-objective optimization genetic algorithm with decomposition. The method is named DM-MOGA and possesses two highlights. First, the boundary correction strategy is designed for the modules obtained in the process of local module detection and pre-simplification. Second, during the evolution, we introduce Davies-Bouldin index and clustering coefficient as fitness functions which are improved and migrated to weighted networks. In order to identify modules that are more relevant to diseases, the above strategies are designed to consider the network topology of genes and the strength of connections with other genes at the same time. Experimental results of different gene expression datasets of non-small cell lung cancer demonstrate that the core modules obtained by DM-MOGA are more effective than those obtained by several other advanced module identification methods. CONCLUSIONS: The proposed method identifies disease-relevant modules by optimizing two novel fitness functions to simultaneously consider the local topology of each gene and its connection strength with other genes. The association of the identified core modules with lung cancer has been confirmed by pathway and gene ontology enrichment analysis. Junliang Shang, Xuhui Zhu, Feng Li 0033, Jin-Xing Liu 0001 |
BMC Bioinform. | 1 |
| 2023 | MSF-LRR: Multi-Similarity Information Fusion Through Low-Rank Representation to Predict Disease-Associated MicrobesabstractAn Increase in microbial activity is shown to be intimately connected with the pathogenesis of diseases. Considering the expense of traditional verification methods, researchers are working to develop high-efficiency methods for detecting potential disease-related microbes. In this article, a new prediction method, MSF-LRR, is established, which uses Low-Rank Representation (LRR) to perform multi-similarity information fusion to predict disease-related microbes. Considering that most existing methods only use one class of similarity, three classes of microbe and disease similarity are added. Then, LRR is used to obtain low-rank structural similarity information. Additionally, the method adaptively extracts the local low-rank structure of the data from a global perspective, to make the information used for the prediction more effective. Finally, a neighbor-based prediction method that utilizes the concept of collaborative filtering is applied to predict unknown microbe-disease pairs. As a result, the AUC value of MSF-LRR is superior to other existing algorithms under 5-fold cross-validation. Furthermore, in case studies, excluding originally known associations, 16 and 19 of the top 20 microbes associated with Bacterial Vaginosis and Irritable Bowel Syndrome, respectively, have been confirmed by the recent literature. In summary, MSF-LRR is a good predictor of potential microbe-disease associations and can contribute to drug discovery and biological research. Jin-Xing Liu 0001, Meng-Meng Yin, Ying-Lian Gao, Junliang Shang, Chun-Hou Zheng 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2023 | A Method Based On Dual-Network Information Fusion to Predict MiRNA-Disease AssociationsabstractMicroRNAs (miRNAs) are single-stranded small RNAs. An increasing number of studies have shown that miRNAs play a vital role in many important biological processes. However, some experimental methods to predict unknown miRNA-disease associations (MDAs) are time-consuming and costly. Only a small percentage of MDAs are verified by researchers. Therefore, there is a great need for high-speed and efficient methods to predict novel MDAs. In this paper, a new computational method based on Dual-Network Information Fusion (DNIF) is developed to predict potential MDAs. Specifically, on the one hand, two enhanced sub-models are integrated to reconstruct an effective prediction framework; on the other hand, the prediction performance of the algorithm is improved by fully fusing multiple omics data information, including validated miRNA-disease associations network, miRNA functional similarity, disease semantic similarity and Gaussian interaction profile (GIP) kernel network associations. As a result, DNIF achieves the excellent performance under situation of 5-fold cross validation (average AUC of 0.9571). In the cases study of three important human diseases, our model has achieved satisfactory performance in predicting potential miRNAs for certain diseases. The reliable experimental results demonstrate that DNIF could serve as an effective calculation method to accelerate the identification of MDAs. Feng Zhou 0021, Meng-Meng Yin, Jing-Xiu Zhao, Junliang Shang, Jin-Xing Liu 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2023 | A Personalized Low-Rank Subspace Clustering Method Based on Locality and Similarity Constraints for scRNA-seq Data AnalysisabstractSingle-cell RNA sequencing (scRNA-seq) technology can provide expression profile of single cells, which propels biological research into a new chapter. Clustering individual cells based on their transcriptome is a critical objective of scRNA-seq data analysis. However, the high-dimensional, sparse and noisy nature of scRNA-seq data pose a challenge to single-cell clustering. Therefore, it is urgent to develop a clustering method targeting scRNA-seq data characteristics. Due to its powerful subspace learning capability and robustness to noise, the subspace segmentation method based on low-rank representation (LRR) is broadly used in clustering researches and achieves satisfactory results. In view of this, we propose a personalized low-rank subspace clustering method, namely PLRLS, to learn more accurate subspace structures from both global and local perspectives. Specifically, we first introduce the local structure constraint to capture the local structure information of the data, while helping our method to obtain better inter-cluster separability and intra-cluster compactness. Then, in order to retain the important similarity information that is ignored by the LRR model, we utilize the fractional function to extract similarity information between cells, and introduce this information as the similarity constraint into the LRR framework. The fractional function is an efficient similarity measure designed for scRNA-seq data, which has theoretical and practical implications. In the end, based on the LRR matrix learned from PLRLS, we perform downstream analyses on real scRNA-seq datasets, including spectral clustering, visualization and marker gene identification. Comparative experiments show that the proposed method achieves superior clustering accuracy and robustness. Tian-Jing Qiao, Jin-Xing Liu 0001, Junliang Shang, Shasha Yuan, Chun-Hou Zheng 0001, Juan Wang 0003 |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | GCCN: Graph Capsule Convolutional Network for Progressive Mild Cognitive Impairment Prediction and Pathogenesis Identification Based on Imaging Genetic DataabstractIn this study, we proposed a novel method called the graph capsule convolutional network (GCCN) to predict the progression from mild cognitive impairment to dementia and identify its pathogenesis. First, we proposed a novel risk gene discovery component to indirectly target genes with higher interactions with others. These risk genes and brain regions were collected as nodes to construct heterogeneous pathogenic information association graphs. Second, the graph capsules were established by projecting heterogeneous pathogenic information into a set of disentangled latent components. The orientation and length of capsules are representations of the format and intensity of pathogenic information. Third, graph capsule convolution network was used to model the information flows among pathogenic factors and elaborates the convergence of primary capsules to advanced capsules. The advanced capsule is a concept that organizes pathogenic information based on its consistency, and the synergistic effects of advanced capsules directed the development of the disease. Finally, discriminative pathogenic information flows were captured by a straightforward built-in interpretation mechanism, i.e., the dynamic routing mechanism, and applied to the identification of pathogenesis. GCCN has been experimentally shown to be significantly advanced on public datasets. Further experiments have shown that the pathogenic factors identified by GCCN are evidential and closely related to progressive mild cognitive impairment. Junliang Shang, Qi Zou 0003, Qianqian Ren, Boxin Guan, Feng Li 0033, Jin-Xing Liu 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2023 | NLRRC: A Novel Clustering Method of Jointing Non-Negative LRR and Random Walk Graph Regularized NMF for Single-Cell Type IdentificationabstractThe development of single-cell RNA sequencing (scRNA-seq) technology has opened up a new perspective for us to study disease mechanisms at the single cell level. Cell clustering reveals the natural grouping of cells, which is a vital step in scRNA-seq data analysis. However, the high noise and dropout of single-cell data pose numerous challenges to cell clustering. In this study, we propose a novel matrix factorization method named NLRRC for single-cell type identification. NLRRC joins non-negative low-rank representation (LRR) and random walk graph regularized NMF (RWNMFC) to accurately reveal the natural grouping of cells. Specifically, we find the lowest rank representation of single-cell samples by non-negative LRR to reduce the difficulty of analyzing high-dimensional samples and capture the global information of the samples. Meanwhile, by using random walk graph regularization (RWGR) and NMF, RWNMFC captures manifold structure and cluster information before generating a cluster allocation matrix. The cluster assignment matrix contains cluster labels, which can be used directly to get the clustering results. The performance of NLRRC is validated on simulated and real single-cell datasets. The results of the experiments illustrate that NLRRC has a significant advantage in single-cell type identification. Juan Wang 0003, Linping Wang, Shasha Yuan, Feng Li 0033, Jin-Xing Liu 0001, Junliang Shang |
IEEE J. Biomed. Health Informatics | 6 |
| 2023 | Automatic Seizure Detection Using Logarithmic Euclidean-Gaussian Mixture Models (LE-GMMs) and Improved Deep Forest LearningabstractAutomatic seizure detection could facilitate early detection, improve treatment planning, and reduce medical workload. This study describes a novel Logarithmic Euclidean-Gaussian Mixture Models (LE-GMMs) and an improved Deep Forest learning algorithm for epileptic seizure detection. The LE-GMMs could map the Riemannian manifold structure of Gaussian models to linear Euclidean space, which fully exploits the ability of GMMs to distinguish non-seizure and seizure EEG signals. The Multi-Pooling and error Screening Forest (MPSForest) learning method based on Deep Forest uses multi-pooling and out-of-bagging (OOB) error screening to reduce memory load and random tree construction. Firstly, variational modal decomposition (VMD) is applied to decompose electroencephalogram (EEG) signals into five layers, and the first three layers are chosen to construct EEG time-frequency distribution. Then Gaussian Mixture Models are estimated, and the LE-GMMs are constructed to extract valid EEG features. These features are input into the MPSForest model to classify seizure and non-seizure samples. After that, the outputs are subjected to post-processing to get the final seizure detection results, including moving average filtering and the adaptive collar technique. The proposed method achieves average sensitivity of 98.22% and specificity of 98.99% on the UPenn and Mayo Clinic dataset, and for the long-term Freiburg EEG dataset with 21 patients, the sensitivity of 98.47% and specificity of 98.57% are yielded respectively with the false detection rate of 0.24/h. The experimental results show that this proposed method has excellent accuracy in distinguishing non-seizure and seizure EEG signals and holds great potential for clinical research and diagnostics. Shasha Yuan, Junliang Shang, Jin-Xing Liu 0001, Juan Wang 0003 |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | scSASSL: Self-attention semi-supervised learning with deep generative models to automatically identify cell typesabstractHigh-throughput single-cell sequencing has distinct advantages over previous bulk sequencing technologies. It provides an opportunity for researchers to study cell heterogeneity from the level of individual cells and explain biological relationships between individual cells from a higher resolution perspective. However, single-cell data are characterized by a large number of samples, high dimensionality, and sparseness, which pose a challenge to traditional methods. Therefore, we develop a semi-supervised deep generative model with a self-attention mechanism. The use of deep learning methods allows the denoising and dimensionality reduction of high-dimensional single-cell data nonlinear. We use a neural network with a self-attention mechanism for cell type prediction. This approach facilitates the neural network to extract cell-to-cell relationship features and enhances the model’s ability to extract features. The model can generate data. We apply this ability of imputation to single-cell datasets, thus solving the sparsity problem of single-cell datasets. We have conducted experiments on several simulated and real datasets, and the experimental results show that our proposed method largely outperforms other existing methods both in terms of identifying cell types and imputation on single-cell data. Our method is scalable because it can handle large-scale single-cell datasets of more than a million quantities. This method is promising in other fields as well. All source codes used in our experiments have been deposited at https://github.com/FengLi12/scSASSL. Hongyu Duan, Feng Li 0033, Xin Chu, Zhensheng Sun, Junliang Shang, Xikui Liu 0001, Yan Li 0041 |
BIBM | 5 |
| 2022 | Artificial bee colony algorithm based on self-adjusting random grouping for high-order epistasis detectionabstractIn the genome-wide association studies (GWAS), epistasis detection is of great significance to study the pathogenesis of complex diseases. Epistasis refers to the effect of interactions between multiple single nucleotide polymorphisms (SNPs) on complex diseases. In this paper, an artificial bee colony algorithm based on self-adjusting random grouping (ABC-SRG) is proposed for high-order epistasis detection. ABC-SRG adopts a new self-adjusting random grouping strategy, which realizes the division of the original data according to the fitness value of each grouping. In addition, a variance-based adaptive iteration strategy is proposed, which implements the adaptive iteration through the variance of the fitness value of each iteration of the algorithm. To demonstrate the effectiveness of the algorithm, the experiments on simulated data and real data were conducted. In the simulation experiments, ABC-SRG was compared with the other five methods for second-order and third-order SNP interaction detection. Age-related macular degeneration (AMD) data were selected for the real data experiment, and most of the SNP interactions detected in the experiment have been confirmed to be related to the AMD disease. Therefore, ABC-SRG is an effective method to detect high-order epistasis. Junliang Shang, Yijun Gu, Feng Li 0033, Jin-Xing Liu 0001, Boxin Guan |
BIBM | 1 |
| 2022 | Identification of cancer driver modules by combining network functional and topology informationabstractAccurate identification of cancer driver modules or pathways is important for controlling disease progression and timely treatment. In recent years, most approaches have been based on mutation data combined with gene interaction networks to identify cancer driver modules, but cancer-related genes tend to interact with each other, and the mutations they experience disruption their neighbors. Therefore, we propose a framework that combines network function and topological information to quantify the extent to which mutated genes disrupt their neighbors. Firstly, similarity in protein-protein interaction networks binds to high coverage and high mutual exclusivity of mutant genes, which are used to obtain the impact of the interaction between two mutant genes on biological function. Secondly, we quantified the degree of gene disruption by mutant genes in their neighborhood using an adaptive spread strength measure to obtain the gene spread strength network (GSSN). Finally, the module is extended using CFinder strategy to obtain the optimal driving module. We apply our method to 12 cancer datasets, and the experimental results show that our method outperforms the other three methods on most datasets. At the same time, we also analyze common and low-frequency driver modules in cancer. Xin Chu, Feng Li 0033, Hongyu Duan, Junliang Shang, Juan Wang 0003, Jin-Xing Liu 0001 |
BIBM | 4 |
| 2022 | Predicting LncRNA-Disease Associations Based on LncRNA-MiRNA-Disease Multilayer Association Network and Bipartite Network RecommendationabstractThe pathogenesis of many human diseases is unclear, but many studies have shown that lncRNAs are deeply involved in the development of diseases. However, the exploration of lncRNA-disease associations in the laboratory requires a lot of time and financial resources, and computational-based methods have obvious advantages and become a promising research direction. But few experiments consider the relationship between other biological factors and lncRNAs and diseases. In this paper, a novel lncRNA-disease association prediction method, MANBNR, is proposed. MiRNAs are introduced by MANBNR to construct a lncRNA-miRNA-disease multilayer association network (MAN). The main innovation of MANBNR is to mine potential lncRNA-disease association information based on miRNA information. For any lncRNA-disease pair, lncRNA-associated miRNAs and disease-associated miRNAs are sorted into two sets, respectively. The association of this lncRNA-disease pair is judged by comparing the number of miRNAs shared in the two sets. This solves the problem that the known lncRNA-disease association matrices are too sparse. Finally, the bipartite network recommendation (BNR) algorithm was used to accurately predict potential lncRNA-disease association. The performance of MANBNR is better than that of many advanced methods at present. Case studies of breast cancer and lung cancer further demonstrate that MANBNR is an effective and reliable method for LDAs prediction. Guozheng Zhang, Shu-Zhen Li, Xu-Ran Dou, Junliang Shang, Qianqian Ren, Ying-Lian Gao |
BIBM | 4 |
| 2022 | MHILDA: identifying disease-associated lncRNAs by extracting key features from integrated heterogeneous networksabstractPredicting disease-related long non-coding RNAs (lncRNAs) can help reveal the genetic mechanisms of complex diseases. Accurately identifying disease-associated lncRNAs is crucial for human diagnosis and therapeutics of complex diseases. However, most computational models ignore the noise of the data and the interference of redundant information. In this study, we build heterogeneous networks by integrating three different data sources of lncRNAs, miRNAs and diseases, and then propose an efficient computational model called MHILDA. MHILDA selects the most helpful features to train the model by Lasso's feature extraction. MHILDA is evaluated by five-fold cross-validation and performs well both on the benchmark dataset and on the independent test set. To further evaluate the performance of MHILDA, two types of case studies are implemented. The experimental results show that MHILDA can predict lncRNAs for unknown diseases. Junliang Shang, Tongdui Zhang, Qianqian Ren, Guozheng Zhang |
BIBM | 2 |
| 2022 | Tensor Robust PCA Based on Transformed Tensor Singular Value Decomposition for Cancer Genomic DataabstractThe mining and analysis of genomics data provides a new idea for exploring the pathogenesis of human disease. Since these data often have the features of small samples, high-dimensional, and high redundancy, the traditional matrix decomposition method cannot fully mine the spatial structure and multiple perspective information of cancer genomics data. Inspired by the recently proposed robust tensor completion method, a tensor robust PCA method (TTTD) was proposed based on U-product and transformed tensor singular value decomposition (t-SVD) to explore the integrated cancer genomics data in this paper. Specifically, the unitary transform matrix is employed to replace the discrete Fourier transform matrix in t-SVD, which contributes to recover a lower tubal rank tensor to a certain extent. Meanwhile, the $\mathrm{L}_{2,1}-$norm is employed to learn the sparse term, and the row sparse constraint generated by it can better detect the abnormal value of the real tensor. In addition, the alternating direction method of the multiplier algorithm is used to optimize the TTTD method. Experimental results on the three integrated cancer multi-omics datasets show that the TTTD method achieves the better performance. Sheng-Nan Zhang, Jin-Xing Liu 0001, Juan Wang 0003, Junliang Shang |
BIBM | 5 |
| 2022 | KSMDB: A classification method in imbalanced COVID dataset based on KmeansSMOTE and DeBERTabstract2022 is already the third year of the COVID-19 outbreak, and public opinion information about the outbreak has always been at the forefront of hot searches. The imbalance problem prevalent in many reviews of COVID-19 causes classification models to favor most categories in training and prediction process, resulting in low accuracy of small sample classification data generated by imbalanced data sets. Therefore, it is suggested here that the text classification model is based on the combination of the KMeansSMOTE method combined with DeBERT. First of all, during data processing, the KmeansSMOTE algorithm is utilized to oversample the imbalance of the COVID dataset, which increases the classification accuracy of the model. Besides, we put a stacked denoising bidirectional transformer encoder (DeBERT) to use, a more abstract and richer hidden feature vector is extracted by adding an embedded layer after the input tag, and the noise data is reconstructed to solve the noise problem in the process of raw data existence and oversampling. Furthermore, on the basis of model training, overfitting can be alleviated by adopting an early stopping strategy. A world of experiments using the COVID dataset demonstrates the effectiveness of the proposed method for solving simple imbalance and noise problems. With an overall accuracy of 87%, which improves the classification effect of minority samples and provides a new feasible method for the war of epidemic prevention. Hua-Hui Gao, Junliang Shang, Ling-Yun Dai |
BIBM | 3 |
| 2022 | Construction of Gene Network Based on Inter-tumor Heterogeneity for Tumor Type Identification
Zhensheng Sun, Junliang Shang, Hongyu Duan, Jin-Xing Liu 0001, Xikui Liu 0001, Yan Li 0041, Feng Li 0033 |
ICIC (2) | 2 |
| 2022 | A Network-Based Voting Method for Identification and Prioritization of Personalized Cancer Driver Genes
Feng Li 0033, Junliang Shang, Xikui Liu 0001, Yan Li 0041 |
ISBRA | 3 |
| 2022 | A Tensor Robust Model Based on Enhanced Tensor Nuclear Norm and Low-Rank Constraint for Multi-view Cancer Genomics Data
Shasha Yuan, Junliang Shang, Jin-Xing Liu 0001 |
ISBRA | 3 |
| 2022 | MLMVFE: A Machine Learning Approach Based on Muli-view Features Extraction for Drug-Disease Associations Prediction
Ying Wang 0143, Ying-Lian Gao, Juan Wang 0003, Junliang Shang, Jin-Xing Liu 0001 |
ISBRA | 4 |
| 2022 | ARGLRR: An Adjusted Random Walk Graph Regularization Sparse Low-Rank Representation Method for Single-Cell RNA-Sequencing Data Clustering
Zhen-Chang Wang, Jin-Xing Liu 0001, Junliang Shang, Ling-Yun Dai, Chun-Hou Zheng 0001, Juan Wang 0003 |
ISBRA | 3 |
| 2022 | TDCOSR: A Multimodality Fusion Framework for Association Analysis Between Genes and ROIs of Alzheimer's Disease
Qi Zou 0003, Feng Li 0033, Juan Wang 0003, Jin-Xing Liu 0001, Junliang Shang |
ISBRA | 6 |
| 2022 | Multi-similarity fusion-based label propagation for predicting microbes potentially associated with diseases
Meng-Meng Yin, Ying-Lian Gao, Junliang Shang, Chun-Hou Zheng 0001, Jin-Xing Liu 0001 |
Future Gener. Comput. Syst. | 3 |
| 2022 | Visualization and Analysis of Single Cell RNA-Seq Data by Maximizing Correntropy Based Non-Negative Low Rank RepresentationabstractThe exploration of single cell RNA-sequencing (scRNA-seq) technology generates a new perspective to analyze biological problems. One of the major applications of scRNA-seq data is to discover subtypes of cells by cell clustering. Nevertheless, it is challengeable for traditional methods to handle scRNA-seq data with high level of technical noise and notorious dropouts. To better analyze single cell data, a novel scRNA-seq data analysis model called Maximum correntropy criterion based Non-negative and Low Rank Representation (MccNLRR) is introduced. Specifically, the maximum correntropy criterion, as an effective loss function, is more robust to the high noise and large outliers existed in the data. Moreover, the low rank representation is proven to be a powerful tool for capturing the global and local structures of data. Therefore, some important information, such as the similarity of cells in the subspace, is also extracted by it. Then, an iterative algorithm on the basis of the half-quadratic optimization and alternating direction method is developed to settle the complex optimization problem. Before the experiment, we also analyze the convergence and robustness of MccNLRR. At last, the results of cell clustering, visualization analysis, and gene markers selection on scRNA-seq data reveal that MccNLRR method can distinguish cell subtypes accurately and robustly. Cui-Na Jiao, Jin-Xing Liu 0001, Juan Wang 0003, Junliang Shang, Chun-Hou Zheng 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | MKL-LP: Predicting Disease-Associated Microbes with Multiple-Similarity Kernel Learning-Based Label Propagation
Ying-Lian Gao, Meng-Meng Yin, Jin-Xing Liu 0001, Junliang Shang, Chun-Hou Zheng 0001 |
ISBRA | 4 |
| 2021 | Multiscale part mutual information for quantifying nonlinear direct associations in networksabstractMOTIVATION: For network-assisted analysis, which has become a popular method of data mining, network construction is a crucial task. Network construction relies on the accurate quantification of direct associations among variables. The existence of multiscale associations among variables presents several quantification challenges, especially when quantifying nonlinear direct interactions. RESULTS: In this study, the multiscale part mutual information (MPMI), based on part mutual information (PMI) and nonlinear partial association (NPA), was developed for effectively quantifying nonlinear direct associations among variables in networks with multiscale associations. First, we defined the MPMI in theory and derived its five important properties. Second, an experiment in a three-node network was carried out to numerically estimate its quantification ability under two cases of strong associations. Third, experiments of the MPMI and comparisons with the PMI, NPA and conditional mutual information were performed on simulated datasets and on datasets from DREAM challenge project. Finally, the MPMI was applied to real datasets of glioblastoma and lung adenocarcinoma to validate its effectiveness. Results showed that the MPMI is an effective alternative measure for quantifying nonlinear direct associations in networks, especially those with multiscale associations. AVAILABILITY AND IMPLEMENTATION: The source code of MPMI is available online at https://github.com/CDMB-lab/MPMI. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Junliang Shang, Feng Li 0033, Jin-Xing Liu 0001, Honghai Zhang |
Bioinform. | 1 |
| 2021 | The Automatic Detection of Seizure Based on Tensor Distance And Bayesian Linear Discriminant AnalysisabstractElectroencephalogram (EEG) plays an important role in recording brain activity to diagnose epilepsy. However, it is not only laborious, but also not very cost effective for medical experts to manually identify the features on EEG. Therefore, automatic seizure detection in accordance with the EEG recordings is significant for the diagnosis and treatment of epilepsy. Here, a new method for detecting seizures using tensor distance (TD) is proposed. First, the time-frequency characteristics of EEG signals are obtained by wavelet transformation, and the tensor representation of EEG signals is then obtained. Tucker decomposition is used to obtain the principal components of the EEG tensor. After, the distances between different categories of EEG tensors are calculated as the EEG features. Finally, the TD features are classified through the Bayesian Linear Discriminant Analysis (Bayesian LDA) classifier. The performance of this method is measured by the sensitivity, specificity, and recognition accuracy. Results indicate 95.12% sensitivity, 97.60% specificity, 97.60% recognition accuracy, and a false detection rate of 0.76 per hour in the invasive EEG dataset, which included 566.57[Formula: see text]h of EEG recording data from 21 patients. Taken together, the results show that TD has a good detection effect for seizure classification and that this method has high computational speed and great potential for real-time diagnosis. Delu Ma, Shasha Yuan, Junliang Shang, Jin-Xing Liu 0001, Ling-Yun Dai, Fangzhou Xu |
Int. J. Neural Syst. | 3 |
| 2021 | DSTPCA: Double-Sparse Constrained Tensor Principal Component Analysis Method for Feature SelectionabstractThe identification of differentially expressed genes plays an increasingly important role biologically. Therefore, the feature selection approach has attracted much attention in the field of bioinformatics. The most popular method of principal component analysis studies two-dimensional data without considering the spatial geometric structure of the data. The recently proposed tensor robust principal component analysis method performs sparse and low-rank decomposition on three-dimensional tensors and effectively preserves the spatial structure. Based on this approach, the$L_{2,1}$- norm regularization term is introduced into the DSTPCA (Double-Sparse Constrained Tensor Principal Component Analysis) method. The DSTPCA method removes the redundant noise by double sparse constraints on the objective function to obtain sufficiently sparse results. After the regularization norm is introduced into the model, the ADMM (alternating direction method of multipliers) algorithm is used to solve the optimal problem. In the experiment of feature selection, while the more redundant genes were filtered out, the more genes closely associated with disease were screened. Experimental results using different datasets indicate that our method outperforms other methods. Yue Hu 0017, Jin-Xing Liu 0001, Ying-Lian Gao, Junliang Shang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2021 | CNV_IFTV: An Isolation Forest and Total Variation-Based Detection of CNVs from Short-Read Sequencing DataabstractAccurate detection of copy number variations (CNVs) from short-read sequencing data is challenging due to the uneven distribution of reads and the unbalanced amplitudes of gains and losses. The direct use of read depths to measure CNVs tends to limit performance. Thus, robust computational approaches equipped with appropriate statistics are required to detect CNV regions and boundaries. This study proposes a new method called CNV_IFTV to address this need. CNV_IFTV assigns an anomaly score to each genome bin through a collection of isolation trees. The trees are trained based on isolation forest algorithm through conducting subsampling from measured read depths. With the anomaly scores, CNV_IFTV uses a total variation model to smooth adjacent bins, leading to a denoised score profile. Finally, a statistical model is established to test the denoised scores for calling CNVs. CNV_IFTV is tested on both simulated and real data in comparison to several peer methods. The results indicate that the proposed method outperforms the peer methods. CNV_IFTV is a reliable tool for detecting CNVs from short-read sequencing data even for low-level coverage and tumor purity. The detection results on tumor samples can aid to evaluate known cancer genes and to predict target drugs for disease diagnosis. Xiguo Yuan, Jianing Xi, Liying Yang 0001, Junliang Shang, Junbo Duan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2020 | Sparse Regularization Tensor Robust PCA Based on t-product and Its Application in Cancer Genomic DataabstractGenetic information becomes more and more important in the process of biological research. Gene analysis is an effective mean in biological research, especially the analysis of differentially expressed genes. Robust principal component analysis (RPCA) is an effective method to identify differentially expressed genes. But tensor robust principal component analysis (TRPCA) performs better than RPCA when processing multi-dimensional data. The traditional TRPCA method also has limitations in restoring low-rank sparse components. To further improve the accuracy of the TRPCA method in restoring low-rank components and sparse components, we propose a novel TRPCA method to obtain high-order correlations information of multi-dimensional data. It uses a new nuclear norm based on t-product operator to approximate the rank function. The L2,1-norm is used to improve the sparsity of tensors and reduce the negative effects caused by noises and outliers. At the same time, the introduction of L2,1-norm enhances the sparsity of error components, and improves the accuracy of low-rank component recovery. The low-rank sparse components are obtained by solving the convex problem of the new tensor nuclear norm. It can well preserve the spatial structure and make full use of complementary information to improve the clustering effect. Alternating direction method of multiplier (ADMM) is used to solve the optimization problem of this method. Experimental results on different cancer genomic datasets indicate that our method is superior to other methods. Hang-Jin Yang, Yu-Ying Zhao, Jin-Xing Liu 0001, Yuxia Lei, Junliang Shang, Xiang-Zhen Kong |
BIBM | 5 |
| 2020 | Automatic Seizure Prediction based on Modified Stockwell Transform and Tensor DecompositionabstractReliable epileptic seizure prediction is significantly important in improving the life of patients and enhancing the therapy effect. In this paper, a novel seizure prediction algorithm is proposed employing the tensor decomposition on long-term intracranial EEG recordings. The modified Stockwell transform (MST) is conducted on the segmented EEG signals to transform into two-dimensional instantaneous power spectra. Then, the third-order tensor representation of the multi-channel EEG signals are structured with the models of time, frequency and space. Tucker decomposition, one valid tensor decomposition method, is applied to obtain the principal components of the EEG tensors and the smaller core tensors after decomposition are extracted as features of interictal EEG and preictal EEG. After that, the classification of preictal and interictal data is achieved by feeding the features into Bayesian Linear Discriminant Analysis (BLDA) classifier. The evaluation of the proposed algorithm is carried out on the Freiburg EEG database and a sensitivity of 88.49% for the seizure occurrence period of 30 min, meanwhile, a sensitivity of 97.62% for the seizure occurrence period of 50 min are yielded with a false alarm rate of 0. 25/h. The results show that this algorithm based on tensor analysis has notable performance for seizure prediction. Shasha Yuan, Jin-Xing Liu 0001, Junliang Shang, Fangzhou Xu, Ling-Yun Dai |
BIBM | 3 |
| 2020 | IDSSIM: an lncRNA functional similarity calculation model based on an improved disease semantic similarity methodabstractBACKGROUND: It has been widely accepted that long non-coding RNAs (lncRNAs) play important roles in the development and progression of human diseases. Many association prediction models have been proposed for predicting lncRNA functions and identifying potential lncRNA-disease associations. Nevertheless, among them, little effort has been attempted to measure lncRNA functional similarity, which is an essential part of association prediction models. RESULTS: In this study, we presented an lncRNA functional similarity calculation model, IDSSIM for short, based on an improved disease semantic similarity method, highlight of which is the introduction of information content contribution factor into the semantic value calculation to take into account both the hierarchical structures of disease directed acyclic graphs and the disease specificities. IDSSIM and three state-of-the-art models, i.e., LNCSIM1, LNCSIM2, and ILNCSIM, were evaluated by applying their disease semantic similarity matrices and the lncRNA functional similarity matrices, as well as corresponding matrices of human lncRNA-disease associations coming from either lncRNADisease database or MNDR database, into an association prediction method WKNKN for lncRNA-disease association prediction. In addition, case studies of breast cancer and adenocarcinoma were also performed to validate the effectiveness of IDSSIM. CONCLUSIONS: Results demonstrated that in terms of ROC curves and AUC values, IDSSIM is superior to compared models, and can improve accuracy of disease semantic similarity effectively, leading to increase the association prediction ability of the IDSSIM-WKNKN model; in terms of case studies, most of potential disease-associated lncRNAs predicted by IDSSIM can be confirmed by databases and literatures, implying that IDSSIM can serve as a promising tool for predicting lncRNA functions, identifying potential lncRNA-disease associations, and pre-screening candidate lncRNAs to perform biological experiments. The IDSSIM code, all experimental data and prediction results are available online at https://github.com/CDMB-lab/IDSSIM . Wenwen Fan, Junliang Shang, Feng Li 0033, Shasha Yuan, Jin-Xing Liu 0001 |
BMC Bioinform. | 2 |
| 2020 | Correntropy induced loss based sparse robust graph regularized extreme learning machine for cancer classificationabstractAbstract Background As a machine learning method with high performance and excellent generalization ability, extreme learning machine (ELM) is gaining popularity in various studies. Various ELM-based methods for different fields have been proposed. However, the robustness to noise and outliers is always the main problem affecting the performance of ELM. Results In this paper, an integrated method named correntropy induced loss based sparse robust graph regularized extreme learning machine (CSRGELM) is proposed. The introduction of correntropy induced loss improves the robustness of ELM and weakens the negative effects of noise and outliers. By using the L2,1-norm to constrain the output weight matrix, we tend to obtain a sparse output weight matrix to construct a simpler single hidden layer feedforward neural network model. By introducing the graph regularization to preserve the local structural information of the data, the classification performance of the new method is further improved. Besides, we design an iterative optimization method based on the idea of half quadratic optimization to solve the non-convex problem of CSRGELM. Conclusions The classification results on the benchmark dataset show that CSRGELM can obtain better classification results compared with other methods. More importantly, we also apply the new method to the classification problems of cancer samples and get a good classification effect. Liangrui Ren, Ying-Lian Gao, Jin-Xing Liu 0001, Junliang Shang, Chun-Hou Zheng 0001 |
BMC Bioinform. | 4 |
| 2020 | Introducing Heuristic Information Into Ant Colony Optimization Algorithm for Identifying EpistasisabstractEpistasis learning, which is aimed at detecting associations between multiple Single Nucleotide Polymorphisms (SNPs) and complex diseases, has gained increasing attention in genome wide association studies. Although much work has been done on mapping the SNPs underlying complex diseases, there is still difficulty in detecting epistatic interactions due to the lack of heuristic information to expedite the search process. In this study, a method EACO is proposed to detect epistatic interactions based on the ant colony optimization (ACO) algorithm, the highlights of which are the introduced heuristic information, fitness function, and a candidate solutions filtration strategy. The heuristic information multi-SURF* is introduced into EACO for identifying epistasis, which is incorporated into ant-decision rules to guide the search with linear time. Two functionally complementary fitness functions, mutual information and the Gini index, are combined to effectively evaluate the associations between SNP combinations and the phenotype. Furthermore, a strategy for candidate solutions filtration is provided to adaptively retain all optimal solutions which yields a more accurate way for epistasis searching. Experiments of EACO, as well as three ACO based methods (AntEpiSeeker, MACOED, and epiACO) and four commonly used methods (BOOST, SNPRuler, TEAM, and epiMODE) are performed on both simulation data sets and a real data set of age-related macular degeneration. Results indicate that EACO is promising in identifying epistasis. Yingxia Sun, Junliang Shang, Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Xiujuan Lei |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2019 | The computational prediction of drug-disease interactions using the dual-network L2,1-CMF methodabstractBACKGROUND: Predicting drug-disease interactions (DDIs) is time-consuming and expensive. Improving the accuracy of prediction results is necessary, and it is crucial to develop a novel computing technology to predict new DDIs. The existing methods mostly use the construction of heterogeneous networks to predict new DDIs. However, the number of known interacting drug-disease pairs is small, so there will be many errors in this heterogeneous network that will interfere with the final results. RESULTS: -norm are introduced in our method to achieve better results than other advanced methods. The network similarities of drugs and diseases with their chemical and semantic similarities are combined in this method. CONCLUSIONS: Cross validation is used to evaluate our method, and simulation experiments are used to predict new interactions using two different datasets. Finally, our prediction accuracy is better than other existing methods. This proves that our method is feasible and effective. Ying-Lian Gao, Jin-Xing Liu 0001, Juan Wang 0003, Junliang Shang, Ling-Yun Dai |
BMC Bioinform. | 5 |
| 2019 | PCA via joint graph Laplacian and sparse constraint: Identification of differentially expressed genes and sample clustering on gene expression dataabstractBACKGROUND: In recent years, identification of differentially expressed genes and sample clustering have become hot topics in bioinformatics. Principal Component Analysis (PCA) is a widely used method in gene expression data. However, it has two limitations: first, the geometric structure hidden in data, e.g., pair-wise distance between data points, have not been explored. This information can facilitate sample clustering; second, the Principal Components (PCs) determined by PCA are dense, leading to hard interpretation. However, only a few of genes are related to the cancer. It is of great significance for the early diagnosis and treatment of cancer to identify a handful of the differentially expressed genes and find new cancer biomarkers. RESULTS: In this study, a new method gLSPCA is proposed to integrate both graph Laplacian and sparse constraint into PCA. gLSPCA on the one hand improves the clustering accuracy by exploring the internal geometric structure of the data, on the other hand identifies differentially expressed genes by imposing a sparsity constraint on the PCs. CONCLUSIONS: Experiments of gLSPCA and its comparison with existing methods, including Z-SPCA, GPower, PathSPCA, SPCArt, gLPCA, are performed on real datasets of both pancreatic cancer (PAAD) and head & neck squamous carcinoma (HNSC). The results demonstrate that gLSPCA is effective in identifying differentially expressed genes and sample clustering. In addition, the applications of gLSPCA on these datasets provide several new clues for the exploration of causative factors of PAAD and HNSC. Chun-Mei Feng 0001, Yong Xu 0001, Mi-Xiao Hou, Ling-Yun Dai, Junliang Shang |
BMC Bioinform. | 5 |
| 2018 | Hypergraph regularized NMF by L2, 1-norm for Clustering and Com-abnormal Expression Genes Selection
Na Yu 0004, Ying-Lian Gao, Jin-Xing Liu 0001, Juan Wang 0003, Junliang Shang |
BIBM | 5 |
| 2018 | Performance Analysis of Non-negative Matrix Factorization Methods on TCGA Data
Mi-Xiao Hou, Jin-Xing Liu 0001, Junliang Shang, Ying-Lian Gao, Ling-Yun Dai |
ICIC (2) | 3 |
| 2018 | acsFSDPC: A Density-Based Automatic Clustering Algorithm with an Adaptive Cuckoo Search
Junliang Shang, Xuhui Zhu, Jin-Xing Liu 0001, Chun-Hou Zheng 0001 |
ICIC (2) | 2 |
| 2018 | An Improved Particle Swarm Optimization with Dynamic Scale-Free Network for Detecting Multi-omics Features
Shengjun Li, Junliang Shang, Jin-Xing Liu 0001, Chun-Hou Zheng 0001 |
ISBRA | 3 |
| 2017 | Robust graph regularized sparse orthogonal nonnegative matrix factorization for identifying differentially expressed genesabstractWith the advent of sequencing technology, numerous gene expression data are generated. Identifying differentially expressed genes play an important role in the gene therapy of cancer patients. As an useful mathematical tool, nonnegative matrix factorization (NMF) has been successfully used for identifying differentially expressed genes. In this paper, a novel method named robust graph regularized sparse orthogonal nonnegative matrix factorization (RGSON) is proposed and used for identifying differentially expressed genes, which introduces manifold learning, L1and orthogonal constraints into the objective function. In particular, L2,1-norm minimization is enforced on the objective function to improve the robustness of the algorithm. To prove the validity of the algorithm, experiments on the real genomic dataset are conducted. The results show that RGSON performs more effective than many other methods for identifying differentially expressed genes. Ling-Yun Dai, Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Junliang Shang, Chun-Mei Feng 0001, Yaxuan Wang |
BIBM | 4 |
| 2017 | A convex multi-view low-rank sparse regression for feature selection and clusteringabstractMany real-world problems involve multi-view high-dimension-small-sample-size data analysis, such as multi-omics data. The combination of multi-view databases is supposed to provide a better biological significance. However, the multi-view data always contain noise and outlying entries that result in inaccurate and unreliable. It has become an urgent need how to effectively analyze these data. We proposed a novel convex multi-view low-rank sparse regression (CMLSR) algorithm to do cluster and feature selection. The model was constructed by imposing L2,1-norm and trace norm constraints on the regularization functions. It can diminish the impact of noises and outliers and produce more precise results. Clustering quality was determined by both sparse constraint and low-rank constraint. Finally, we selected characteristic genes based on the projection matrix. The method was used in TCGA multi-view genes expression data sets, annotated according to Gene Ontology (GO). In this paper, we demonstrated the effectiveness of the proposed algorithm through comparing it with the existing methods. Yao Lu 0008, Jin-Xing Liu 0001, Junliang Shang |
BIBM | 4 |
| 2017 | A joint-L2, 1-norm-constraint-based semi-supervised feature extraction for RNA-Seq data analysis
Jin-Xing Liu 0001, Dong Wang 0019, Ying-Lian Gao, Chun-Hou Zheng 0001, Junliang Shang, Feng Liu 0013, Yong Xu 0001 |
Neurocomputing | 5 |
| 2016 | Sparse singular value decomposition-based feature extraction for identifying differentially expressed genesabstractRecently, feature extraction and dimensionality reduction have become fundamental tools for many data mining tasks, especially for processing high-dimensional data such as genome data. In this paper, a new feature extraction method based on sparse singular value decomposition (SSVD) is developed. SSVD algorithm is applied to extract differentially expressed genes from two different genome datasets that are all from The Cancer Genome Atlas (TCGA), and then the extracted genes are evaluated by the tools based on Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis. As a gene extraction method, SSVD is also compared with some existing feature extraction methods such as independent component analysis, the p-norm robust feature extraction and sparse principal component analysis. The experimental GO analysis results show that SSVD method outperforms the competitive algorithms. The KEGG analysis results demonstrate the genes which participate in the pathways in cancer. The elaborate experiments prove that SSVD is an effective feature selection method compared with the competitive methods. The KEGG analysis results may provide a meaningful reference to carry out further study for professionals in the field of biomedical science. Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Junliang Shang |
BIBM | 4 |
| 2016 | L21-iPaD: An efficient method for drug-pathway association pairs inferenceabstractPathway-based drug discovery overcomes the disadvantages of the “one drug-one target” method, which aims to find the effective drugs to act on single targets. The current method “iPaD” identities the drug-pathway association pairs by taking the lasso-type penalty on the drug-pathway association matrix. In order to enhance the robustness of the methods and be more effective to find the novel drug-pathway association pairs, we introduce a new method named “L2,1-iPaD”. Compared with the iPaD method, we impose the L2,1-norm constraint on the drug-pathway association coefficient matrix. By applying our method to a real widely datasets (CCLE dataset), we demonstrate that our method is superior to the iPaD method. And our method can obtain the smaller P-values than the iPaD method by performing permutation test to assess the significance of the identified drug-pathway association pairs. More importantly, compared with the iPaD method, our method can identify larger numbers of validated drug-pathway association pairs. The experimental results on the real dataset demonstrate the effectiveness of our method. Dong-Qin Wang, Chun-Hou Zheng 0001, Ying-Lian Gao, Jin-Xing Liu 0001, Shasha Wu, Junliang Shang |
BIBM | 6 |
| 2016 | Comparison of Non-negative Matrix Factorization Methods for Clustering Genomic Data
Mi-Xiao Hou, Ying-Lian Gao, Jin-Xing Liu 0001, Junliang Shang, Chun-Hou Zheng 0001 |
ICIC (2) | 4 |
| 2016 | Gene Extraction Based on Sparse Singular Value Decomposition
Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Junliang Shang |
ICIC (1) | 4 |
| 2016 | A Compressed Sensing Based Feature Extraction Method for Identifying Characteristic Genes
Shengjun Li, Junliang Shang, Jin-Xing Liu 0001 |
ICIC (2) | 2 |
| 2016 | An Improved Ant Colony Optimization Algorithm for the Detection of SNP-SNP Interactions
Yingxia Sun, Junliang Shang, Jin-Xing Liu 0001, Shengjun Li |
ICIC (3) | 2 |
| 2016 | SIPSO: Selectively Informed Particle Swarm Optimization Based on Mutual Information to Determine SNP-SNP Interactions
Wenxiang Zhang, Junliang Shang, Yingxia Sun, Jin-Xing Liu 0001 |
ICIC (1) | 2 |
| 2016 | CINOEDV: a co-information based method for detecting and visualizing n-order epistatic interactionsabstractBACKGROUND: Detecting and visualizing nonlinear interaction effects of single nucleotide polymorphisms (SNPs) or epistatic interactions are important topics in bioinformatics since they play an important role in unraveling the mystery of "missing heritability". However, related studies are almost limited to pairwise epistatic interactions due to their methodological and computational challenges. RESULTS: We develop CINOEDV (Co-Information based N-Order Epistasis Detector and Visualizer) for the detection and visualization of epistatic interactions of their orders from 1 to n (n ≥ 2). CINOEDV is composed of two stages, namely, detecting stage and visualizing stage. In detecting stage, co-information based measures are employed to quantify association effects of n-order SNP combinations to the phenotype, and two types of search strategies are introduced to identify n-order epistatic interactions: an exhaustive search and a particle swarm optimization based search. In visualizing stage, all detected n-order epistatic interactions are used to construct a hypergraph, where a real vertex represents the main effect of a SNP and a virtual vertex denotes the interaction effect of an n-order epistatic interaction. By deeply analyzing the constructed hypergraph, some hidden clues for better understanding the underlying genetic architecture of complex diseases could be revealed. CONCLUSIONS: Experiments of CINOEDV and its comparison with existing state-of-the-art methods are performed on both simulation data sets and a real data set of age-related macular degeneration. Results demonstrate that CINOEDV is promising in detecting and visualizing n-order epistatic interactions. CINOEDV is implemented in R and is freely available from R CRAN: http://cran.r-project.org and https://sourceforge.net/projects/cinoedv/files/ . Junliang Shang, Yingxia Sun, Jin-Xing Liu 0001, Junfeng Xia, Chun-Hou Zheng 0001 |
BMC Bioinform. | 1 |
| 2015 | Semi-supervised Feature Extraction for RNA-Seq Data Analysis
Jin-Xing Liu 0001, Yong Xu 0001, Ying-Lian Gao, Dong Wang 0019, Chun-Hou Zheng 0001, Junliang Shang |
ICIC (3) | 6 |
| 2015 | Hypergraph Supervised Search for Inferring Multiple Epistatic Interactions with Different Orders
Junliang Shang, Shengjun Li, Jin-Xing Liu 0001, Yuanke Zhang |
ICIC (2) | 1 |
| 2011 | Performance analysis of novel methods for detecting epistasisabstractBACKGROUND: Epistasis is recognized fundamentally important for understanding the mechanism of disease-causing genetic variation. Though many novel methods for detecting epistasis have been proposed, few studies focus on their comparison. Undertaking a comprehensive comparison study is an urgent task and a pathway of the methods to real applications. RESULTS: This paper aims at a comparison study of epistasis detection methods through applying related software packages on datasets. For this purpose, we categorize methods according to their search strategies, and select five representative methods (TEAM, BOOST, SNPRuler, AntEpiSeeker and epiMODE) originating from different underlying techniques for comparison. The methods are tested on simulated datasets with different size, various epistasis models, and with/without noise. The types of noise include missing data, genotyping error and phenocopy. Performance is evaluated by detection power (three forms are introduced), robustness, sensitivity and computational complexity. CONCLUSIONS: None of selected methods is perfect in all scenarios and each has its own merits and limitations. In terms of detection power, AntEpiSeeker performs best on detecting epistasis displaying marginal effects (eME) and BOOST performs best on identifying epistasis displaying no marginal effects (eNME). In terms of robustness, AntEpiSeeker is robust to all types of noise on eME models, BOOST is robust to genotyping error and phenocopy on eNME models, and SNPRuler is robust to phenocopy on eME models and missing data on eNME models. In terms of sensitivity, AntEpiSeeker is the winner on eME models and both SNPRuler and BOOST perform well on eNME models. In terms of computational complexity, BOOST is the fastest among the methods. In terms of overall performance, AntEpiSeeker and BOOST are recommended as the efficient and effective methods. This comparison study may provide guidelines for applying the methods and further clues for epistasis detection. Junliang Shang, Daojun Ye, Yaling Yin |
BMC Bioinform. | 1 |