EDBT 2026 Demo / reviewers in the wild / expert
Cheng Liang 0001
dblp:81/9078-1
· DBLP profile ↗
54ranked-venue papers
9as first author
39since 2021 · last 2027
0000-0003-3832-0969ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 29 · 4 first-author · 18 since 2021Artificial intelligence and machine learning · 19 · 3 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | K2A: A knowledge-to-action paradigm for breast cancer analysis via orchestrated pharmacokinetic and clinical semantics
Tianyu Liu 0006, Feiyan Feng, Cheng Liang 0001, Hong Wang 0015, Yanshen Sun |
Expert Syst. Appl. | 4 |
| 2026 | PharmaQA: Prompt-Based Molecular Representation Learning via Pharmacophore-Oriented Question AnsweringabstractMolecular representation plays a central role in computational drug discovery. Pharmacophores, functional groups responsible for molecular bioactivity, have been widely studied in cheminformatics. However, their incorporation into molecular representation learning, particularly in a context reasoning or generalization, remains relatively limited. To address this gap, we propose PharmaQA, a pharmacophore oriented question answering framework that formulates tailored prompts to extract context-aware molecular semantics. Rather than encoding pharmacophore features, PharmaQA learns to answer pharmacophore related queries. This design enables flexible reasoning across diverse tasks, including molecular property prediction, compound-target interaction prediction, and binding affinity estimation. Experimental results on benchmark datasets demonstrate that PharmaQA achieves competitive performance. In a ligand discovery case study using FDA-approved compounds, the framework identified potential inhibitors for three therapeutic targets, with strong docking performance. As a generalizable and modular solution, PharmaQA incorporates pharmacophoric knowledge into molecular embeddings, enhancing both predictive accuracy and interpretability in drug discovery applications. Chengwei Ai, Qiaozhen Meng, Mengwei Sun, Ruihan Dong, Hongpeng Yang, Shiqiang Ma, Cheng Liang 0001, Fei Guo 0001 |
AAAI | 8 |
| 2026 | Geometry-Aware Variational Information Maximization for Deep Incomplete Multi-view ClusteringabstractIncomplete multi-view clustering (IMVC) aims to group data into meaningful clusters when each sample is only partially observed across multiple views. Most existing methods either rely on imputation strategies that may introduce noise and distort the underlying data distribution, or adopt cross-view alignment techniques that focus on pairwise relationships, often resulting in suboptimal representations and unstable clustering performance. In this paper, we propose Geometry-Aware Variational Information Maximization for Deep Incomplete Multi-view Clustering (GAVIM), a novel imputation-free variational framework that enables robust and coherent incomplete multi-view clustering. Specifically, GAVIM leverages mutual information maximization to preserve the high mutual information between the available multi-view data and the shared embedding. Moreover, we explicitly retain local geometric consistency within each view-specific latent space under the guidance of an adaptive global supervision signal. Lastly, GAVIM aligns all views simultaneously using a Gramian representation alignment measure, ensuring coherent structure across modalities and promoting unified, semantically meaningful representations. Extensive experiments on five benchmark IMVC datasets with varying levels of view incompleteness demonstrate that GAVIM consistently outperforms state-of-the-art methods in clustering accuracy and representation quality. Wenlan Chen, Daoyuan Wang, Fei Guo 0001, Cheng Liang 0001 |
AAAI | 5 |
| 2026 | High-order correlation and consistency-aware multi-view clustering via anchor graph learning
Cheng Liang 0001, Wenchao Zang, Daoyuan Wang, Fei Guo 0001 |
Neural Networks | 1 |
| 2026 | Graph-Embedded Deep Generative Clustering for Single-Cell Multi-Omics Data IntegrationabstractThe advancement of sequencing technologies has generated an unprecedented volume of single-cell multi-omics data, providing new opportunities for biological discovery and medical research. However, due to the high heterogeneity across different omics types, effective integration of single-cell multi-omics data remains a critical challenge. Existing methods generally ignore the graph structure information among cells or resort to additional knowledge to construct the cell graphs, leading to suboptimal performance and potentially limited practical utility. In this study, we propose a novel Graph-embedded Deep Generative Clustering model (GeDGC) for single-cell multi-omics data integration. Specifically, GeDGC simultaneously learns the shared latent representations and cluster factors across multiple omics by leveraging Gaussian mixture models. Moreover, we impose the graph embedding constraint on both the latent representations and the cluster assignments to ensure the preservation of intrinsic local data structure among cells. As a result, our model captures complex correlations across omics and obtains informative shared latent embeddings for downstream tasks. Extensive experimental results with seventeen competing methods on ten datasets confirm the superiority of GeDGC in single-cell multi-omics data integration. Cheng Liang 0001, Wenlan Chen, Chang-Dong Wang 0001, Shichao Zhang 0001, Fei Guo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | MultiPert: An adversarial alignment and dual attention framework for single-cell multi-omics perturbation predictionabstractPrecise prediction of perturbation responses is essential in systems biology research, as it plays a pivotal role in characterizing cellular identities and elucidating the regulatory mechanisms of biological pathways. Existing perturbation-responses prediction approaches are predominantly confined to single-modality transcriptomic data, limiting their capacity to capture cross-layer molecular effects. Here, we present MultiPert, a deep learning framework specifically designed for predicting perturbation responses in single-cell multi-omics data. MultiPert employs modality-specific encoders with dedicated pretraining, integrates perturbation through a dual-attention mechanism, and achieves cross-modal alignment via adversarial training. Benchmarking on human THP-1 and kidney multi-omics datasets demonstrates that MultiPert reliably predicts both perturbed gene expression and protein abundance profiles, achieving superior accuracy and stability compared to state-of-the-art strategies. MultiPert generalizes to unseen perturbations and uncovers regulatory mechanisms of immune checkpoint molecules based on perturbed proteomic predictions. In addition, enrichment analyses of perturbed transcriptomic predictions reveal immune-related pathways. By providing an integrated and interpretable framework, MultiPert expands the scope of perturbation modeling at the multi-omics level, thereby offering a robust methodological foundation for comprehensive research into pathogenesis and drug discovery. Xinyue Tang, Jiawei Li 0018, Cheng Liang 0001, Jijun Tang, Fei Guo 0001 |
PLoS Comput. Biol. | 4 |
| 2026 | Structured knowledge-inspired two-stage knowledge alignment framework for Alzheimer's disease diagnosis
Fulin Zheng, Hong Wang 0015, Tianyu Liu 0006, Feiyan Feng, Cheng Liang 0001 |
Pattern Recognit. | 6 |
| 2026 | Incomplete Multi-View Clustering via Robust Representation Learning and Tensor-Based Co-RegularizationabstractAs incompleteness is common in real-world data, incomplete multi-view clustering is of great significance in the unsupervised learning field because it allows the partitioning of multi-view data with missing information into distinct groups. In this paper, we propose a novel generalized framework for incomplete multi-view clustering based on robust representation learning and tensor-based co-regularization (RRLTCR). Specifically, a robust principal component analysis is first used to learn a robust representation for each view. To explore high-order relationships among views, the view-specific spectral embeddings are stacked into a third-order tensor with a Schattenp-norm constraint. By spreading the complementary information of the high-quality available data from each view on a global scale, our model is able to alleviate the adverse effects of data noise and uncover the underlying common cluster structure. An effective iterative optimization strategy is developed to efficiently solve our model. According to the experimental results on seven datasets, our proposed framework has the potential to improve the clustering performance for a variety of incomplete multi-view clustering problems. Our research work brings a generalized framework for incomplete multi-view clustering, which can also assist in exploring the large cohort of existing incomplete multimodality datasets for other downstream tasks. Cheng Liang 0001, Daoyuan Wang, Fei Guo 0001, Shichao Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Boosting the No-Reference Image Quality Assessment via Low-Quality Pseudo ReferencesabstractNo-reference image quality assessment (NR-IQA) aims to predict perceptual image quality without access to pristine references, which remains challenging due to diverse and complex distortions. Recent pseudo-reference-based methods attempt to mitigate this challenge but often rely on highfidelity pseudo-reference reconstruction. In contrast, this work shows that improving NR-IQA performance does not depend on reconstruction quality, but on effective representation learning, feature alignment, and deviation modeling between distorted images and pseudo references. To this end, we propose a novel NR-IQA framework that leverages low-quality pseudo references generated by a masked autoencoder with a lightweight decoder. Rather than pursuing detailed reconstruction, the pseudo reference is used to facilitate representation-level deviation modeling in a shared latent space via a cross-attention-based mechanism. Extensive experiments on multiple benchmark datasets demonstrate that the proposed method consistently outperforms state-of-the-art NR-IQA approaches while maintaining modest computational complexity. Our source code will be available at: https://github.com/jianjin008/L-IQA. Lili Meng, Yingnan Wang, Miaohui Wang, Guosheng Lin, Cheng Liang 0001, Jiande Sun 0001, Huaxiang Zhang 0001, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Prototype-Calibrated Multimodal Fusion with Relational Consistency for Spatial Domain Identification in Spatial TranscriptomicsabstractSpatial transcriptomics technologies produce rich multimodal data, including gene expression profiles, spatial coordinates, and histological images, offering unprecedented opportunities to characterize tissue organization. However, effectively integrating these heterogeneous modalities remains challenging due to the complex molecular and morphological interplay. To address this, we propose PCMF-RC, a novel Prototype-Calibrated Multimodal Fusion framework with Relational Consistency, designed for accurate spatial domain identification. Our approach first partitions high-resolution histopathological images into patches centered on spatial transcriptomics spots and extracts spatially aware visual features using a pre-trained visual state space model. We then obtain modality-specific latent embeddings where we separately encode gene expression and image features via graph convolutional networks. To improve semantic consistency, a cross-view prototype matching module that aligns cluster prototypes across modalities is introduced to mitigate prototype shift. We further design a relational consistency contrastive learning module to effectively enforce alignment of spatial relational structures, reducing distributional discrepancies and enhancing robustness. Extensive experiments on multiple spatial transcriptomics datasets demonstrate that PCMF-RC outperforms existing methods in spatial domain delineation, offering a robust and generalizable solution for multimodal spatial omics integration. Daoyuan Wang, Guanghui Li 0003, Cheng Liang 0001 |
BIBM | 4 |
| 2025 | Image-Enhanced Hybrid Encoding with Reinforced Contrastive Learning for Spatial Domain Identification in Spatial TranscriptomicsabstractSpatial transcriptomics integrates spatial, gene expression, and multichannel immunohistochemistry image data, enabling advanced insights into cellular organization. However, existing methods often struggle to effectively fuse these multimodal data, limiting their potential for accurate spatial domain identification. Here, we propose IE-HERCL (Image-Enhanced Hybrid Encoding with Reinforced Contrastive Learning), a novel framework designed to address this challenge. Specifically, IE-HERCL employs hybrid encoding to capture both the non-spatial features and spatial dependencies for both gene and image modalities via autoencoders and GraphSAGE, respectively. These features are then fused using cross-view attention mechanisms to generate the unified informative embedding. To enhance the representation learning capability, we introduce a reinforced contrastive learning strategy to mitigate the influences of false negative samples, where we detect potential positive counterparts with high-order random walks. In addition, the cluster alignment is dynamically refined through optimal transport, which ensures that the fused consensus representation is coherent and robust, enabling accurate spatial domain identification. Our approach achieves state-of-the-art performance on five image-enhanced spatial transcriptomics datasets, demonstrating its robustness and effectiveness in multimodal integration and spatial domain identification. IE-HERCL offers a powerful and innovative solution for advancing spatial transcriptomics analysis. The code is released on https://github.com/wdyi701/IE-HERCL. Daoyuan Wang, Wenlan Chen, Cheng Liang 0001, Fei Guo 0001 |
IJCAI | 4 |
| 2025 | Deep Variational Incomplete Multi-View Clustering with Information-Theoretic Guidance
Wenlan Chen, Cheng Liang 0001, Fei Guo 0001 |
ACM Multimedia | 3 |
| 2025 | Dual-Level Distribution Alignment for Deep Incomplete Multi-View ClusteringabstractIncomplete Multi-view Clustering (IMvC) aims to perform effective clustering in the presence of missing views by exploiting the available information. While many existing approaches demonstrate satisfactory performance, their failure to adequately optimize the recovered data often limits the quality of learned representations and thus hampers clustering performance. To address this challenge, we propose a novel method, Dual-Level Distribution Alignment for Deep Incomplete Multi-View Clustering (DDAIMVC). To effectively address missing data, DDAIMVC employs a fusion-fill strategy to recover incomplete views. The recovered data from each view are then concatenated and processed through an attention mechanism to generate a unified high-level representation. To ensure consistent information across views, the framework performs distribution alignment at both the instance and cluster levels. Specifically, instance-level distribution alignment is conducted by minimizing the maximum mean discrepancy among views, while cluster-level distribution alignment is enhanced via prototypical contrastive learning, which encourages coherent cluster assignments across different modalities. Through the co-optimization of dual-level distribution alignment, the common representation reveals a clear clustering structure. Experimental results on benchmark multi-view datasets demonstrate that DDAIMVC consistently achieves state-of-the-art clustering performance. Fujian Ren, Wenlan Chen, Fei Guo 0001, Cheng Liang 0001 |
ACM Multimedia | 5 |
| 2025 | Disentangled Cross-Modal Representation Learning with Enhanced Mutual SupervisionabstractCross-modal representation learning aims to extract semantically aligned representations from heterogeneous modalities such as images and text. Existing multimodal VAE-based models often suffer from limited capability to align heterogeneous modalities or lack sufficient structural constraints to clearly separate the modality-specific and shared factors. In this work, we propose a novel framework, termed **D**isentangled **C**ross-**M**odal Representation Learning with **E**nhanced **M**utual Supervision (DCMEM). Specifically, our model disentangles the common and distinct information across modalities and regularizes the shared representation learned from each modality in a mutually supervised manner. Moreover, we incorporate the information bottleneck principle into our model to ensure that the shared and modality-specific factors encode exclusive yet complementary information. Notably, our model is designed to be trainable on both complete and partial multimodal datasets with a valid Evidence Lower Bound. Extensive experimental results demonstrate significant improvements of our model over existing methods on various tasks including cross-modal generation, clustering, and classification. Wenlan Chen, Daoyuan Wang, Fei Guo 0001, Cheng Liang 0001 |
NeurIPS | 5 |
| 2025 | DynaPhArM: Adaptive and Physics-Constrained Modeling for Target-Drug Complexes with Drug-Specific AdaptationsabstractAccurately modeling the target-drug complex at atom level presents a significant challenge in the computer-aided drug design. Traditional methods that rely solely on rigid transformations often fail to capture the adaptive interactions between targets and drugs, particularly during substantial conformational changes in targets upon ligand binding, which becomes especially critical when learning target-drug interactions in drug design. Accurately modeling these changes is crucial for understanding target-drug interactions and improving drug efficacy. To address these challenges, we introduce DynaPhArM, an SE(3)-Equivariant Transformer model specifically designed to capture adaptive alterations occurring within target-drug interactions. DynaPhArM utilizes the cooperative scalar-vector representation, drug-specific embeddings, and a diffusion process to effectively model the evolving dynamics of interactions between targets and drugs. Furthermore, we integrate physical information and energetic principles that maintain essential geometric constraints, such as bond lengths, bond angles, van der Waals forces (vdW), within a multi-task learning (MTL) framework to enhance accuracy. Experimental results demonstrate that DynaPhArM achieves state-of-the-art performance with an overall root mean square deviation (RMSD) of 2.01 Å and a sc-RMSD of 0.29 Å while exhibiting higher success rates compared to existing methodologies. Additionally, DynaPhArM shows promise in enhancing drug specificity, thereby simulating how targets adapt to various drugs through precise modeling of atomic-level interactions and conformational flexibility. Diya Zhang, Mengwei Sun, Xingdan Wang, Cheng Liang 0001, Qiaozhen Meng, Shiqiang Ma, Fei Guo 0001 |
NeurIPS | 4 |
| 2025 | EDDINet: Enhancing drug-drug interaction prediction via information flow and consensus constrained multi-graph contrastive learningabstractPredicting drug–drug interactions (DDIs) is crucial for understanding and preventing adverse drug reactions (ADRs). However, most existing methods inadequately explore the interactive information between drugs in a self-supervised manner, limiting our comprehension of drug–drug associations. This paper introduces EDDINet : E nhancing D rug- D rug I nteraction Prediction via Information Flow and Consensus-Constrained Multi-Graph Contrastive Learning for precise DDI prediction. We first present a cross-modal information-flow mechanism to integrate diverse drug features, enriching the structural insights conveyed by the drug feature vector. Next, we employ contrastive learning to filter various biological networks, enhancing the model’s robustness. Additionally, we propose a consensus regularization framework that collaboratively trains multi-view models, producing high-quality drug representations. To unify drug representations derived from different biological information, we utilize an attention mechanism for DDI prediction. Extensive experiments demonstrate that EDDINet surpasses state-of-the-art unsupervised models and outperforms some supervised baseline models in DDI prediction tasks. Our approach shows significant advantages and holds promising potential for advancing DDI research and improving drug safety assessments. Our codes are available at: https://github.com/95LY/EDDINet_code . • EDDINet is proposed via Information Flow and Consensus-Constrained Multi-Graph Contrastive Learning. • Introduce the cross-modal information-flow mechanism to facilitate the information across various drug features. • EDDINet surpasses state-of-the-art unsupervised models and even outperforms some supervised baseline models in DDI prediction tasks. Hong Wang 0015, Luhe Zhuang, Yijie Ding, Prayag Tiwari, Cheng Liang 0001 |
Artif. Intell. Medicine | 5 |
| 2025 | PCLSurv: a prototypical contrastive learning-based multi-omics data integration model for cancer survival predictionabstractAccurate cancer survival prediction remains a critical challenge in clinical oncology, largely due to the complex and multi-omics nature of cancer data. Existing methods often struggle to capture the comprehensive range of informative features required for precise predictions. Here, we introduce PCLSurv, an innovative deep learning framework designed for cancer survival prediction using multi-omics data. PCLSurv integrates autoencoders to extract omics-specific features and employs sample-level contrastive learning to identify distinct yet complementary characteristics across data views. Then, features are fused via a bilinear fusion module to construct a unified representation. To further enhance the model's capacity to capture high-level semantic relationships, PCLSurv aligns similar samples with shared prototypes while separating unrelated ones via prototypical contrastive learning. As a result, PCLSurv effectively distinguishes patient groups with varying survival outcomes at different semantic similarity levels, providing a robust framework for stratifying patients based on clinical and molecular features. We conduct extensive experiments on 11 cancer datasets. The comparison results confirm the superior performance of PCLSurv over existing alternatives. The source code of PCLSurv is freely available at https://github.com/LiangSDNULab/PCLSurv. Wenlan Chen, Hai Zhong, Cheng Liang 0001 |
Briefings Bioinform. | 4 |
| 2025 | Cancer survival prediction based on soft-label guided contrastive learning and global feature fusionabstractMOTIVATION: The high complexity and heterogeneity of cancer pose significant challenges to personalized treatment, making the improvement of cancer survival prediction accuracy crucial for clinical decision-making. The integration of multi-omics data enables a more comprehensive capture of multi-layered information in complex biological processes. However, existing survival analysis models still face limitations in accurately extracting and effectively integrating the unique and shared information from multi-omics data. RESULTS: In this article, we propose a novel prediction model for cancer survival based on soft-label guided contrastive learning and global feature fusion, namely SLCGF. Our model first extracts paired feature representations for each omics using Siamese encoders. We then perform intra-view and inter-view contrastive learning simultaneously, employing a neighborhood-based paradigm to enhance feature discrimination and alignment across omics. To ensure reliable neighbor retention and improve model robustness, we treat the affinities between samples and their high-order neighbors as soft labels to guide the contrastive learning process at both levels. In addition, we adopt a global self-attention mechanism to obtain the unified representation for cancer survival prediction, where the cross-omics connections are fully exploited and complementary information is adaptively integrated. We comprehensively evaluate the performance of our model on 13 cancer multi-omics datasets, and the experimental results demonstrate its superiority over existing approaches. AVAILABILITY AND IMPLEMENTATION: Source code is available at https://github.com/LiangSDNULab/SLCGF. Huiying Jiang, Wenlan Chen, Fei Guo 0001, Cheng Liang 0001 |
Bioinform. | 4 |
| 2025 | Unsupervised multi-view feature selection based on weighted low-rank tensor learning and its application in multi-omics datasets
Daoyuan Wang, Lianzhi Wang, Wenlan Chen, Hong Wang 0015, Cheng Liang 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Robust high-order graph learning for incomplete multi-view clustering
Daoyuan Wang, Fujian Ren, Yuntang Zhuang, Cheng Liang 0001 |
Expert Syst. Appl. | 4 |
| 2024 | Deep learning model for protein multi-label subcellular localization and function prediction based on multi-task collaborative trainingabstractThe functional study of proteins is a critical task in modern biology, playing a pivotal role in understanding the mechanisms of pathogenesis, developing new drugs, and discovering novel drug targets. However, existing computational models for subcellular localization face significant challenges, such as reliance on known Gene Ontology (GO) annotation databases or overlooking the relationship between GO annotations and subcellular localization. To address these issues, we propose DeepMTC, an end-to-end deep learning-based multi-task collaborative training model. DeepMTC integrates the interrelationship between subcellular localization and the functional annotation of proteins, leveraging multi-task collaborative training to eliminate dependence on known GO databases. This strategy gives DeepMTC a distinct advantage in predicting newly discovered proteins without prior functional annotations. First, DeepMTC leverages pre-trained language model with high accuracy to obtain the 3D structure and sequence features of proteins. Additionally, it employs a graph transformer module to encode protein sequence features, addressing the problem of long-range dependencies in graph neural networks. Finally, DeepMTC uses a functional cross-attention mechanism to efficiently combine upstream learned functional features to perform the subcellular localization task. The experimental results demonstrate that DeepMTC outperforms state-of-the-art models in both protein function prediction and subcellular localization. Moreover, interpretability experiments revealed that DeepMTC can accurately identify the key residues and functional domains of proteins, confirming its superior performance. The code and dataset of DeepMTC are freely available at https://github.com/ghli16/DeepMTC. Peihao Bai, Guanghui Li 0003, Jiawei Luo 0001, Cheng Liang 0001 |
Briefings Bioinform. | 4 |
| 2024 | Drug repositioning based on residual attention network and free multiscale adversarial trainingabstractBACKGROUND: Conducting traditional wet experiments to guide drug development is an expensive, time-consuming and risky process. Analyzing drug function and repositioning plays a key role in identifying new therapeutic potential of approved drugs and discovering therapeutic approaches for untreated diseases. Exploring drug-disease associations has far-reaching implications for identifying disease pathogenesis and treatment. However, reliable detection of drug-disease relationships via traditional methods is costly and slow. Therefore, investigations into computational methods for predicting drug-disease associations are currently needed. RESULTS: This paper presents a novel drug-disease association prediction method, RAFGAE. First, RAFGAE integrates known associations between diseases and drugs into a bipartite network. Second, RAFGAE designs the Re_GAT framework, which includes multilayer graph attention networks (GATs) and two residual networks. The multilayer GATs are utilized for learning the node embeddings, which is achieved by aggregating information from multihop neighbors. The two residual networks are used to alleviate the deep network oversmoothing problem, and an attention mechanism is introduced to combine the node embeddings from different attention layers. Third, two graph autoencoders (GAEs) with collaborative training are constructed to simulate label propagation to predict potential associations. On this basis, free multiscale adversarial training (FMAT) is introduced. FMAT enhances node feature quality through small gradient adversarial perturbation iterations, improving the prediction performance. Finally, tenfold cross-validations on two benchmark datasets show that RAFGAE outperforms current methods. In addition, case studies have confirmed that RAFGAE can detect novel drug-disease associations. CONCLUSIONS: The comprehensive experimental results validate the utility and accuracy of RAFGAE. We believe that this method may serve as an excellent predictor for identifying unobserved disease-drug associations. Guanghui Li 0003, Shuwen Li, Cheng Liang 0001, Qiu Xiao, Jiawei Luo 0001 |
BMC Bioinform. | 3 |
| 2024 | Sparse graph cascade multi-kernel fusion contrastive learning for microbe-disease association prediction
Shengpeng Yu, Hong Wang 0015, Meifang Hua, Cheng Liang 0001, Yanshen Sun |
Expert Syst. Appl. | 4 |
| 2024 | A Multi-Relational Graph Encoder Network for Fine-Grained Prediction of MiRNA-Disease AssociationsabstractMicroRNAs (miRNAs) are critical in diagnosing and treating various diseases. Automatically demystifying the interdependent relationships between miRNAs and diseases has recently made remarkable progress, but their fine-grained interactive relationships still need to be explored. We propose a multi-relational graph encoder network for fine-grained prediction of miRNA-disease associations (MRFGMDA), which uses practical and current datasets to construct a multi-relational graph encoder network to predict disease-related miRNAs and their specific relationship types (upregulation, downregulation, or dysregulation). We evaluated MRFGMDA and found that it accurately predicted miRNA-disease associations, which could have far-reaching implications for clinical medical analysis, early diagnosis, prevention, and treatment. Case analyses, Kaplan-Meier survival analysis, expression difference analysis, and immune infiltration analysis further demonstrated the effectiveness and feasibility of MRFGMDA in uncovering potential disease-related miRNAs. Overall, our work represents a significant step toward improving the prediction of miRNA-disease associations using a fine-grained approach could lead to more accurate diagnosis and treatment of diseases. Shengpeng Yu, Hong Wang 0015, Jing Li 0156, Jun Zhao 0017, Cheng Liang 0001, Yanshen Sun |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2024 | Robust Tensor Subspace Learning for Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering has represented a significant role in grouping real images. In this study, a novel robust tensor subspace learning (RTSL) is proposed for incomplete multi-view clustering. Specifically, the missing samples within views are first recovered by matrix factorization. The recovered information is utilized for latent representations learning. And then, the obtained latent representations are organized from all views into a third-order tensor and the intrinsic sample relations are captured with tensor linear representation. Moreover, a low-rank sample coefficient tensor is sought to capture high-order connections among views by imposing the tensor nuclear norm. Compared with traditional learning paradigms in the vector space, the sample relations within each view as well as across views could be preserved with the aid of robust tensor subspace learning. As a result, our model can simultaneously handle the missing samples and exploit the intrinsic correlations, leading to enhanced representation capability and better quality of the recovered data. We design an efficient iterative optimization strategy to solve the proposed method. Experimental results on eight datasets show that our model outperforms other competing approaches. Cheng Liang 0001, Daoyuan Wang, Huaxiang Zhang 0001, Shichao Zhang 0001, Fei Guo 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Deep multi-view contrastive learning for cancer subtype identificationabstractCancer heterogeneity has posed great challenges in exploring precise therapeutic strategies for cancer treatment. The identification of cancer subtypes aims to detect patients with distinct molecular profiles and thus could provide new clues on effective clinical therapies. While great efforts have been made, it remains challenging to develop powerful computational methods that can efficiently integrate multi-omics datasets for the task. In this paper, we propose a novel self-supervised learning model called Deep Multi-view Contrastive Learning (DMCL) for cancer subtype identification. Specifically, by incorporating the reconstruction loss, contrastive loss and clustering loss into a unified framework, our model simultaneously encodes the sample discriminative information into the extracted feature representations and well preserves the sample cluster structures in the embedded space. Moreover, DMCL is an end-to-end framework where the cancer subtypes could be directly obtained from the model outputs. We compare DMCL with eight alternatives ranging from classic cancer subtype identification methods to recently developed state-of-the-art systems on 10 widely used cancer multi-omics datasets as well as an integrated dataset, and the experimental results validate the superior performance of our method. We further conduct a case study on liver cancer and the analysis results indicate that different subtypes might have different responses to the selected chemotherapeutic drugs. Wenlan Chen, Hong Wang 0015, Cheng Liang 0001 |
Briefings Bioinform. | 3 |
| 2023 | Identifying spatial domains of spatially resolved transcriptomics via multi-view graph convolutional networksabstractMOTIVATION: Recent advances in spatially resolved transcriptomics (ST) technologies enable the measurement of gene expression profiles while preserving cellular spatial context. Linking gene expression of cells with their spatial distribution is essential for better understanding of tissue microenvironment and biological progress. However, effectively combining gene expression data with spatial information to identify spatial domains remains challenging. RESULTS: To deal with the above issue, in this paper, we propose a novel unsupervised learning framework named STMGCN for identifying spatial domains using multi-view graph convolution networks (MGCNs). Specifically, to fully exploit spatial information, we first construct multiple neighbor graphs (views) with different similarity measures based on the spatial coordinates. Then, STMGCN learns multiple view-specific embeddings by combining gene expressions with each neighbor graph through graph convolution networks. Finally, to capture the importance of different graphs, we further introduce an attention mechanism to adaptively fuse view-specific embeddings and thus derive the final spot embedding. STMGCN allows for the effective utilization of spatial context to enhance the expressive power of the latent embeddings with multiple graph convolutions. We apply STMGCN on two simulation datasets and five real spatial transcriptomics datasets with different resolutions across distinct platforms. The experimental results demonstrate that STMGCN obtains competitive results in spatial domain identification compared with five state-of-the-art methods, including spatial and non-spatial alternatives. Besides, STMGCN can detect spatially variable genes with enriched expression patterns in the identified domains. Overall, STMGCN is a powerful and efficient computational framework for identifying spatial domains in spatial transcriptomics data. Xuejing Shi, Juntong Zhu, Yahui Long, Cheng Liang 0001 |
Briefings Bioinform. | 4 |
| 2023 | CasANGCL: pre-training and fine-tuning model based on cascaded attention network and graph contrastive learning for molecular property predictionabstractMOTIVATION: Molecular property prediction is a significant requirement in AI-driven drug design and discovery, aiming to predict the molecular property information (e.g. toxicity) based on the mined biomolecular knowledge. Although graph neural networks have been proven powerful in predicting molecular property, unbalanced labeled data and poor generalization capability for new-synthesized molecules are always key issues that hinder further improvement of molecular encoding performance. RESULTS: We propose a novel self-supervised representation learning scheme based on a Cascaded Attention Network and Graph Contrastive Learning (CasANGCL). We design a new graph network variant, designated as cascaded attention network, to encode local-global molecular representations. We construct a two-stage contrast predictor framework to tackle the label imbalance problem of training molecular samples, which is an integrated end-to-end learning scheme. Moreover, we utilize the information-flow scheme for training our network, which explicitly captures the edge information in the node/graph representations and obtains more fine-grained knowledge. Our model achieves an 81.9% ROC-AUC average performance on 661 tasks from seven challenging benchmarks, showing better portability and generalizations. Further visualization studies indicate our model's better representation capacity and provide interpretability. Zixi Zheng, Yanyan Tan, Hong Wang 0015, Shengpeng Yu, Tianyu Liu 0006, Cheng Liang 0001 |
Briefings Bioinform. | 6 |
| 2023 | Protein-protein interaction site prediction by model ensembling with hybrid feature and self-attentionabstractBACKGROUND: Protein-protein interactions (PPIs) are crucial in various biological functions and cellular processes. Thus, many computational approaches have been proposed to predict PPI sites. Although significant progress has been made, these methods still have limitations in encoding the characteristics of each amino acid in sequences. Many feature extraction methods rely on the sliding window technique, which simply merges all the features of residues into a vector. The importance of some key residues may be weakened in the feature vector, leading to poor performance. RESULTS: We propose a novel sequence-based method for PPI sites prediction. The new network model, PPINet, contains multiple feature processing paths. For a residue, the PPINet extracts the features of the targeted residue and its context separately. These two types of features are processed by two paths in the network and combined to form a protein representation, where the two types of features are of relatively equal importance. The model ensembling technique is applied to make use of more features. The base models are trained with different features and then ensembled via stacking. In addition, a data balancing strategy is presented, by which our model can get significant improvement on highly unbalanced data. CONCLUSION: The proposed method is evaluated on a fused dataset constructed from Dset186, Dset_72, and PDBset_164, as well as the public Dset_448 dataset. Compared with current state-of-the-art methods, the performance of our method is better than the others. In the most important metrics, such as AUPRC and recall, it surpasses the second-best programmer on the latter dataset by 6.9% and 4.7%, respectively. We also demonstrated that the improvement is essentially due to using the ensemble model, especially, the hybrid feature. We share our code for reproducibility and future research at https://github.com/CandiceCong/StackingPPINet . Hanhan Cong, Hong Liu 0013, Cheng Liang 0001, Yuehui Chen |
BMC Bioinform. | 4 |
| 2023 | EMPPNet: Enhancing Molecular Property Prediction via Cross-modal Information Flow and Hierarchical Attention
Zixi Zheng, Hong Wang 0015, Yanyan Tan, Cheng Liang 0001, Yanshen Sun |
Expert Syst. Appl. | 4 |
| 2023 | Incomplete multi-view clustering by simultaneously learning robust representations and optimal graph structures
Mingchao Shang, Cheng Liang 0001, Jiawei Luo 0001, Huaxiang Zhang 0001 |
Inf. Sci. | 2 |
| 2023 | Multi-view unsupervised feature selection with tensor robust principal component analysis and consensus graph learningabstractRecently, multi-view unsupervised feature selection has attracted much attention due to its efficiency and better interpretability in processing high-dimensional multi-view datasets. Most existing methods rely on the constructed similarity matrices to obtain reliable pseudo labels to guide the feature selection. However, the considerable adverse noise in the raw data inevitably impedes the exploration of true underlying similarity structures. Besides, the inter-view correlations are often ignored during the common representation learning , which limits the effective fusion of the essential information from multiple views. To solve these issues, we design a novel robust multi-view unsupervised feature selection framework. Specifically, our method seeks a set of noise-free view-specific similarity matrices by leveraging tensor robust principal component analysis , where the high-order connections among different views are well exploited through the constructed low-rank tensor. Meanwhile, a high-quality consensus similarity matrix is adaptively learned from the view-specific representations within the same unified framework to capture the shared local structures. To enhance the discriminative ability of the feature selection matrix, we further impose a rank constraint on the consensus similarity matrix to obtain reliable pseudo cluster indicators. We present an efficient optimization algorithm ground on the alternating direction method of multipliers to solve the proposed model. Experimental results on six multi-view datasets confirm the superiority of our method. Cheng Liang 0001, Lianzhi Wang, Li Liu 0031, Huaxiang Zhang 0001, Fei Guo 0001 |
Pattern Recognit. | 1 |
| 2023 | Multiview Robust Graph-Based Clustering for Cancer Subtype IdentificationabstractCancer subtype identification is to classify cancer into groups according to their molecular characteristics and clinical manifestations and is the basis for more personalized diagnosis and therapy. Public datasets such as The Cancer Genome Atlas (TCGA) have collected a massive number of multi-omics data. The accumulation of these datasets provides unprecedented opportunities to study the mechanism of cancers and further identify cancer subtypes at a comprehensive level. In this paper, we propose a multi-view robust graph-based clustering (MRGC) method to effectively identify cancer subtypes. Our method first learns robust latent representations from the raw omics data to alleviate the influences of the noise, where a set of similarity matrices are then adaptively learned based on these new representations. Finally, a global similarity graph is obtained by exploiting the consensus structure from the graphs. As a result, the three parts in our method can reinforce each other in a mutual iterative manner. We conduct extensive experiments on both generic machine learning datasets and cancer datasets. The experimental results confirm that our model can achieve satisfactory clustering performance compared to several state-of-the-art approaches. Moreover, we convey the practicability of MRGC by carrying out a case study on hepatocellular carcinoma. Cheng Liang 0001, Hong Wang 0015 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | Multi-view Unsupervised Feature Selection via Consensus Guided Low-rank Tensor LearningabstractRecently, with the exponentially increased amount of multi-view data in various fields such as Multimedia and bioinformatics, multi-view unsupervised feature selection has attracted much attention due to its necessity in dealing with high-dimensional features. Although previous approaches have achieved great success, they generally ignore the consistent information and the high-order connections among views. In this paper, we present a general multi-view unsupervised feature selection model which integrates the common graph learning and feature selection into a unified framework. Specifically, our approach first learns a pseudo label matrix for each view by preserving the local data structure, and then stack them into a third-order tensor with low-rank constraint to explore the high-order connections among the views. In order to exploit the consistent information among different views, we seek a consensus graph matrix with optimal cluster structure by taking advantage of the view-specific pseudo label matrices and the rank constraint. Meanwhile, we adopt the sparse regression model to select discriminative features under the guidance of the final pseudo labels obtained from the learned consensus graph. We introduce an alternate optimization algorithm ground on the alternating direction method of multipliers (ADMM) to optimize the presented method. Extensive experiments on both machine learning and single-cell multi-omics datasets prove the effectiveness of our method. Moreover, the case study carried out on an ovarian cancer dataset further confirms the applicability of our method in identifying non-redundant and representative features. Lianzhi Wang, Cheng Liang 0001, Wenjiao Dong, Wenlan Chen |
BIBM | 2 |
| 2022 | A knowledge-driven network for fine-grained relationship detection between miRNA and diseaseabstractIncreasing biological evidence indicated that microRNAs (miRNAs) play a vital role in exploring the pathogenesis of various human diseases (especially in tumors). Mining disease-related miRNAs is of great significance for the clinical diagnosis and treatment of diseases. Compared with the traditional experimental methods with the significant limitations of high cost, long cycle and small scale, the methods based on computing have the advantages of being cost-effective. However, although the current methods based on computational biology can accurately predict the correlation between miRNAs and disease, they can not predict the detailed association information at a fine level. We propose a knowledge-driven approach to the fine-grained prediction of disease-related miRNAs (KDFGMDA). Different from the previous methods, this method can finely predict the clear associations between miRNA and disease, such as upregulation, downregulation or dysregulation. Specifically, KDFGMDA extracts triple information from massive experimental data and existing datasets to construct a knowledge graph and then trains a depth graph representation learning model based on knowledge graph to complete fine-grained prediction tasks. Experimental results show that KDFGMDA can predict the relationship between miRNA and disease accurately, which is of far-reaching significance for medical clinical research and early diagnosis, prevention and treatment of diseases. Additionally, the results of case studies on three types of cancers, Kaplan-Meier survival analysis and expression difference analysis further provide the effectiveness and feasibility of KDFGMDA to detect potential candidate miRNAs. Availability: Our work can be downloaded from https://github.com/ShengPengYu/KDFGMDA. Shengpeng Yu, Hong Wang 0015, Tianyu Liu 0006, Cheng Liang 0001, Jiawei Luo 0001 |
Briefings Bioinform. | 4 |
| 2022 | Predicting miRNA-disease associations based on graph attention network with multi-source informationabstractBACKGROUND: There is a growing body of evidence from biological experiments suggesting that microRNAs (miRNAs) play a significant regulatory role in both diverse cellular activities and pathological processes. Exploring miRNA-disease associations not only can decipher pathogenic mechanisms but also provide treatment solutions for diseases. As it is inefficient to identify undiscovered relationships between diseases and miRNAs using biotechnology, an explosion of computational methods have been advanced. However, the prediction accuracy of existing models is hampered by the sparsity of known association network and single-category feature, which is hard to model the complicated relationships between diseases and miRNAs. RESULTS: In this study, we advance a new computational framework (GATMDA) to discover unknown miRNA-disease associations based on graph attention network with multi-source information, which effectively fuses linear and non-linear features. In our method, the linear features of diseases and miRNAs are constructed by disease-lncRNA correlation profiles and miRNA-lncRNA correlation profiles, respectively. Then, the graph attention network is employed to extract the non-linear features of diseases and miRNAs by aggregating information of each neighbor with different weights. Finally, the random forest algorithm is applied to infer the disease-miRNA correlation pairs through fusing linear and non-linear features of diseases and miRNAs. As a result, GATMDA achieves impressive performance: an average AUC of 0.9566 with five-fold cross validation, which is superior to other previous models. In addition, case studies conducted on breast cancer, colon cancer and lymphoma indicate that 50, 50 and 48 out of the top fifty prioritized candidates are verified by biological experiments. CONCLUSIONS: The extensive experimental results justify the accuracy and utility of GATMDA and we could anticipate that it may regard as a utility tool for identifying unobserved disease-miRNA relationships. Guanghui Li 0003, Yuejin Zhang, Cheng Liang 0001, Qiu Xiao, Jiawei Luo 0001 |
BMC Bioinform. | 4 |
| 2021 | Cancer Subtype Identification based on Multi-view Subspace Clustering with Adaptive Local Structure LearningabstractIdentifying cancer subtypes is of key significance for the application of personalized medicine. It is devoted to using the unsupervised clustering methods to stratify cancer patients into diverse subgroups and provide specific treatments accordingly. Recently, the abundant multi-omics data have provided unprecedented opportunities to discover intrinsic subtypes at an integrative level. However, it still remains a challenging task to combine these heterogeneous data and reveal the underlying clustering patterns. In this paper, we propose a novel method to identify cancer subtypes based on multi-view subspace clustering with adaptive local structure learning. Specifically, our model can effectively explore the local similarity structure from each omics data and simultaneously discover the unified consensus similarity structure among all omics data. Moreover, we design an efficient Augmented Lagrangian Multiplier algorithm to optimize the proposed framework. The experimental results on both cancer datasets and image datasets confirmed the superior performance of the proposed method over existing alternatives. In addition, we conduct a case study on melanoma and identify a set of molecular functions specific to each subtype. Mingchao Shang, Huaxiang Zhang 0001, Cheng Liang 0001 |
BIBM | 4 |
| 2021 | Cancer subtype identification by consensus guided graph autoencodersabstractMOTIVATION: Cancer subtype identification aims to divide cancer patients into subgroups with distinct clinical phenotypes and facilitate the development for subgroup specific therapies. The massive amount of multi-omics datasets accumulated in the public databases have provided unprecedented opportunities to fulfill this task. As a result, great computational efforts have been made to accurately identify cancer subtypes via integrative analysis of these multi-omics datasets. RESULTS: In this article, we propose a Consensus Guided Graph Autoencoder (CGGA) to effectively identify cancer subtypes. First, we learn for each omic a new feature matrix by using graph autoencoders, where both structure information and node features can be effectively incorporated during the learning process. Second, we learn a set of omic-specific similarity matrices together with a consensus matrix based on the features obtained in the first step. The learned omic-specific similarity matrices are then fed back to the graph autoencoders to guide the feature learning. By iterating the two steps above, our method obtains a final consensus similarity matrix for cancer subtyping. To comprehensively evaluate the prediction performance of our method, we compare CGGA with several approaches ranging from general-purpose multi-view clustering algorithms to multi-omics-specific integrative methods. The experimental results on both generic datasets and cancer datasets confirm the superiority of our method. Moreover, we validate the effectiveness of our method in leveraging multi-omics datasets to identify cancer subtypes. In addition, we investigate the clinical implications of the obtained clusters for glioblastoma and provide new insights into the treatment for patients with different subtypes. AVAILABILITYAND IMPLEMENTATION: The source code of our method is freely available at https://github.com/alcs417/CGGA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Cheng Liang 0001, Mingchao Shang, Jiawei Luo 0001 |
Bioinform. | 1 |
| 2021 | Inferring Synergistic Drug Combinations Based on Symmetric Meta-Path in a Novel Heterogeneous NetworkabstractCombinatorial drug therapy is a promising way for treating cancers, which can reduce drug side effects and improve drug efficacy. However, due to the large-scale combinatorial space, it is difficult to quickly and effectively identify novel synergistic drug combinations for further implementing combinatorial drug therapy. The computational method of fusing multi-source knowledge is a time- and cost-efficient strategy to infer synergistic drug combinations for testing. However, for the existing computational methods of inferring synergistic drug combinations, it still remains a challenging to effectively combine multi-source information to achieve the desired results. Hence, in this study, we developed a novel Inference method of Synergistic Drug Combinations based on Symmetric Meta-Path (ISDCSMP), which can systematically and accurately prioritize synergistic drug combinations in a novel drug-target heterogeneous network integrating multi-source information. In the experiment, ISDCSMP outperformed the state-of-the-art methods in terms of AUC and precision on the benchmark dataset in five-fold cross validation. Moreover, we further illustrated performances of different ways for obtaining the combination coefficients, and analyzed the influences of the maximum meta-path length. The performances of various single meta-paths were described in five-fold cross validation. Finally, we confirmed the practical usefulness of ISDCSMP with the predicted novel synergistic drug combinations. The source code of ISDCSMP is available at https://github.com/KDDing/ISDCSMP. Pingjian Ding, Cheng Liang 0001, Wenjue Ouyang, Guanghui Li 0003, Qiu Xiao, Jiawei Luo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2020 | Multi-view feature selection via Nonnegative Structured Graph Learning
Xiangpin Bai, Lei Zhu 0002, Cheng Liang 0001, Jingjing Li 0001, Xiushan Nie, Xiaojun Chang |
Neurocomputing | 3 |
| 2020 | Robust optimal graph clustering
Fei Wang 0055, Lei Zhu 0002, Cheng Liang 0001, Jingjing Li 0001, Xiaojun Chang, Ke Lu 0001 |
Neurocomputing | 3 |
| 2020 | Potential circRNA-disease association prediction using DeepWalk and network consistency projection
Guanghui Li 0003, Jiawei Luo 0001, Diancheng Wang, Cheng Liang 0001, Qiu Xiao, Pingjian Ding, Hailin Chen |
J. Biomed. Informatics | 4 |
| 2020 | Identifying lncRNA and mRNA Co-Expression Modules from Matched Expression Data in Ovarian CancerabstractLong non-coding RNAs (lncRNAs) have been shown to be involved in multiple biological processes and play critical roles in tumorigenesis. Numerous lncRNAs have been discovered in diverse species, but the functions of most lncRNAs still remain unclear. Meanwhile, their expression patterns and regulation mechanisms are also far from being fully understood. With the advances of high-throughput technologies, the increasing availability of genomic data creates opportunities for deciphering the molecular mechanism and underlying pathogenesis of human diseases. Here, we develop an integrative framework called JONMF to identify lncRNA-mRNA co-expression modules based on the sample-matched lncRNA and mRNA expression profiles. We formulate the module detection task as an optimization problem with joint orthogonal non-negative matrix factorization that could effectively prevent multicollinearity and produce a good modularity interpretation. The constructed lncRNA-mRNA co-expression network and the gene-gene interaction network are used as the network-regularized constraints to improve the module accuracy, while the sparsity constraints are simultaneously utilized to achieve modular sparse solutions. We applied JONMF to human ovarian cancer dataset and the experiment results demonstrate that the proposed method can effectively discover biologically functional co-expression modules, which may provide insights into the function of lncRNAs and molecular mechanism of human diseases. Qiu Xiao, Jiawei Luo 0001, Cheng Liang 0001, Guanghui Li 0003, Pingjian Ding, Ying Liu 0027 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2019 | CeModule: an integrative framework for discovering regulatory patterns from genomic data in cancerabstractBACKGROUND: Non-coding RNAs (ncRNAs) are emerging as key regulators and play critical roles in a wide range of tumorigenesis. Recent studies have suggested that long non-coding RNAs (lncRNAs) could interact with microRNAs (miRNAs) and indirectly regulate miRNA targets through competing interactions. Therefore, uncovering the competing endogenous RNA (ceRNA) regulatory mechanism of lncRNAs, miRNAs and mRNAs in post-transcriptional level will aid in deciphering the underlying pathogenesis of human polygenic diseases and may unveil new diagnostic and therapeutic opportunities. However, the functional roles of vast majority of cancer specific ncRNAs and their combinational regulation patterns are still insufficiently understood. RESULTS: Here we develop an integrative framework called CeModule to discover lncRNA, miRNA and mRNA-associated regulatory modules. We fully utilize the matched expression profiles of lncRNAs, miRNAs and mRNAs and establish a model based on joint orthogonality non-negative matrix factorization for identifying modules. Meanwhile, we impose the experimentally verified miRNA-lncRNA interactions, the validated miRNA-mRNA interactions and the weighted gene-gene network into this framework to improve the module accuracy through the network-based penalties. The sparse regularizations are also used to help this model obtain modular sparse solutions. Finally, an iterative multiplicative updating algorithm is adopted to solve the optimization problem. CONCLUSIONS: We applied CeModule to two cancer datasets including ovarian cancer (OV) and uterine corpus endometrial carcinoma (UCEC) obtained from TCGA. The modular analysis indicated that the identified modules involving lncRNAs, miRNAs and mRNAs are significantly associated and functionally enriched in cancer-related biological processes and pathways, which may provide new insights into the complex regulatory mechanism of human diseases at the system level. Qiu Xiao, Jiawei Luo 0001, Cheng Liang 0001, Guanghui Li 0003, Buwen Cao |
BMC Bioinform. | 3 |
| 2019 | Adaptive multi-view multi-label learning for identifying disease-associated candidate miRNAsabstractIncreasing evidence has indicated that microRNAs(miRNAs) play vital roles in various pathological processes and thus are closely related with many complex human diseases. The identification of potential disease-related miRNAs offers new opportunities to understand disease etiology and pathogenesis. Although there have been numerous computational methods proposed to predict reliable miRNA-disease associations, they suffer from various limitations that affect the prediction accuracy and their applicability. In this study, we develop a novel method to discover disease-related candidate miRNAs based on Adaptive Multi-View Multi-Label learning(AMVML). Specifically, considering the inherent noise existed in the current dataset, we propose to learn a new affinity graph adaptively for both diseases and miRNAs from multiple similarity profiles. We then simultaneously update the miRNA-disease association predicted from both spaces based on multi-label learning. In particular, we prove the convergence of AMVML theoretically and the corresponding analysis indicates that it has a fast convergence rate. To comprehensively illustrate the prediction performance of our method, we compared AMVML with four state-of-the-art methods under different validation frameworks. As a result, our method achieved comparable performance under various evaluation metrics, which suggests that our method is capable of discovering greater number of true miRNA-disease associations. The case study conducted on thyroid neoplasms further identified a potential diagnostic biomarker. Together, the experimental results confirms the utility of our method and we anticipate that our method could serve as a reliable and efficient tool for uncovering novel disease-related miRNAs. Cheng Liang 0001, Shengpeng Yu, Jiawei Luo 0001 |
PLoS Comput. Biol. | 1 |
| 2018 | A graph regularized non-negative matrix factorization method for identifying microRNA-disease associationsabstractMOTIVATION: MicroRNAs (miRNAs) play crucial roles in post-transcriptional regulations and various cellular processes. The identification of disease-related miRNAs provides great insights into the underlying pathogenesis of diseases at a system level. However, most existing computational approaches are biased towards known miRNA-disease associations, which is inappropriate for those new diseases or miRNAs without any known association information. RESULTS: In this study, we propose a new method with graph regularized non-negative matrix factorization in heterogeneous omics data, called GRNMF, to discover potential associations between miRNAs and diseases, especially for new diseases and miRNAs or those diseases and miRNAs with sparse known associations. First, we integrate the disease semantic information and miRNA functional information to estimate disease similarity and miRNA similarity, respectively. Considering that there is no available interaction observed for new diseases or miRNAs, a preprocessing step is developed to construct the interaction score profiles that will assist in prediction. Next, a graph regularized non-negative matrix factorization framework is utilized to simultaneously identify potential associations for all diseases. The results indicated that our proposed method can effectively prioritize disease-associated miRNAs with higher accuracy compared with other recent approaches. Moreover, case studies also demonstrated the effectiveness of GRNMF to infer unknown miRNA-disease associations for those novel diseases and miRNAs. AVAILABILITY AND IMPLEMENTATION: The code of GRNMF is freely available at https://github.com/XIAO-HN/GRNMF/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qiu Xiao, Jiawei Luo 0001, Cheng Liang 0001, Pingjian Ding |
Bioinform. | 3 |
| 2018 | Semi-supervised prediction of human miRNA-disease association based on graph regularization framework in heterogeneous networks
Jiawei Luo 0001, Pingjian Ding, Cheng Liang 0001, Xiangtao Chen |
Neurocomputing | 3 |
| 2018 | Human disease MiRNA inference by combining target information based on heterogeneous manifolds
Pingjian Ding, Jiawei Luo 0001, Cheng Liang 0001, Qiu Xiao, Buwen Cao |
J. Biomed. Informatics | 3 |
| 2018 | Predicting microRNA-disease associations using label propagation based on linear neighborhood similarity
Guanghui Li 0003, Jiawei Luo 0001, Qiu Xiao, Cheng Liang 0001, Pingjian Ding |
J. Biomed. Informatics | 4 |
| 2017 | Collective Prediction of Disease-Associated miRNAs Based on Transduction LearningabstractThe discovery of human disease-related miRNA is a challenging problem for complex disease biology research. For existing computational methods, it is difficult to achieve excellent performance with sparse known miRNA-disease association verified by biological experiment. Here, we develop CPTL, a Collective Prediction based on Transduction Learning, to systematically prioritize miRNAs related to disease. By combining disease similarity, miRNA similarity with known miRNA-disease association, we construct a miRNA-disease network for predicting miRNA-disease association. Then, CPTL calculates relevance score and updates the network structure iteratively, until a convergence criterion is reached. The relevance score of node including miRNA and disease is calculated by the use of transduction learning based on its neighbors. The network structure is updated using relevance score, which increases the weight of important links. To show the effectiveness of our method, we compared CPTL with existing methods based on HMDD datasets. Experimental results indicate that CPTL outperforms existing approaches in terms of AUC, precision, recall, and F1-score. Moreover, experiments performed with different number of iterations verify that CPTL has good convergence. Besides, it is analyzed that the varying of weighted parameters affect predicted results. Case study on breast cancer has further confirmed the identification ability of CPTL. Jiawei Luo 0001, Pingjian Ding, Cheng Liang 0001, Buwen Cao, Xiangtao Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2016 | Detecting overlapping protein complexes in weighted protein-protein interaction networks using pseudo-clique extension based on fuzzy relationabstractDetecting overlapping protein complexes in protein-protein interaction (PPI) networks can provide insight into cellular functional organization and thus elucidate underlying cellular mechanisms. Recently, various algorithms for protein complex detection have been developed for PPI networks. However, the majority of algorithms primarily depend on network topological features and/or gene expression profile, failing to consider the inherent biological meanings between protein pairs. In this paper, we propose a method of pseudo-clique extension based on fuzzy relation (PCE-FR) that detects protein complexes from PPI networks weighted with the biological significance hidden in protein pairs. The proposed algorithm operates in three stages: it first forms the non-overlapping protein sub-structure based on fuzzy relation and then expands each sub-structure by adding neighbor proteins to maximize the cohesive score. Finally, highly overlapped candidate protein complexes are merged to form the final protein complex set. We apply PCE-FR to two yeast PPI networks and a human PPI network and validate our results by using CYC2008 and CHPC2012, respectively. Experimental results show that our method outperforms classical algorithms such as CFinder, ClusterONE, CMC, RRW, HC-PIN and ProRank+, and that it achieves ideal overall performance in terms of Precision, Accuracy, and Separation. Buwen Cao, Jiawei Luo 0001, Cheng Liang 0001, Shulin Wang |
IJCNN | 3 |
| 2016 | A Novel Method to Detect Functional microRNA Regulatory Modules by Bicliques MergingabstractUNLABELLED: MicroRNAs (miRNAs) are post-transcriptional regulators that repress the expression of their targets. They are known to work cooperatively with genes and play important roles in numerous cellular processes. Identification of miRNA regulatory modules (MRMs) would aid deciphering the combinatorial effects derived from the many-to-many regulatory relationships in complex cellular systems. Here, we develop an effective method called BiCliques Merging (BCM) to predict MRMs based on bicliques merging. By integrating the miRNA/mRNA expression profiles from The Cancer Genome Atlas (TCGA) with the computational target predictions, we construct a weighted miRNA regulatory network for module discovery. The maximal bicliques detected in the network are statistically evaluated and filtered accordingly. We then employed a greedy-based strategy to iteratively merge the remaining bicliques according to their overlaps together with edge weights and the gene-gene interactions. Comparing with existing methods on two cancer datasets from TCGA, we showed that the modules identified by our method are more densely connected and functionally enriched. Moreover, our predicted modules are more enriched for miRNA families and the miRNA-mRNA pairs within the modules are more negatively correlated. Finally, several potential prognostic modules are revealed by Kaplan-Meier survival analysis and breast cancer subtype analysis. AVAILABILITY: BCM is implemented in Java and available for download in the supplementary materials, which can be found on the Computer Society Digital Library at http://doi.ieeecomputersociety.org/10.1109/ TCBB.2015.2462370. Cheng Liang 0001, Yue Li 0017, Jiawei Luo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2015 | A novel motif-discovery algorithm to identify co-regulatory motifs in large transcription factor and microRNA co-regulatory networks in humanabstractMOTIVATION: Interplays between transcription factors (TFs) and microRNAs (miRNAs) in gene regulation are implicated in various physiological processes. It is thus important to identify biologically meaningful network motifs involving both types of regulators to understand the key co-regulatory mechanisms underlying the cellular identity and function. However, existing motif finders do not scale well for large networks and are not designed specifically for co-regulatory networks. RESULTS: In this study, we propose a novel algorithm CoMoFinder to accurately and efficiently identify composite network motifs in genome-scale co-regulatory networks. We define composite network motifs as network patterns involving at least one TF, one miRNA and one target gene that are statistically significant than expected. Using two published disease-related co-regulatory networks, we show that CoMoFinder outperforms existing methods in both accuracy and robustness. We then applied CoMoFinder to human TF-miRNA co-regulatory network derived from The Encyclopedia of DNA Elements project and identified 44 recurring composite network motifs of size 4. The functional analysis revealed that genes involved in the 44 motifs are enriched for significantly higher number of biological processes or pathways comparing with non-motifs. We further analyzed the identified composite bi-fan motif and showed that gene pairs involved in this motif structure tend to physically interact and are functionally more similar to each other than expected. AVAILABILITY AND IMPLEMENTATION: CoMoFinder is implemented in Java and available for download at http://www.cs.utoronto.ca/∼yueli/como.html. Cheng Liang 0001, Yue Li 0017, Jiawei Luo 0001, Zhaolei Zhang |
Bioinform. | 1 |
| 2014 | Mirsynergy: detecting synergistic miRNA regulatory modules by overlapping neighbourhood expansionabstractMOTIVATION: Identification of microRNA regulatory modules (MiRMs) will aid deciphering aberrant transcriptional regulatory network in cancer but is computationally challenging. Existing methods are stochastic or require a fixed number of regulatory modules. RESULTS: We propose Mirsynergy, an efficient deterministic overlapping clustering algorithm adapted from a recently developed framework. Mirsynergy operates in two stages: it first forms MiRMs based on co-occurring microRNA (miRNA) targets and then expands each MiRM by greedily including (excluding) mRNAs into (from) the MiRM to maximize the synergy score, which is a function of miRNA-mRNA and gene-gene interactions. Using expression data for ovarian, breast and thyroid cancer from The Cancer Genome Atlas, we compared Mirsynergy with internal controls and existing methods. Mirsynergy-MiRMs exhibit significantly higher functional enrichment and more coherent miRNA-mRNA expression anti-correlation. Based on Kaplan-Meier survival analysis, we proposed several prognostically promising MiRMs and envisioned their utility in cancer research. AVAILABILITY AND IMPLEMENTATION: Mirsynergy is implemented/available as an R/Bioconductor package at www.cs.utoronto.ca/∼yueli/Mirsynergy.html. Yue Li 0017, Cheng Liang 0001, Ka-Chun Wong, Jiawei Luo 0001, Zhaolei Zhang |
Bioinform. | 2 |