VLDB 2026 Research / reviewers in the wild / expert
Zhu-Hong You
dblp:35/8343 · also Zhuhong You
· DBLP profile ↗
228ranked-venue papers
14as first author
126since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 187 · 11 first-author · 104 since 2021Artificial intelligence and machine learning · 39 · 3 first-author · 20 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | GeoMoE: Geometry-Driven prompts with adaptive Mixture-of-Experts for remote sensing image semantic segmentation
Zhen Wang 0020, Nan Xu 0008, Zhu-Hong You |
Expert Syst. Appl. | 6 |
| 2026 | Dual-Channel Learning Framework for Zero-Shot CircRNA-miRNA Interaction Prediction via State Space ModelingabstractCircRNA-miRNA interaction (CMI) plays a pivotal role in disease therapeutics and drug discovery. However, existing methods face several challenges in modeling complex biological networks and zero-shot learning scenarios. Biological networks encapsulate rich biological information, yet current approaches often fail to fully exploit this depth. Moreover, zero-shot prediction requires models to identify new interactions without relying on previously observed samples, imposing stringent requirements on generalization capabilities. To address these limitations, we propose a dual-channel learning framework leveraging State space modeling for Zero-shot CMI prediction (ZeroStem). ZeroStem first enhances the biological relevance of node using prior knowledge, and employs a graph Transformer to extract macro-topological representations. Subsequently, it generates semantic subgraphs based on meta-paths to focus on specific biological relationships, utilizing the Mamba to extract micro-semantic representations via state space modeling. Finally, macro-topological and micro-semantic representations are seamlessly integrated through linear transformation and residual connections, enabling high-precision zero-shot CMI prediction. Extensive experiments on multiple benchmark datasets demonstrate that ZeroStem significantly outperforms existing methods, validating its efficiency and robust generalization in CMI prediction. Case studies further illustrate that ZeroStem offers novel insights into the molecular mechanisms underlying intricate disease-associated networks. Mengmeng Wei, Lei Wang 0121, Zhu-Hong You, Pengwei Hu 0001, Bo-Wei Zhao, Zhi-an Huang |
AAAI | 3 |
| 2026 | PEGNet-CDA: A Propagation-Enhanced Graph Network for CircRNA-Disease Association Prediction
Yue-Chao Li, Yao-Lu Li, Chen-Yv Yang, Mengmeng Wei, Xinfei Wang 0001, Lei Wang 0121, Zhi-an Huang, Zhu-Hong You |
ICIC (27) | 9 |
| 2026 | MSCANet: Multi-scale Cross-Attention Fusion Network for Multimodal Remote Sensing Image Semantic Segmentation
Zhen Wang 0020, Zhu-Hong You, Yu Li 0030 |
ICIC (8) | 4 |
| 2026 | DualEnc-CDA: Predicting CircRNA-Disease Associations via Complementary Structural Encoding
Hai-Ru You, Yue-Chao Li, Meng-Meng Wei, Xinfei Wang 0001, Yu Li 0030, Bo-Lin Chen, Zhu-Hong You |
ICIC (27) | 8 |
| 2026 | Spatial-spectral fusion enables drug repositioning by capturing indirect and long-range associations in biological networks
Lei Wang 0121, Runzhou Tang, Zhi-an Huang, Feng Tan 0002, Lun Hu, Zhu-Hong You, Pengwei Hu 0001 |
Bioinform. | 8 |
| 2026 | DeShiftNet: a deformable-shifted cross-attention network for lightweight and robust organoid image segmentationabstractBACKGROUND: Organoid image segmentation is essential for quantitative analysis in disease modeling and drug screening, yet remains highly challenging due to substantial morphological variability and blurred boundaries in organoid images. Existing approaches often struggle to achieve a favorable balance between segmentation accuracy and computational efficiency. RESULTS: In this paper, DeShiftNet, a lightweight segmentation framework, is proposed to extract discriminative features with high accuracy while maintaining low computational overhead. The model incorporates a deformable-shifted encoding strategy that adaptively samples local structures. It also includes a cross-attention-guided decoder for selective multi-scale feature alignment. Furthermore, a deformable multi-scale contextual refinement module enhances boundary coherence and contextual consistency. Extensive experiments on the multi-type OrganoID dataset show that DeShiftNet achieves competitive performance compared with recent segmentation models, while maintaining only 1.78M parameters and 2.65 GFLOPs. Notably, DeShiftNet achieves a Dice score of 0.961 on the Lung subset. CONCLUSION: These results indicate its potential practical value for efficient organoid segmentation in high-throughput experimental workflows. Le Tong, Tao Shu, Xinru Zhuang, Jingrui Bai, Lun Hu, Feng Tan 0002, Zhu-Hong You, Pengwei Hu 0001 |
BMC Bioinform. | 8 |
| 2026 | Three-dimensional geometric deep learning for reaction prediction with equivariant graph transformer
Zhouxiang Wang, Zhu-Hong You, Qiangguo Jin |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | KA-DDI: A knowledge-adaptive contrastive learning framework for drug-drug interaction prediction
Yu Li 0030, Jia-Ming Liu, Yue-Chao Li, Hai-Ru You, Zhu-Hong You, Chenggang Mi 0001 |
Expert Syst. Appl. | 5 |
| 2026 | CM-PHI: combining multi-hop attention graph neural network with sequence semantic analysis to predict phage-host interaction
Jie Pan 0007, Rui Wang 0102, Weiping Ding 0001, Yuechao Li, Zhu-Hong You, Qinghua Huang, Dawei Wei, Yanmei Sun |
Expert Syst. Appl. | 5 |
| 2026 | Cos-UMamba: Optimizing salient object detection with cosine scanning and bias-corrected feature fusion in optical remote sensing images
Zhen Wang 0020, Fu-Lin He, Nan Xu 0008, Zhu-Hong You |
Expert Syst. Appl. | 4 |
| 2026 | HpMiX: A Disease ceRNA biomarker prediction framework driven by graph topology-constrained Mixup and hypergraph residual enhancement
Xinfei Wang 0001, Lan Huang 0002, Yan Wang 0028, Renchu Guan, Zhu-Hong You, Fengfeng Zhou |
Neural Networks | 5 |
| 2026 | Multi-hop graph structural modeling for cancer-related circRNA-miRNA interaction prediction
Mengmeng Wei, Lei Wang 0121, Xiao-Rui Su 0001, Bo-Wei Zhao, Zhu-Hong You |
Pattern Recognit. | 5 |
| 2026 | MuGNet-CMI: Multi-Head Hybrid Graph Neural Network for Predicting circRNA-miRNA Interactions With Global High-Order and Local Low-Order InformationabstractCircular RNAs (circRNAs) are non-coding RNA molecules that play a crucial role in regulating genes and contributing to disease progression. CircRNAs can function as sponges for microRNAs (miRNAs), thereby regulating gene expression and influencing disease outcomes. Identifying associations between circRNAs and miRNAs through computational methods enhances the understanding of complex disease mechanisms and offers a reliable tool for pre-selecting candidates for experimental validation. Existing models, however, are limited in their ability to capture either global or local node information, the prediction of circRNA and miRNA interactions is still challenging. In order to effectively deal with this problem, we propose a novel framework for predicting circRNA-miRNA interactions (CMIs), known as MuGNet-CMI, which leverages multi-head hybrid graph neural network and global high-order and local low-order information. The model employs the MetaPath2Vec algorithm to generate high-quality node embeddings within the circRNA-miRNA heterogeneous matrix. The multi-head dynamic attention mechanism, combined with GraphSAGE, is incorporated to efficiently capture both global high-order and local low-order node information. Additionally, we integrate neural aggregators into the multi-head dynamic attention mechanism to aggregate feature information from the captured nodes. Validation using three real datasets demonstrates that MuGNet-CMI delivers good performance in predicting CMIs, offering valuable insights to guide experimental research in gene regulation. Lei Wang 0121, Zhu-Hong You, Xinfei Wang 0001, Mengmeng Wei, Mianshuo Lu |
IEEE Trans. Big Data | 4 |
| 2026 | scProGraph: A Cell Bagging Strategy for Cell Type Annotation With Gene Interaction-Aware Explainability
Yue-Chao Li, Hai-Ru You, Xuequn Shang 0001, Leon Wong, Zhi-an Huang, Zhu-Hong You |
IEEE Trans. Big Data | 7 |
| 2026 | scALGSL: Active Learning and Graph Structure Learning for Cell Type Annotation From Single-Cell RNA-Seq DataabstractThe breakthrough development of single-cell RNA sequencing technology enables tissue heterogeneity analysis at single-cell resolution, where accurate cell type annotation is crucial for unlocking its full potential. To address three key challenges in current annotation methods-scarce labeled data, suboptimal graph topology, and missing cell state information-we propose scALGSL, an innovative framework integrating dynamic graph optimization with active learning. Our core contributions are threefold: (1) A graph-guided active learning mechanism adaptively selects high-value training samples, significantly alleviating label scarcity; (2) A learnable graph structure optimization module dynamically refines adjacency matrices to eliminate spurious connections caused by data sparsity; (3) A novel cell state auxiliary pathway extracts critical functional features via pre-trained models to enhance type discrimination. The systematic review showed that the average accuracy and f1 of scALGSL on the cancer dataset were 0.896 and 0.771, respectively, and it showed good robustness in cross-platform tasks. Integration of cell state information substantially boosts performance, while ablation studies validate the necessity of node selection and edge optimization modules. This framework provides a scalable solution for precise cell annotation, facilitating tumor microenvironment analysis and precision medicine applications. Zhihua Du, Jia-Le Yi, Wei-Lin Hu, Jianqiang Li 0001, Hai-Ru You, Zhu-Hong You, Zhi-an Huang |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2026 | Fuzzy Mixture-of-Experts Aggregation for Organoid Identification With Multiscale State Space FeaturesabstractAccurate and automated identification of organoids from bright-field images is essential for enabling high-throughput drug screening and precision medicine. Organoids, as 3-D in vitro cellular models, closely recapitulate the functional and structural characteristics of their tissue or organ of origin, presenting an unprecedented opportunity for biomedical research. However, the complexity of bright-field microscopy images, including heterogeneous backgrounds and diverse organoid morphologies, poses significant challenges for existing computational methods, often hindering robust feature extraction and high-throughput analysis. To address these issues at the intersection of computational vision and organoid biology, we propose FEMSSorg, a novel organoid recognition framework designed to adaptively aggregate multiscale scan-selected state space features through a fuzzy mixture-of-experts (FuzzyMoE) scoring mechanism. FEMSSorg introduces a fuzzy expert soft routing mechanism (fuzzy route), implemented via Fuzzy C-Means-based soft routing assignments, forming a new class of fuzzy MoE that leverages fuzzy expert clustering scores to dynamically integrate local (LocalSS) and global (GlobalSS) state space features. This approach enables effective balancing of global pixel dependencies and local texture information, thereby substantially reducing background interference and image noise in bright-field images and improving the accuracy of organoid identification. Furthermore, we incorporate a Dual Downsampling Adaptive Pooling Feature Fusion module, which combines original backbone features with parallel downsampled features and utilizes content-aware pooling for adaptive multilevel and multiscale feature fusion. Experimental results on multiclass organoid bright-field image datasets demonstrate that FEMSSorg achieves state-of-the-art performance in both organoid detection and morphological texture classification, highlighting its value as a robust computational tool for advancing real-time, high-throughput organoid research. Pengwei Hu 0001, Thomas Herget, Feng Tan 0002, Jun Zhang 0003, Lun Hu, Zhu-Hong You, Xin Luo 0001 |
IEEE Trans. Fuzzy Syst. | 9 |
| 2026 | MRGCDDI: Multi-Relation Graph Contrastive Learning Without Data Augmentation for Drug-Drug Interaction Events PredictionabstractPredicting drug-drug interactions (DDIs) is a significant concern in the field of deep learning. It can effectively reduce potential adverse consequences and improve therapeutic safety. Graph neural network (GNN)-based models have made satisfactory progress in DDI event prediction. However, most existing models overlook crucial drug structure and interaction information, which is necessary for accurate DDI event prediction. To tackle this issue, we introduce a new method called MRGCDDI. This approach employs contrastive learning, but unlike conventional methods, it does not require data augmentation, thereby avoiding additional noise. MRGCDDI maintains the semantics of the graphical data during encoder perturbation through a simple yet effective contrastive learning approach, without the need for manual trial and error, tedious searching, or expensive domain knowledge to select enhancements. The approach presented in this study effectively integrates drug features extracted from drug molecular graphs and information from multi-relational drug-drug interaction (DDI) networks. Extensive experimental results demonstrate that MRGCDDI outperforms state-of-the-art methods on both datasets. Specifically, on Deng's dataset, MRGCDDI achieves an average increase of 4.33% in accuracy, 11.57% in Macro-F1, 10.97% in Macro-Recall, and 10.64% in Macro-Precision. Similarly, on Ryu's dataset, the model shows improvements with an average increase of 2.42% in accuracy, 3.86% in Macro-F1, 3.49% in Macro-Recall, and 2.75% in Macro-Precision. Yu Li 0030, Lin-Xuan Hou, Zhu-Hong You, Chenggang Mi 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2026 | A Dynamic Multi-Scale Hypergraph Learning Framework Driven by Features and Structures for ceRNA-Disease Association PredictionabstractCompetitive endogenous RNA (ceRNA) networks are pivotal for uncovering disease molecular mechanisms. Graph representation learning is a cornerstone for modeling biological regulatory networks and predicting disease-related biomarkers. However, current methods face challenges: traditional graph neural network (GNN) rely on low-order graph structures, which struggle to capture high-order molecular interactions, resulting in topological information loss; shallow GNN fail to model long-range dependencies, while deep architectures suffer from over-smoothing, limiting complex regulatory expression; static embeddings overlook dynamic molecular interactions, reducing biomarker accuracy. These limitations highlight the need for advanced graph learning frameworks. To address these challenges, we propose DMHLF, a Dynamic Multi-scale Hypergraph Learning Framework for predicting disease-associated ceRNA biomarkers. The framework first integrates multiple regulatory relationships among miRNAs, lncRNAs, circRNAs, mRNAs, and diseases to construct disease-specific ceRNA regulatory networks, capturing local and global regulatory patterns through multi-Hop hyperedges. Subsequently, we devise a Hypergraph-Weighted Dynamic Random Walk (HEDRW) method to dynamically extract node meta-embeddings that encode high-order regulatory information. Concurrently, we extend Eigen-GNN spectral analysis to hypergraph structures, incorporating a residual-enhanced hypergraph neural network to preserve the global topological properties of shallow hypergraphs. Finally, a cross-scale attention mechanism aligns and fuses multi-scale features to generate high-quality node embeddings for disease-ceRNA association prediction. Experiments on diverse datasets demonstrate that DMHLF significantly outperforms existing methods. Case study further validates the framework's efficacy in identifying disease-related ceRNA biomarkers, providing a reliable predictive tool for biomedical research. Xinfei Wang 0001, Lan Huang 0002, Yan Wang 0028, Renchu Guan, Zhu-Hong You, Fengfeng Zhou, Yu-Qing Li |
IEEE J. Biomed. Health Informatics | 5 |
| 2026 | Graph-Based Prediction of miRNA-Drug Associations With Multisource Information and Metapath Enhancement MatricesabstractRecent studies have demonstrated that miRNA expression dysregulation is closely related to the occurrence of various diseases; thus, miRNA-based drug development strategies have received increasing research interest. Most existing computational methods focus on the attribute information of individual nodes and are limited to the direct associations between nodes, thereby ignoring the complex associations inherent in the network. This limitation may lead to the loss of key potential information, which impacts the prediction accuracy. To address these issues, we propose a multisource information fusion and metapath enhancement matrix based graph autoencoder (MSMP-GAE) to predict the potential associations between miRNAs and drugs. The proposed MSMP-GAE model comprises a metapath instance extraction module, a metapath feature enhanced encoder module, a weighted feature fusion module, and a graph autoencoder. First, we construct an miRNA-drug heterogeneous network using experimentally validated miRNA-drug interactions and integrate various miRNA and drug features into an initial feature matrix to comprehensively represent their intrinsic property information. Then, we extract metapath instances from the interaction network, generate multiple metapath enhancement matrices, and fuse them with the initial feature matrix to generate high-quality node feature embeddings. Finally, we employ the graph autoencoder for fivefold cross-validation on a public dataset and test it on an independent test set. Experimental results demonstrate that the proposed MSMP-GAE model obtained an area under the curve (AUC) and AUPR values of 98.61% and 98.23%, respectively, which is considerably better than the several state-of-the-art methods. This highlights the importance of the higher-order complex associations between nodes in the miRNA-drug association (MDA) prediction task and provides a new method and approach to advance MDA prediction. Ming-Yang Wu, Pengwei Hu 0001, Zhu-Hong You, Jun Zhang 0003, Lun Hu, Xin Luo 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2026 | scBIT: Integrating Single-Cell Transcriptomic Data Into fMRI-Based Prediction for Alzheimer's Disease DiagnosisabstractFunctional MRI (fMRI) and single-cell transcriptomics are pivotal in Alzheimer's disease (AD) research, each providing unique insights into neural function and molecular mechanisms. However, integrating these complementary modalities remains largely unexplored. Here, we introduce scBIT, a novel method for enhancing AD prediction by combining fMRI with single-nucleus RNA (snRNA). scBIT leverages snRNA as an auxiliary modality, significantly improving fMRI-based prediction models and providing comprehensive interpretability. It employs a sampling strategy to segment snRNA data into cell-type-specific gene networks and utilizes a self-explainable graph neural network to extract critical subgraphs. Additionally, we use demographic and genetic similarities to pair snRNA and fMRI data across individuals, enabling robust cross-modal learning. Extensive experiments validate scBIT's effectiveness in revealing intricate brain region-gene associations and enhancing diagnostic prediction accuracy. By advancing brain imaging transcriptomics to the single-cell level, scBIT sheds new light on biomarker discovery in AD research. Experimental results show that incorporating snRNA data into the scBIT model significantly boosts accuracy, improving binary classification by 3.39% and five-class classification by 26.59%. The codes were implemented in Python and have been released on GitHub (https://github.com/77YQ77/scBIT) and Zenodo (https://zenodo.org/records/11599030) with detailed instructions. Yao Hu 0001, Yue-Chao Li, Xiyue Cao, Kay Chen Tan, Zhu-Hong You, Zhi-an Huang |
IEEE Trans. Medical Imaging | 7 |
| 2025 | MTGCDA: Enabling Accurate CircRNA-Disease Association Prediction Through Transformer-Guided Multi-Source Graph LearningabstractCircular RNA (circRNA) has a stable structure and tissue-specific expression, which is of great value in the diagnosis and treatment of diseases. However, complex biological relationships and heterogeneous data hinder the prediction of circRNA-disease associations, resulting in challenges such as weak semantic representation and information loss. To address this problem, we propose a transformer-based multi-source heterogeneous graph model MTGCDA. The model first aggregates various biological data sources related to circRNA and disease to construct a heterogeneous graph containing various types of nodes and relationships. By applying a specialized heterogeneous graph neural network, the unique structural and contextual properties of different biological entities are captured. Subsequently, the circular RNA and disease node embeddings derived from multi-layer heterogeneous graph convolutional networks are combined to form a comprehensive joint representation. The fused embeddings are then processed by a CatBoost classifier to accurately estimate the likelihood of potential associations. Experiments on the CircR2Disease dataset show that MTGCDA achieves an AUC of 0.9756, outperforming existing methods. In addition, 9 of the 10 best predictions have been validated by literature, demonstrating the effectiveness and biological relevance of the model. Si-Zhe Liang, Lei Wang 0232, Zhu-Hong You, Tailong Shi 0002 |
BIBM | 3 |
| 2025 | Predicting MiRNA-MRNA Interactions via Multi-Scale Feature Integration and Dual-Layer Graph Attention NetworksabstractThe prediction of miRNA-mRNA interactions is fundamental to elucidating gene regulatory mechanisms and disease pathogenesis. This study proposes GMLA, a computational framework for this predictive task. The architecture first utilizes an autoencoder to derive compressed, low-dimensional feature embeddings for miRNAs and mRNAs. These embeddings then populate a heterogeneous graph, where a dual-layer Graph Attention Network (DLGAT), augmented with residual connections, is employed to capture intricate topological dependencies. A Jumping Knowledge Network (JKNet) that leverages a multi-head self-attention mechanism aggregates these layer-specific representations, enhancing the expressive capacity of the model. The efficacy of the model was systematically evaluated across several benchmark datasets characterized by diverse scales and distributions. On the principal MTIS-10317 dataset, GMLA yielded an AUC of 0.8867 and an AUPR of 0.8865. Ablation studies and parameter sensitivity analyses validated the functional contribution of each architectural component and delineated optimal hyperparameter configurations. Case studies involving the PTEN gene and miR-21-5p were conducted to evaluate the utility of the model in a practical setting, substantiating its capacity to identify biologically relevant interactions. Tailong Shi 0002, Lei Wang 0232, Zhu-Hong You, Si-Zhe Liang |
BIBM | 3 |
| 2025 | Unsupervised Cell Clustering in Single-Cell RNA Sequencing Data Using a Multi-Graph TransformerabstractSingle-cell RNA sequencing (scRNA-seq) technology reveals cellular heterogeneity and functional diversity, but its high dimensionality and sparsity pose challenges for analysis. Cell clustering is a crucial task in scRNA-seq data analysis. While supervised methods rely on extensive annotations, traditional unsupervised clustering methods often ignore intercellular relationships between cells. Recent works have shown that pathway-aware, multi-view graph constructions improve robustness and annotation accuracy across platforms and tissues [1] [2]. We propose scMGTC, a novel unsupervised multi-graph transformer-based cell clustering model. scMGTC integrates biological prior knowledge into representation learning by extracting gene sets from KEGG pathways and constructing distinct cell-cell graphs for each pathway. Extensive analysis conducted on six real scRNA-seq datasets demonstrates that scMGTC achieves promising performance in scRNA-seq clustering. Chu-Xuan Zhang, Yue-Chao Li, Zhu-Hong You |
BIBM | 4 |
| 2025 | TriM-DTA: a Tri-Modal Fusion Framework for Drug-Target Binding Affinity PredictionabstractThe process of discovering new therapeutic drugs is often time-consuming and resource-intensive, making the precise estimation of the binding strength between drug candidates and their target proteins an essential step in modern computational pharmacology. Traditional methods mostly rely on molecular data from a single modality such as sequence and structure, making it difficult to effectively obtain the spatial and biochemical characteristics of complementarity in molecular interactions. In this work, we propose TriM-DTA, a tri-modal information fusion framework to accurate predict drug-target binding affinity that integrates sequence features, topological graphs, and geometric structures of both drugs and targets. Through a dedicated encoder for each modality and a hierarchical fusion scheme, our model achieves a more holistic understanding of drug-target complexes, as evidenced by its competitive performance on benchmark datasets. Ablation studies confirm the distinct contribution of each module to overall performance. Han-Wu Zhu, Zhu-Hong You |
BIBM | 3 |
| 2025 | Bridging Knowledge Gaps: Fine-Tuned RAG Frameworks for Biomedical Evidence-Based Question Answering
Xiang-Yun Wang, Zhu-Hong You, Yu Li 0030, Yun-Hui Yan, Yue-Chao Li |
ICIC (23) | 2 |
| 2025 | CALM-AcPEP: Predicting Anticancer Peptides Using Cross-Attention and Pre-Trained Language Model
Xinke Zhan, Tiantao Liu, Pratiti Bhadra, Zhu-Hong You, Shirley W. I. Siu |
ICIC (26) | 5 |
| 2025 | Predicting MiRNA-Disease Associations Using Chebyshev Graph Convolution and Graph
Lihao Zhou, Ru Nie, Zhu-Hong You |
ICIC (25) | 5 |
| 2025 | Prediction of Budd-Chiari syndrome based on attention mechanisms of high-risk factors in multi-hop graph learning
Mengmeng Wei, Bo-Wei Zhao, Maoheng Zu, Qingqiao Zhang, Zhu-Hong You |
Sci. China Inf. Sci. | 9 |
| 2025 | Regulation-aware graph learning for drug repositioning over heterogeneous biological network
Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Dongxu Li 0002, Pengwei Hu 0001, Zhu-Hong You, Xin Luo 0001, Lun Hu |
Inf. Sci. | 7 |
| 2025 | scExGraph: Explainable graph neural network for predicting tumor environment components with single-cell sequencing data
Zhihua Du, Jiale Yi, Jianqiang Li 0001, Hai-Ru You, Zhu-Hong You, Zhi-an Huang |
Knowl. Based Syst. | 5 |
| 2025 | Fine-Tuning SAM for Forward-Looking Sonar With Collaborative Prompts and EmbeddingabstractThe Segment Anything Model (SAM) represents a significant advancement in semantic segmentation, particularly for natural images, but encounters notable limitations when applied to forward-looking sonar (FLS) images. The primary challenges lie in the inherent boundary ambiguity of FLS images, which complicates the use of prompt strategies for accurate boundary delineation, and the lack of effective interaction between prompts and image features. In this letter, we introduce a collaborative prompting strategy to address these issues by generating dense prompt embeddings and sonar tokens that focus on contour and boundary features, thereby replacing the original dense prompt embedding and IoU token. To further enhance segmentation, we employ embedding compensation techniques based on Mamba and KAN, which increase boundary information to image embedings and improve the fusion of prompts within image embeddings. We conducted comprehensive experiments, including comparative analyses and ablation studies, to validate the superiority of our proposed approach. Results show that our method significantly improves segmentation performance for FLS images, effectively addressing boundary ambiguity and optimizing prompt utilization. The source code and dataset will be available on https://github.com/darkseid-arch/FLSSAM. Zhen Wang 0020, Nan Xu 0008, Zhu-Hong You |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | Hypergraph representation learning for identifying circRNA-disease associations
Yang Li 0111, Xuegang Hu, Pei-Pei Li 0001, Lei Wang 0121, Zhu-Hong You |
Pattern Recognit. | 5 |
| 2025 | scTECTA: Asymmetric Deep Transfer Learning for Cross-Patient Tumor Microenvironment Single-Cell AnnotationabstractCellular heterogeneity and dynamic interactions within the tumor microenvironment are critical drivers of cancer initiation and progression. Single-cell RNA sequencing, with its high-resolution capabilities, has significantly advanced the study of cellular heterogeneity in the tumor microenvironment. However, existing single-cell annotation methods are limited by data sparsity, biological heterogeneity, and batch effects, which hinder their broader application in this context. To address this, we propose scTECTA, an innovative graph neural network-based method that employs transfer learning to seamlessly transfer cell-type annotation knowledge from a well-annotated source domain to an unannotated target domain. This approach leverages graph domain adaptation, integrating novel asymmetric neural network architecture and domain-adversarial learning framework. By harnessing the generalization capabilities of graph convolutional network to correct distribution shifts and employing adversarial training to further align expression profiles across batches, scTECTA substantially enhances predictive precision and robustness. We performed a systematic evaluation across multiple datasets from diverse sources, encompassing six cancer types from 34 patients, to compare the cell-type classification performance of scTECTA against 10 benchmark methods. The results demonstrate that scTECTA markedly outperforms benchmark methods in cell-type classification and exhibits robust batch-effect correction, establishing it as an efficient and powerful tool for tumor microenvironment cell-type annotation. Zi-Yi Zeng, Xiyue Cao, Yue-Chao Li, Hai-Ru You, Zhu-Hong You |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2025 | Vision Foundation Model-Driven Multiscale Expert Tuning for Multimodal Remote Sensing Semantic SegmentationabstractMultimodal remote sensing semantic segmentation based on Optical and Digital Surface Model (Opt-DSM) data is pivotal for comprehensive scene interpretation. However, prevailing methodologies often lack a unified vision foundation model and encounter significant challenges in bridging modality gaps and achieving effective feature fusion. Conventional models, such as the Segment Anything Model (SAM), exhibit inherent limitations when addressing the unique complexities of multimodal remote sensing, particularly in managing cross-modal discrepancies and intricate surface structures. In this study, we present VF-MET (Vision Foundation Model-Driven Multi-Scale Expert Tuning), an innovative framework meticulously tailored for Opt-DSM semantic segmentation tasks. VF-MET incorporates an adaptive Multi-Scale Expert Tuning (AMET) strategy, which substantially enhances the feature extraction capabilities of vision foundation models. This enables the robust capture of cross-scale and morphologically irregular objects, while simultaneously preserving superior generalization ability. To further address the segmentation of densely distributed and weakly correlated regions, we propose a collaborative Box-Point Prompt Mechanism (CBPM), which significantly improves spatial localization and contextual discrimination. Moreover, we introduce a Two-Stage Mask Decoder (TSMD) that facilitates efficient multimodal feature fusion and augments contextual understanding. Extensive experiments conducted on public Opt-DSM benchmark datasets unequivocally demonstrate that VF-MET achieves state-of-the-art performance. Comprehensive ablation studies further substantiate the indispensable contributions of each constituent module within the proposed architecture. The source code and datasets are publicly accessible at https://github.com/NWPUFranklee/VF-MET.git. Zhen Wang 0020, Nan Xu 0008, Zhu-Hong You, De-Shuang Huang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | FDMamba: Frequency-Driven Dual-Branch Mamba Network for Road Extraction From Remote Sensing ImagesabstractRoad extraction from remote sensing imagery is crucial for a variety of applications, including transportation monitoring, disaster response, and urban planning. However, existing methods often fail to accurately delineate sparse, curvilinear, and boundary-blurred road structures in high-resolution images, leading to incomplete detail preservation and inadequate contextual understanding. To address these challenges, we propose a novel Frequency-Driven Dual-Branch Mamba Network (FDMamba) for precise road extraction from remote sensing imagery. The proposed FDMamba integrates frequency-aware modeling with a dual-branch architecture, enabling collaborative learning of fine-grained edge details and global spatial dependencies. Specifically, FDMamba comprises three key modules: a Fourier Reconstruction Attention Mechanism (FRAM) to enhance high-frequency boundary information and low-frequency structural representation; a Rotation-aware Mamba Module (RAMamba) that leverages multi-path state space modeling for robust directional perception of road structures; and a Phase-guided Feature Fusion Module (PFFM) for effective cross-scale alignment and fusion of high- and low-frequency features. Furthermore, to mitigate the issue of blurred or ambiguous boundaries, we introduce a hybrid loss function that combines binary cross-entropy, focal loss, and frequency-aware loss, explicitly guiding the model to focus on edge structure and multi-frequency complementary information. Extensive experiments on three benchmark datasets, CHN6-CUG, DeepGlobe, and Massachusetts, demonstrate that FDMamba consistently outperforms state-of-the-art methods in terms of F1-score and IoU, achieving superior boundary clarity and structural continuity while preserving overall geometric integrity. The code is available at https://github.com/darkseid-arch/RE-FDMamba. Zhen Wang 0020, Shen-Ao Yuan, Nan Xu 0008, Zhu-Hong You, De-Shuang Huang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | UAVSeg: Dual-Encoder Cross-Scale Attention Network for UAV Images' Semantic SegmentationabstractBenefiting from the powerful feature extraction and feature correlation modeling capabilities of convolutional neural networks (CNNs) and Transformer models, these techniques have been widely used in unmanned aerial vehicle (UAV) aerial image semantic segmentation tasks. However, the ground objects in aerial images contain feature information with different scales, and existing methods directly cascade low-level visual features and high-level semantic features without processing, resulting in low semantic segmentation precision. To address these challenges, we propose a dual-encoder cross-scale attention network, which efficiently extracts local and global context information from aerial images and performs fine-grained fusion of multiscale features to improve semantic segmentation performance. First, we introduce the dual-CNN-Transformer encoder, which embeds the scan-focus window Transformer (SFWT) into CNNs as an auxiliary encoder to supplement the local feature information lost in the global context information extraction process. Second, the cross-scale lightweight integration (CSLI) module is designed, which uses a light dot-product attention mechanism (DPAM) to fusion multiscale features and reduce model calculation parameters. Finally, the linear multilayer perceptron (LMLP) is used to restore the feature map resolution while expanding the deconvolution receptive field. To validate the effectiveness of the proposed method, we conducted extensive experiments on real aerial scene datasets, including UAVid, Urban Drone, and AeroScapes. The experimental results show that our method achieves state-of-the-art performance while maintaining superior real-time efficiency. Implementation codes will be available athttps://github.com/darkseid-arch/UAVSeg. Zhen Wang 0020, Zhu-Hong You, Nan Xu 0008, Chuanlei Zhang, De-Shuang Huang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Knowledge Graph Neural Network With Spatial-Aware Capsule for Drug-Drug Interaction PredictionabstractUncovering novel drug-drug interactions (DDIs) plays a pivotal role in advancing drug development and improving clinical treatment. The outstanding effectiveness of graph neural networks (GNNs) has garnered significant interest in the field of DDI prediction. Consequently, there has been a notable surge in the development of network-based computational approaches for predicting DDIs. However, current approaches face limitations in capturing the spatial relationships between neighboring nodes and their higher-level features during the aggregation of neighbor representations. To address this issue, this study introduces a novel model, KGCNN, designed to comprehensively tackle DDI prediction tasks by considering spatial relationships between molecules within the biomedical knowledge graph (BKG). KGCNN is built upon a message-passing GNN framework, consisting of propagation and aggregation. In the context of the BKG, KGCNN governs the propagation of information based on semantic relationships, which determine the flow and exchange of information between different molecules. In contrast to traditional linear aggregators, KGCNN introduces a spatial-aware capsule aggregator, which effectively captures the spatial relationships among neighboring molecules and their higher-level features within the graph structure. The ultimate goal is to leverage these learned drug representations to predict potential DDIs. To evaluate the effectiveness of KGCNN, it undergoes testing on two datasets. Extensive experimental results demonstrate its superiority in DDI predictions and quantified performance. Xiao-Rui Su 0001, Bo-Wei Zhao, Jun Zhang 0003, Pengwei Hu 0001, Zhu-Hong You, Lun Hu |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Integrating Transformer and Graph Attention Network for circRNA-miRNA Interaction PredictionabstractCircRNA-miRNA interaction (CMI) plays a crucial role in the gene regulatory network of the cell. Numerous experiments have shown that abnormalities in CMI can impact molecular functions and physiological processes, leading to the occurrence of specific diseases. Current computational models for predicting CMI typically focus on local molecular entity relationships, thereby neglecting inherent molecular attributes and global structural information. To address these limitations, we propose a multi-feature fusion prediction model based on the transformer and graph attention network, named EGATCMI. Specifically, EGATCMI combines the transformer architecture with Word2vec to pre-train the sequence of circRNA and miRNA, capturing their sequence feature representation and sequence similarity. By leveraging the self-attention mechanism, EGATCMI extracts global structural feature from the CMI network. EGATCMI effectively integrates the obtained multi-feature for prediction, achieving AUC values of 0.9106 and 0.9470 on the CMI-9905 and CircBank datasets, respectively, outperforming existing methods. In case studies that the prediction of interactions between three miRNAs that are closely related to diseases and circRNAs, 8 out of 10 pairs were accurately predicted and validated. Extensive experimental results demonstrate the potential of EGATCMI as a reliable tool for candidate screening in biological investigations. Mengmeng Wei, Lei Wang 0121, Bo-Wei Zhao, Xiao-Rui Su 0001, Zhu-Hong You, De-Shuang Huang |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Equivariant 3D-Conditional Diffusion Model for De Novo Drug DesignabstractDe novo drug design speeds up drug discovery, mitigating its time and cost burdens with advanced computational methods. Previous work either insufficiently utilized the 3D geometric structure of the target proteins, or generated ligands in an order that was inconsistent with real physics. Here we propose an equivariant 3D-conditional diffusion model, named DiffFBDD, for generating new pharmaceutical compounds based on 3D geometric information of specific target protein pockets. DiffFBDD overcomes the underutilization of geometric information by integrating full atomic information of pockets to backbone atoms using an equivariant graph neural network. Moreover, we develop a diffusion approach to generate drugs by generating ligand fragments for specific protein pockets, which requires fewer computational resources and less generation time (65.98% 96.10% lower). DiffFBDD offers better performance than state-of-the-art models in generating ligands with strong binding affinity to specific protein pockets, while maintaining high validity, uniqueness, and novelty, with clear potential for exploring the drug-like chemical space. Zhu-Hong You |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Local-Global Structure-Aware Geometric Equivariant Graph Representation Learning for Predicting Protein-Ligand Binding AffinityabstractPredicting protein-ligand binding affinities is a critical problem in drug discovery and design. A majority of existing methods fail to accurately characterize and exploit the geometrically invariant structures of protein-ligand complexes for predicting binding affinities. In this study, we propose Geo-protein-ligand binding affinity (PLA), a geometric equivariant graph representation learning framework with local-global structure awareness, to predict binding affinity by capturing the geometric information of protein-ligand complexes. Specifically, the local structural information of 3-D protein-ligand complexes is extracted by using an equivariant graph neural network (EGNN), which iteratively updates node representations while preserving the equivariance of coordinate transformations. Meanwhile, a graph transformer is utilized to capture long-range interactions among atoms, offering a global view that adaptively focuses on complex regions with a significant impact on binding affinities. Furthermore, the multiscale information from the two channels is integrated to enhance the predictive capability of the model. Extensive experimental studies on two benchmark datasets confirm the superior performance of Geo-PLA. Moreover, the visual interpretation of the learned protein-ligand complexes further indicates that our model offers valuable biological insights for virtual screening and drug repositioning. Zhu-Hong You, Xuequn Shang 0001, Lei Wang 0121, Zhen Wang 0020 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Toward Multilabel Classification for Multiple Disease Prediction Using Gut Microbiota ProfilesabstractAdvancements in high-throughput technologies have yielded large-scale human gut microbiota profiles, sparking considerable interest in exploring the relationship between the gut microbiome and complex human diseases. Through extracting and integrating knowledge from complex microbiome data, existing machine learning (ML)-based studies have demonstrated their effectiveness in the precise identification of high-risk individuals. However, these approaches struggle to address the heterogeneity and sparsity of microbial features and explore the intrinsic relatedness among human diseases. In this work, we reframe human gut microbiome-based disease detection as a multilabel classification (MLC) problem and integrate a range of innovative techniques within the proposed MLC framework, aptly named GutMLC. Specifically, the entity semantic similarity as priori knowledge is incorporated into multilabel feature selection and loss functions by capturing the shared attributes and inherent associations among diseases and microbes. To tackle the issue of label imbalance, both within and between labels, we adapt the focal loss (FL) function for MLC using debiased inverse weighting. Extensive experiment results consistently demonstrate the competitive performance of GutMLC in comparison with commonly used MLC and single-label classification (SLC) algorithms. This work seeks to unlock the potential of gut microbiota as robust biomarkers for multiple disease prediction. Zhi-an Huang, Pengwei Hu 0001, Lun Hu, Zhu-Hong You, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Dual-Channel Learning Framework for Drug-Drug Interaction Prediction via Relation-Aware Heterogeneous Graph TransformerabstractIdentifying novel drug-drug interactions (DDIs) is a crucial task in pharmacology, as the interference between pharmacological substances can pose serious medical risks. In recent years, several network-based techniques have emerged for predicting DDIs. However, they primarily focus on local structures within DDI-related networks, often overlooking the significance of indirect connections between pairwise drug nodes from a global perspective. Additionally, effectively handling heterogeneous information present in both biomedical knowledge graphs and drug molecular graphs remains a challenge for improved performance of DDI prediction. To address these limitations, we propose a Transformer-based relatIon-aware Graph rEpresentation leaRning framework (TIGER) for DDI prediction. TIGER leverages the Transformer architecture to effectively exploit the structure of heterogeneous graph, which allows it direct learning of long dependencies and high-order structures. Furthermore, TIGER incorporates a relation-aware self-attention mechanism, capturing a diverse range of semantic relations that exist between pairs of nodes in heterogeneous graph. In addition to these advancements, TIGER enhances predictive accuracy by modeling DDI prediction task using a dual-channel network, where drug molecular graph and biomedical knowledge graph are fed into two respective channels. By incorporating embeddings obtained at graph and node levels, TIGER can benefit from structural properties of drugs as well as rich contextual information provided by biomedical knowledge graph. Extensive experiments conducted on three real-world datasets demonstrate the effectiveness of TIGER in DDI prediction. Furthermore, case studies highlight its ability to provide a deeper understanding of underlying mechanisms of DDIs. Xiao-Rui Su 0001, Pengwei Hu 0001, Zhu-Hong You, Philip S. Yu, Lun Hu |
AAAI | 3 |
| 2024 | DNMDA: Deep Non-negative Matrix Factorization with Multi-level Integration for MiRNA-Drug Interaction PredictionabstractNumerous studies have demonstrated that the interaction between miRNAs and drugs plays a pivotal role in regulating gene expression and cellular function. Therefore, predicting these interactions is crucial for the development of novel drugs and personalized therapies. Existing methods for predicting miRNA-drug interactions often fail to leverage the full spectrum of molecular and biological features and overlook complex high-dimensional patterns. Deep non-negative matrix factorization (DNMF) addresses these limitations by extracting higher-level representations, thereby enhancing prediction accuracy and robustness. Building on this, this paper proposes a model called DNMDA. In this model, we integrate multiple similarity networks for both miRNAs and drugs and then extract their features through three key modules. Moreover, autoencoders are used to combine various feature sets, allowing for the capture of complementary information and enhancing the model’s capacity for making precise and reliable predictions. The resulting features are consolidated into a unified feature vector for each miRNA-drug pair. Ultimately, these feature vectors and their associated labels are provided to the classifier for training. To verify the predictions, a five-fold cross-validation was conducted. The five-fold cross-validation demonstrated a clear advantage in DNMDA’s metrics, underscoring its reliability in predicting potential miRNA-drug interactions. This claim is further supported by the predictive results section in the paper, offering concrete evidence of DNMDA’s efficacy in this field. Yujie Qi, Zhu-Hong You, Zimai Zhang, Lun Hu, Xi Zhou 0007, Pengwei Hu 0001 |
BIBM | 4 |
| 2024 | EELMCDA: Combining evolutionary ensemble learning with matrix feature decomposition for predicting circRNA-disease associationsabstractRecent studies have indicated that circular RNAs (circRNAs) play a significant role in the diagnosis and treatment of disease. However, the prediction of associations between circRNAs and diseases using conventional biological methods is constrained by numerous factors. In this study, we proposed a novel computational model called EELMCDA that combines evolutionary ensemble learning (EEL) approach and matrix feature decomposition method to predict potential circRNA-disease associations. The model firstly integrates circRNA function information, disease semantic information, and circRNA and disease gaussian interaction profile kernel (GIPK) information into an integrated matrix and constructed the corresponding feature matrix, then uses the matrix feature decomposition algorithm to obtain its important feature, and finally adopted evolutionary ensemble learning module to predict circRNA-disease associations. The average accuracy of the EELMCDA model by 5-fold cross-validation on CircR2Disease, CircAtlasv2.0, Circ2Disease, and CircRNADisease datasets were 92.40%, 92.90%, 88.91%, and 90.74%, respectively. Moreover, in case studies, the 21 of the top 30 circRNA-disease pairs with the highest EELMCDA scores were validated in recent literatures. These results further demonstrate the effectiveness of EELMCDA in predicting circRNA-disease associations. Zheng Wang 0065, Lei Wang 0030, Zhu-Hong You, Lei Wang 0121, Yang Li 0111 |
BIBM | 3 |
| 2024 | Drug-Drug Interaction Prediction Based on Probability Transfer Multi-modal Feature Representation LearningabstractIn drug discovery and combination therapy, drug-drug interactions can lead to adverse reactions, affecting not only disease treatment but also risking the market withdrawal of new drugs. Traditional experiments in vitro and in vivo are labor-intensive and time-consuming for identifying potential DDIs. Although existing computational methods offer new perspectives for DDIs identification, they still have limitations. This paper innovatively uses the probability transfer matrix combined with Stacked Denoising Autoencoder to propose a model named MultiPT-DDI to calculate the correlation of edge nodes in the adjacency matrix, which effectively learns the multi-level representation of nodes and mitigates the probabilistic bias of the edge nodes in the sparse matrices and the noise of the original data. Specifically, the method first samples multiple bipartite graph networks using random surfing thus obtaining multiple probabilistic transfer matrices. Subsequently, multiple denoising autoencoder modules are employed for layer-wise unsupervised pre-training of the network. Finally, we infer the relationships between drug pairs using the Random Forest algorithm. The experiment obtains the AUC score of 0.9433 and the AUPR score of 0.9372 in the 5-fold cross-validation, significantly outperforming existing models. In the case studies, 26 of the top 30 drug pairs with the highest scores were validated. The empirical evidence indicates that MultiPT-DDI is an effective complementary model for predicting potential DDIs, providing a reliable reference for traditional experimental methods. Chang-Qin Yu, Mengmeng Wei, Zhu-Hong You |
BIBM | 6 |
| 2024 | Lightweight Coal Flow Foreign Object Detection Algorithm
Ru Nie, Xiaobing Shen, Zhengwei Li 0001, Yanxia Jiang, Hongmei Liao, Zhu-Hong You |
ICIC (3) | 6 |
| 2024 | LAROD-HD: Low-Cost Adaptive Real-Time Object Detection for High-Resolution Video Surveillance
Yanglin Pu, Chen Gao 0005, Bo Liu 0112, Si Liu 0001, Junhua Xiao, Shiliang Pu, Zhu-Hong You |
ICIC (8) | 8 |
| 2024 | Context-Aware Relative Distinctive Feature Learning for Person Re-identification
Hangyuan Yang, Yanglin Pu, Zhu-Hong You |
ICIC (8) | 5 |
| 2024 | Likelihood-based feature representation learning combined with neighborhood information for predicting circRNA-miRNA associationsabstractConnections between circular RNAs (circRNAs) and microRNAs (miRNAs) assume a pivotal position in the onset, evolution, diagnosis and treatment of diseases and tumors. Selecting the most potential circRNA-related miRNAs and taking advantage of them as the biological markers or drug targets could be conducive to dealing with complex human diseases through preventive strategies, diagnostic procedures and therapeutic approaches. Compared to traditional biological experiments, leveraging computational models to integrate diverse biological data in order to infer potential associations proves to be a more efficient and cost-effective approach. This paper developed a model of Convolutional Autoencoder for CircRNA-MiRNA Associations (CA-CMA) prediction. Initially, this model merged the natural language characteristics of the circRNA and miRNA sequence with the features of circRNA-miRNA interactions. Subsequently, it utilized all circRNA-miRNA pairs to construct a molecular association network, which was then fine-tuned by labeled samples to optimize the network parameters. Finally, the prediction outcome is obtained by utilizing the deep neural networks classifier. This model innovatively combines the likelihood objective that preserves the neighborhood through optimization, to learn the continuous feature representation of words and preserve the spatial information of two-dimensional signals. During the process of 5-fold cross-validation, CA-CMA exhibited exceptional performance compared to numerous prior computational approaches, as evidenced by its mean area under the receiver operating characteristic curve of 0.9138 and a minimal SD of 0.0024. Furthermore, recent literature has confirmed the accuracy of 25 out of the top 30 circRNA-miRNA pairs identified with the highest CA-CMA scores during case studies. The results of these experiments highlight the robustness and versatility of our model. Lu-Xiang Guo, Lei Wang 0121, Zhu-Hong You, Meng-Lei Hu, Bo-Wei Zhao, Yang Li 0111 |
Briefings Bioinform. | 3 |
| 2024 | Biolinguistic graph fusion model for circRNA-miRNA association predictionabstractEmerging clinical evidence suggests that sophisticated associations with circular ribonucleic acids (RNAs) (circRNAs) and microRNAs (miRNAs) are a critical regulatory factor of various pathological processes and play a critical role in most intricate human diseases. Nonetheless, the above correlations via wet experiments are error-prone and labor-intensive, and the underlying novel circRNA-miRNA association (CMA) has been validated by numerous existing computational methods that rely only on single correlation data. Considering the inadequacy of existing machine learning models, we propose a new model named BGF-CMAP, which combines the gradient boosting decision tree with natural language processing and graph embedding methods to infer associations between circRNAs and miRNAs. Specifically, BGF-CMAP extracts sequence attribute features and interaction behavior features by Word2vec and two homogeneous graph embedding algorithms, large-scale information network embedding and graph factorization, respectively. Multitudinous comprehensive experimental analysis revealed that BGF-CMAP successfully predicted the complex relationship between circRNAs and miRNAs with an accuracy of 82.90% and an area under receiver operating characteristic of 0.9075. Furthermore, 23 of the top 30 miRNA-associated circRNAs of the studies on data were confirmed in relevant experiences, showing that the BGF-CMAP model is superior to others. BGF-CMAP can serve as a helpful model to provide a scientific theoretical basis for the study of CMA prediction. Lu-Xiang Guo, Lei Wang 0121, Zhu-Hong You, Meng-Lei Hu, Bo-Wei Zhao, Yang Li 0111 |
Briefings Bioinform. | 3 |
| 2024 | A microbial knowledge graph-based deep learning model for predicting candidate microbes for target hostsabstractPredicting interactions between microbes and hosts plays critical roles in microbiome population genetics and microbial ecology and evolution. How to systematically characterize the sophisticated mechanisms and signal interplay between microbes and hosts is a significant challenge for global health risks. Identifying microbe-host interactions (MHIs) can not only provide helpful insights into their fundamental regulatory mechanisms, but also facilitate the development of targeted therapies for microbial infections. In recent years, computational methods have become an appealing alternative due to the high risk and cost of wet-lab experiments. Therefore, in this study, we utilized rich microbial metagenomic information to construct a novel heterogeneous microbial network (HMN)-based model named KGVHI to predict candidate microbes for target hosts. Specifically, KGVHI first built a HMN by integrating human proteins, viruses and pathogenic bacteria with their biological attributes. Then KGVHI adopted a knowledge graph embedding strategy to capture the global topological structure information of the whole network. A natural language processing algorithm is used to extract the local biological attribute information from the nodes in HMN. Finally, we combined the local and global information and fed it into a blended deep neural network (DNN) for training and prediction. Compared to state-of-the-art methods, the comprehensive experimental results show that our model can obtain excellent results on the corresponding three MHI datasets. Furthermore, we also conducted two pathogenic bacteria case studies to further indicate that KGVHI has excellent predictive capabilities for potential MHI pairs. Jie Pan 0007, Jiaoyang Yu, Zhu-Hong You, Shixu Wang, Fengzhi Ren, Xuexia Zhang, Yanmei Sun |
Briefings Bioinform. | 5 |
| 2024 | Multi-view learning framework for predicting unknown types of cancer markers via directed graph neural networks fitting regulatory networksabstractThe discovery of diagnostic and therapeutic biomarkers for complex diseases, especially cancer, has always been a central and long-term challenge in molecular association prediction research, offering promising avenues for advancing the understanding of complex diseases. To this end, researchers have developed various network-based prediction techniques targeting specific molecular associations. However, limitations imposed by reductionism and network representation learning have led existing studies to narrowly focus on high prediction efficiency within single association type, thereby glossing over the discovery of unknown types of associations. Additionally, effectively utilizing network structure to fit the interaction properties of regulatory networks and combining specific case biomarker validations remains an unresolved issue in cancer biomarker prediction methods. To overcome these limitations, we propose a multi-view learning framework, CeRVE, based on directed graph neural networks (DGNN) for predicting unknown type cancer biomarkers. CeRVE effectively extracts and integrates subgraph information through multi-view feature learning. Subsequently, CeRVE utilizes DGNN to simulate the entire regulatory network, propagating node attribute features and extracting various interaction relationships between molecules. Furthermore, CeRVE constructed a comparative analysis matrix of three cancers and adjacent normal tissues through The Cancer Genome Atlas and identified multiple types of potential cancer biomarkers through differential expression analysis of mRNA, microRNA, and long noncoding RNA. Computational testing of multiple types of biomarkers for 72 cancers demonstrates that CeRVE exhibits superior performance in cancer biomarker prediction, providing a powerful tool and insightful approach for AI-assisted disease biomarker discovery. Xinfei Wang 0001, Lan Huang 0002, Yan Wang 0028, Renchu Guan, Zhu-Hong You, Nan Sheng, Xuping Xie, Wenju Hou |
Briefings Bioinform. | 5 |
| 2024 | A multichannel graph neural network based on multisimilarity modality hypergraph contrastive learning for predicting unknown types of cancer biomarkersabstractIdentifying potential cancer biomarkers is a key task in biomedical research, providing a promising avenue for the diagnosis and treatment of human tumors and cancers. In recent years, several machine learning-based RNA-disease association prediction techniques have emerged. However, they primarily focus on modeling relationships of a single type, overlooking the importance of gaining insights into molecular behaviors from a complete regulatory network perspective and discovering biomarkers of unknown types. Furthermore, effectively handling local and global topological structural information of nodes in biological molecular regulatory graphs remains a challenge to improving biomarker prediction performance. To address these limitations, we propose a multichannel graph neural network based on multisimilarity modality hypergraph contrastive learning (MML-MGNN) for predicting unknown types of cancer biomarkers. MML-MGNN leverages multisimilarity modality hypergraph contrastive learning to delve into local associations in the regulatory network, learning diverse insights into the topological structures of multiple types of similarities, and then globally modeling the multisimilarity modalities through a multichannel graph autoencoder. By combining representations obtained from local-level associations and global-level regulatory graphs, MML-MGNN can acquire molecular feature descriptors benefiting from multitype association properties and the complete regulatory network. Experimental results on predicting three different types of cancer biomarkers demonstrate the outstanding performance of MML-MGNN. Furthermore, a case study on gastric cancer underscores the outstanding ability of MML-MGNN to gain deeper insights into molecular mechanisms in regulatory networks and prominent potential in cancer biomarker prediction. Xinfei Wang 0001, Lan Huang 0002, Yan Wang 0028, Renchu Guan, Zhu-Hong You, Nan Sheng, Xuping Xie, Qixing Yang |
Briefings Bioinform. | 5 |
| 2024 | MHESMMR: a multilevel model for predicting the regulation of miRNAs expression by small moleculesabstractAccording to the expression of miRNA in pathological processes, miRNAs can be divided into oncogenes or tumor suppressors. Prediction of the regulation relations between miRNAs and small molecules (SMs) becomes a vital goal for miRNA-target therapy. But traditional biological approaches are laborious and expensive. Thus, there is an urgent need to develop a computational model. In this study, we proposed a computational model to predict whether the regulatory relationship between miRNAs and SMs is up-regulated or down-regulated. Specifically, we first use the Large-scale Information Network Embedding (LINE) algorithm to construct the node features from the self-similarity networks, then use the General Attributed Multiplex Heterogeneous Network Embedding (GATNE) algorithm to extract the topological information from the attribute network, and finally utilize the Light Gradient Boosting Machine (LightGBM) algorithm to predict the regulatory relationship between miRNAs and SMs. In the fivefold cross-validation experiment, the average accuracies of the proposed model on the SM2miR dataset reached 79.59% and 80.37% for up-regulation pairs and down-regulation pairs, respectively. In addition, we compared our model with another published model. Moreover, in the case study for 5-FU, 7 of 10 candidate miRNAs are confirmed by related literature. Therefore, we believe that our model can promote the research of miRNA-targeted therapy. Yongjian Guan, Liping Li 0003, Zhu-Hong You, Weixiao Meng 0001, Xinfei Wang 0001, Lu-Xiang Guo |
BMC Bioinform. | 4 |
| 2024 | BEROLECMI: a novel prediction method to infer circRNA-miRNA interaction from the role definition of molecular attributes and biological networksabstractCircular RNA (CircRNA)-microRNA (miRNA) interaction (CMI) is an important model for the regulation of biological processes by non-coding RNA (ncRNA), which provides a new perspective for the study of human complex diseases. However, the existing CMI prediction models mainly rely on the nearest neighbor structure in the biological network, ignoring the molecular network topology, so it is difficult to improve the prediction performance. In this paper, we proposed a new CMI prediction method, BEROLECMI, which uses molecular sequence attributes, molecular self-similarity, and biological network topology to define the specific role feature representation for molecules to infer the new CMI. BEROLECMI effectively makes up for the lack of network topology in the CMI prediction model and achieves the highest prediction performance in three commonly used data sets. In the case study, 14 of the 15 pairs of unknown CMIs were correctly predicted. Xinfei Wang 0001, Zhu-Hong You, Yan Wang 0028, Lan Huang 0002, Yan Qiao 0002, Lei Wang 0121, Zhengwei Li 0001 |
BMC Bioinform. | 3 |
| 2024 | BioKG-CMI: a multi-source feature fusion model based on biological knowledge graph for predicting circRNA-miRNA interactions
Mengmeng Wei, Lei Wang 0121, Bo-Wei Zhao, Xiao-Rui Su 0001, Zhu-Hong You |
Sci. China Inf. Sci. | 8 |
| 2024 | Attention-Based Learning for Predicting Drug-Drug Interactions in Knowledge Graph Embedding Based on Multisource Fusion InformationabstractDrug combinations can reduce drug resistance and side effects and enable the improvement of disease treatment efficacy. Therefore, how to effectively identify drug-drug interactions (DDIs) is a challenging problem. Currently, there exist several approaches that leverage advanced representation learning and graph-based techniques for DDIs prediction. While these methods have demonstrated promising results, a limited number of approaches effectively utilize the potential of knowledge graphs (KGs), which provide information on drug attributes and multirelation among entities. In this work, we introduce a novel attention-based KGs representation learning framework. To encode drug SMILES sequence, a pretrained model is used, while molecular structure information is mapped as the initialization of nodes within the KG using a message-passing neural network. Additionally, the knowledge-aware graph attention network is employed to capture the drug and its topological neighbor representation in the KG representation module. To prevent the oversmoothing problem, the residual layer is used in the DDI prediction module. Comprehensive experiments on several datasets have demonstrated that the proposed method outperforms the state-of-the-art algorithms on the DDI prediction task across a range of evaluation metrics. It achieves an accuracy of 0.924 and an AUC of 0.9705 on the KEGG dataset and attains an ACC of 0.9777 and an AUC of 0.9959 on the OGB-biokg dataset. These experimental findings affirm that our approach is a dependable model for predicting the association of drugs. Yu Li 0030, Zhu-Hong You, Shu-Min Wang, Chenggang Mi 0001, Meineng Wang |
Int. J. Intell. Syst. | 2 |
| 2024 | A hierarchical GNN across semantic and topological domains for predicting circRNA-microRNA interactionsabstractIdentifying circRNA-microRNA interactions (CMI) is a significant biomedical issue in recent years. This problem provides insights into using circRNA as biomarkers, developing cancer therapies and producing cancer vaccines. Using computational methods for identification is a more time-efficient and cost-effective approach. In computational methods, using graphs to represent and explore the CMI is a mainstream approach. However, existing relevant methods do not achieve optimal results by utilizing both the semantic information extracted from sequences and the topological information extracted from graph structures. To address this issue, we propose HGLMALLM, a graph contrastive learning method that learns node representation crossing both the semantic domain generated via motif-aware pre-trained LLMs and the topological domain extracted from hierarchical graph structures. Our method effectively addresses the issue in existing Message Passing Neural Network (MPNN) method that edge components losing heterogeneity after multiple iterations. Moreover, this method utilizes the heterogeneity of graph which is extended from the traditional bipartite graph to heterogeneous through the semantic domain . Two commonly used datasets were partitioned based on the distribution of node degrees. Then, we benchmarked our method against existing methods. In the independent testing set evaluation, it achieved a 3 % and 1 % improvement on two datasets. Our method demonstrated the best stability in ten-fold cross-validation on the training set. A test conducted on the peripheral components reveals robust performance of our model. A dataset collected from real scenarios was used to demonstrate the strong predictive ability of our method for identifying unidentified CMI. Ji-Ren Zhou, Rui Niu, Xuequn Shang 0001, Zhu-Hong You |
Knowl. Based Syst. | 5 |
| 2024 | AMDECDA: Attention Mechanism Combined With Data Ensemble Strategy for Predicting CircRNA-Disease AssociationabstractAccumulating evidence from recent research reveals that circRNA is tightly bound to human complex disease and plays an important regulatory role in disease progression. Identifying disease-associated circRNA occupies a key role in the research of disease pathogenesis. In this study, we propose a new model AMDECDA for predicting circRNA-disease association (CDA) by combining attention mechanism and data ensemble strategy. Firstly, we fuse the heterogeneous information including circRNA Gaussian interaction profile (GIP), disease semantics and disease GIP, and then use the attention mechanism of Graph Attention Network (GAT) to focus on the critical information of data, reasonably allocate resources and extract their essential features. Finally, the ensemble deep RVFL network (edRVFL) is utilized to quickly and accurately predict CDA in the non-iterative manner of closed-form solutions. In the five-fold cross-validation experiment on the benchmark data set, AMDECDA achieves an accuracy of 93.10% with a sensitivity of 97.56% in 0.9235 AUC. In comparison with previous models, AMDECDA exhibits highly competitiveness. Furthermore, 26 of the top 30 unknown CDAs of AMDECDA predicted scores are proved by the related literature. These results indicate that AMDECDA can effectively anticipate latent CDA and provide help for further biological wet experiments. Lei Wang 0121, Leon Wong, Zhu-Hong You, De-Shuang Huang |
IEEE Trans. Big Data | 3 |
| 2024 | Predicting miRNA-Disease Associations Based on Spectral Graph Transformer With Dynamic Attention and RegularizationabstractExtensive research indicates that microRNAs (miRNAs) play a crucial role in the analysis of complex human diseases. Recently, numerous methods utilizing graph neural networks have been developed to investigate the complex relationships between miRNAs and diseases. However, these methods often face challenges in terms of overall effectiveness and are sensitive to node positioning. To address these issues, the researchers introduce DARSFormer, an advanced deep learning model that integrates dynamic attention mechanisms with a spectral graph Transformer effectively. In the DARSFormer model, a miRNA-disease heterogeneous network is constructed initially. This network undergoes spectral decomposition into eigenvalues and eigenvectors, with the eigenvalue scalars being mapped into a vector space subsequently. An orthogonal graph neural network is employed to refine the parameter matrix. The enhanced features are then input into a graph Transformer, which utilizes a dynamic attention mechanism to amalgamate features by aggregating the enhanced neighbor features of miRNA and disease nodes. A projection layer is subsequently utilized to derive the association scores between miRNAs and diseases. The performance of DARSFormer in predicting miRNA-disease associations (MDAs) is exemplary. It achieves an AUC of 94.18% in a five-fold cross-validation on the HMDD v2.0 database. Similarly, on HMDD v3.2, it records an AUC of 95.27%. Case studies involving colorectal, esophageal, and prostate tumors confirm 27, 28, and 26 of the top 30 associated miRNAs against the dbDEMC and miR2Disease databases, respectively. Zhengwei Li 0001, Ru Nie, Lei Zhang 0029, Zhu-Hong You |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | GSLCDA: An Unsupervised Deep Graph Structure Learning Method for Predicting CircRNA-Disease AssociationabstractGrowing studies reveal that Circular RNAs (circRNAs) are broadly engaged in physiological processes of cell proliferation, differentiation, aging, apoptosis, and are closely associated with the pathogenesis of numerous diseases. Clarification of the correlation among diseases and circRNAs is of great clinical importance to provide new therapeutic strategies for complex diseases. However, previous circRNA-disease association prediction methods rely excessively on the graph network, and the model performance is dramatically reduced when noisy connections occur in the graph structure. To address this problem, this paper proposes an unsupervised deep graph structure learning method GSLCDA to predict potential CDAs. Concretely, we first integrate circRNA and disease multi-source data to constitute the CDA heterogeneous network. Then the network topology is learned using the graph structure, and the original graph is enhanced in an unsupervised manner by maximize the inter information of the learned and original graphs to uncover their essential features. Finally, graph space sensitive k-nearest neighbor (KNN) algorithm is employed to search for latent CDAs. In the benchmark dataset, GSLCDA obtained 92.67% accuracy with 0.9279 AUC. GSLCDA also exhibits exceptional performance on independent datasets. Furthermore, 14, 12 and 14 of the top 16 circRNAs with the most points GSLCDA prediction scores were confirmed in the relevant literature in the breast cancer, colorectal cancer and lung cancer case studies, respectively. Such results demonstrated that GSLCDA can validly reveal underlying CDA and offer new perspectives for the diagnosis and therapy of complex human diseases. Lei Wang 0121, Zhengwei Li 0001, Zhu-Hong You, De-Shuang Huang, Leon Wong |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | MAGCDA: A Multi-Hop Attention Graph Neural Networks Method for CircRNA-Disease Association PredictionabstractWith a growing body of evidence establishing circular RNAs (circRNAs) are widely exploited in eukaryotic cells and have a significant contribution in the occurrence and development of many complex human diseases. Disease-associated circRNAs can serve as clinical diagnostic biomarkers and therapeutic targets, providing novel ideas for biopharmaceutical research. However, available computation methods for predicting circRNA-disease associations (CDAs) do not sufficiently consider the contextual information of biological network nodes, making their performance limited. In this work, we propose a multi-hop attention graph neural network-based approach MAGCDA to infer potential CDAs. Specifically, we first construct a multi-source attribute heterogeneous network of circRNAs and diseases, then use a multi-hop strategy of graph nodes to deeply aggregate node context information through attention diffusion, thus enhancing topological structure information and mining data hidden features, and finally use random forest to accurately infer potential CDAs. In the four gold standard data sets, MAGCDA achieved prediction accuracy of 92.58%, 91.42%, 83.46% and 91.12%, respectively. MAGCDA has also presented prominent achievements in ablation experiments and in comparisons with other models. Additionally, 18 and 17 potential circRNAs in top 20 predicted scores for MAGCDA prediction scores were confirmed in case studies of the complex diseases breast cancer and Almozheimer's disease, respectively. These results suggest that MAGCDA can be a practical tool to explore potential disease-associated circRNAs and provide a theoretical basis for disease diagnosis and treatment. Lei Wang 0121, Zhengwei Li 0001, Zhu-Hong You, De-Shuang Huang, Leon Wong |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Motif-Aware miRNA-Disease Association Prediction via Hierarchical Attention NetworkabstractAs post-transcriptional regulators of gene expression, micro-ribonucleic acids (miRNAs) are regarded as potential biomarkers for a variety of diseases. Hence, the prediction of miRNA-disease associations (MDAs) is of great significance for an in-depth understanding of disease pathogenesis and progression. Existing prediction models are mainly concentrated on incorporating different sources of biological information to perform the MDA prediction task while failing to consider the fully potential utility of MDA network information at the motif-level. To overcome this problem, we propose a novel motif-aware MDA prediction model, namely MotifMDA, by fusing a variety of high- and low-order structural information. In particular, we first design several motifs of interest considering their ability to characterize how miRNAs are associated with diseases through different network structural patterns. Then, MotifMDA adopts a two-layer hierarchical attention to identify novel MDAs. Specifically, the first attention layer learns high-order motif preferences based on their occurrences in the given MDA network, while the second one learns the final embeddings of miRNAs and diseases through coupling high- and low-order preferences. Experimental results on two benchmark datasets have demonstrated the superior performance of MotifMDA over several state-of-the-art prediction models. This strongly indicates that accurate MDA prediction can be achieved by relying solely on MDA network information. Furthermore, our case studies indicate that the incorporation of motif-level structure information allows MotifMDA to discover novel MDAs from different perspectives. Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Pengwei Hu 0001, Zhu-Hong You, Lun Hu |
IEEE J. Biomed. Health Informatics | 8 |
| 2023 | CAMPEOD: A Cross Attention-Based Multi-Scale Patch Embedding Organoid Detection ModelabstractIn medical research, organoids, which exhibit structural and functional resemblances to authentic organs, offer a substantial avenue for delving into aspects encompassing physiology, pathophysiology, diseases, and pharmaceutical screening. Critical insights into drug responsiveness are often gleaned from these organoids' dimensional and configurational disparities. However, conventional detection methodologies reliant upon fluorescent labeling engender potential hazards to organoid integrity, thereby impinging upon their intrinsic dynamic attributes. Traditional bounding-box detection methodologies fall short in encapsulating intricate morphological particulars, and certain deep-learning approaches grapple with the intricate task of capturing multi-scale data, particularly when tasked with discerning organoid structures characterized by marked shape and size heterogeneities. In a bid to surmount these constraints, our study introduces CAMPEOD, an innovative framework that synergistically amalgamates multi-scale attributes derived from organoid specimens, employing cross-attention mechanisms. This novel approach effectively obviates superfluous background interference and image noise, thereby endowing an automated, finely-tuned dissection of organoid samples. Significantly, this segmentation process ensures congruence with authentic organoid quantities and morphological characteristics. By facilitating comprehensive scrutiny of microscopy images of organoid samples on a large scale, CAMPEOD assumes considerable implications for the realm of pharmaceutical screening and ailment emulation. Lun Hu, Zhu-Hong You, Pengwei Hu 0001 |
BIBM | 3 |
| 2023 | Deep-USIpred: identifying substrates of ubiquitin protein ligases E3 and deubiquitinases with pretrained protein embedding and bayesian neural networkabstractIdentifying the substrates of ubiquitin protein ligase (E3) and deubiquitinases (DUB) contributes to the discovery of potential therapeutic targets for diseases. However, experimental identification of E3/DUB-substrate interactions is costly and time-consuming. Current computational methods for predicting E3/DUB-substrate interactions rely heavily on specific domain knowledge and involve complex and diverse biological data processing. To address this challenge, we proposed a deep learning prediction model, named Deep-USIpred, which predicts E3/DUB-substrate interactions using protein sequences. The proposed Deep-USIpred model encodes protein sequences with a pretrained model and utilizes 1DCNN-BNN deep learning algorithm to make a robust prediction model. We evaluated the performance of the proposed model on real datasets, and our experimental results show that it can achieve excellent prediction performance on the tasks of ESI and DSI. Our proposed method provides a promising alternative for the prediction of E3/DUB-substrate interactions, which has the potential to accelerate drug discovery for various diseases. The source code and dataset are available at https://github.com/PGTSING/Deep-USIpred. Jia Wang 0008, Gui-Qing Pan, Jianqiang Li 0001, Xuequn Shang 0001, Zhu-Hong You |
BIBM | 5 |
| 2023 | A Novel Graph Representation Learning Model for Drug Repositioning Using Graph Transition Probability Matrix Over Heterogenous Information Networks
Dongxu Li 0002, Bo-Wei Zhao, Xiao-Rui Su 0001, Zhu-Hong You, Pengwei Hu 0001, Lun Hu |
ICIC (3) | 6 |
| 2023 | TransOrga: End-To-End Multi-modal Transformer-Based Organoid Segmentation
Jiajia Li 0004, Zhu-Hong You, Lun Hu, Pengwei Hu 0001, Feng Tan 0002 |
ICIC (3) | 6 |
| 2023 | Multi-level Subgraph Representation Learning for Drug-Disease Association Prediction Over Heterogeneous Biological Information Network
Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Dongxu Li 0002, Pengwei Hu 0001, Zhu-Hong You, Lun Hu |
ICIC (3) | 6 |
| 2023 | PTBGRP: predicting phage-bacteria interactions with graph representation learning on microbial heterogeneous information networkabstractIdentifying the potential bacteriophages (phage) candidate to treat bacterial infections plays an essential role in the research of human pathogens. Computational approaches are recognized as a valid way to predict bacteria and target phages. However, most of the current methods only utilize lower-order biological information without considering the higher-order connectivity patterns, which helps to improve the predictive accuracy. Therefore, we developed a novel microbial heterogeneous interaction network (MHIN)-based model called PTBGRP to predict new phages for bacterial hosts. Specifically, PTBGRP first constructs an MHIN by integrating phage-bacteria interaction (PBI) and six bacteria-bacteria interaction networks with their biological attributes. Then, different representation learning methods are deployed to extract higher-level biological features and lower-level topological features from MHIN. Finally, PTBGRP employs a deep neural network as the classifier to predict unknown PBI pairs based on the fused biological information. Experiment results demonstrated that PTBGRP achieves the best performance on the corresponding ESKAPE pathogens and PBI dataset when compared with state-of-art methods. In addition, case studies of Klebsiella pneumoniae and Staphylococcus aureus further indicate that the consideration of rich heterogeneous information enables PTBGRP to accurately predict PBI from a more comprehensive perspective. The webserver of the PTBGRP predictor is freely available at http://120.77.11.78/PTBGRP/. Jie Pan 0007, Zhu-Hong You, Wencai You, Chenlu Feng, Xuexia Zhang, Fengzhi Ren, Sanxing Ma, Yanmei Sun |
Briefings Bioinform. | 2 |
| 2023 | A feature extraction method based on noise reduction for circRNA-miRNA interaction prediction combining multi-structure features in the association networksabstractMOTIVATION: A large number of studies have shown that circular RNA (circRNA) affects biological processes by competitively binding miRNA, providing a new perspective for the diagnosis, and treatment of human diseases. Therefore, exploring the potential circRNA-miRNA interactions (CMIs) is an important and urgent task at present. Although some computational methods have been tried, their performance is limited by the incompleteness of feature extraction in sparse networks and the low computational efficiency of lengthy data. RESULTS: In this paper, we proposed JSNDCMI, which combines the multi-structure feature extraction framework and Denoising Autoencoder (DAE) to meet the challenge of CMI prediction in sparse networks. In detail, JSNDCMI integrates functional similarity and local topological structure similarity in the CMI network through the multi-structure feature extraction framework, then forces the neural network to learn the robust representation of features through DAE and finally uses the Gradient Boosting Decision Tree classifier to predict the potential CMIs. JSNDCMI produces the best performance in the 5-fold cross-validation of all data sets. In the case study, seven of the top 10 CMIs with the highest score were verified in PubMed. AVAILABILITY: The data and source code can be found at https://github.com/1axin/JSNDCMI. Xinfei Wang 0001, Zhu-Hong You, Liping Li 0003, Wenzhun Huang, Zhong-Hao Ren, Yue-Chao Li, Weixiao Meng 0001 |
Briefings Bioinform. | 3 |
| 2023 | SPRDA: a link prediction approach based on the structural perturbation to infer disease-associated Piwi-interacting RNAsabstractpiRNA and PIWI proteins have been confirmed for disease diagnosis and treatment as novel biomarkers due to its abnormal expression in various cancers. However, the current research is not strong enough to further clarify the functions of piRNA in cancer and its underlying mechanism. Therefore, how to provide large-scale and serious piRNA candidates for biological research has grown up to be a pressing issue. In this study, a novel computational model based on the structural perturbation method is proposed to predict potential disease-associated piRNAs, called SPRDA. Notably, SPRDA belongs to positive-unlabeled learning, which is unaffected by negative examples in contrast to previous approaches. In the 5-fold cross-validation, SPRDA shows high performance on the benchmark dataset piRDisease, with an AUC of 0.9529. Furthermore, the predictive performance of SPRDA for 10 diseases shows the robustness of the proposed method. Overall, the proposed approach can provide unique insights into the pathogenesis of the disease and will advance the field of oncology diagnosis and treatment. Kai Zheng 0020, Xin-Lu Zhang, Lei Wang 0121, Zhu-Hong You, Zhengwei Li 0001 |
Briefings Bioinform. | 4 |
| 2023 | Adversarial dense graph convolutional networks for single-cell classificationabstractMOTIVATION: In single-cell transcriptomics applications, effective identification of cell types in multicellular organisms and in-depth study of the relationships between genes has become one of the main goals of bioinformatics research. However, data heterogeneity and random noise pose significant difficulties for scRNA-seq data analysis. RESULTS: We have proposed an adversarial dense graph convolutional network architecture for single-cell classification. Specifically, to enhance the representation of higher-order features and the organic combination between features, dense connectivity mechanism and attention-based feature aggregation are introduced for feature learning in convolutional neural networks. To preserve the features of the original data, we use a feature reconstruction module to assist the goal of single-cell classification. In addition, HNNVAT uses virtual adversarial training to improve the generalization and robustness. Experimental results show that our model outperforms the existing classical methods in terms of classification accuracy on benchmark datasets. AVAILABILITY AND IMPLEMENTATION: The source code of HNNVAT is available at https://github.com/DisscLab/HNNVAT. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kangwei Wang, Zhengwei Li 0001, Zhu-Hong You, Pengyong Han, Ru Nie |
Bioinform. | 3 |
| 2023 | iGRLDTI: an improved graph representation learning method for predicting drug-target interactions over heterogeneous biological information networkabstractMOTIVATION: The task of predicting drug-target interactions (DTIs) plays a significant role in facilitating the development of novel drug discovery. Compared with laboratory-based approaches, computational methods proposed for DTI prediction are preferred due to their high-efficiency and low-cost advantages. Recently, much attention has been attracted to apply different graph neural network (GNN) models to discover underlying DTIs from heterogeneous biological information network (HBIN). Although GNN-based prediction methods achieve better performance, they are prone to encounter the over-smoothing simulation when learning the latent representations of drugs and targets with their rich neighborhood information in HBIN, and thereby reduce the discriminative ability in DTI prediction. RESULTS: In this work, an improved graph representation learning method, namely iGRLDTI, is proposed to address the above issue by better capturing more discriminative representations of drugs and targets in a latent feature space. Specifically, iGRLDTI first constructs an HBIN by integrating the biological knowledge of drugs and targets with their interactions. After that, it adopts a node-dependent local smoothing strategy to adaptively decide the propagation depth of each biomolecule in HBIN, thus significantly alleviating over-smoothing by enhancing the discriminative ability of feature representations of drugs and targets. Finally, a Gradient Boosting Decision Tree classifier is used by iGRLDTI to predict novel DTIs. Experimental results demonstrate that iGRLDTI yields better performance that several state-of-the-art computational methods on the benchmark dataset. Besides, our case study indicates that iGRLDTI can successfully identify novel DTIs with more distinguishable features of drugs and targets. AVAILABILITY AND IMPLEMENTATION: Python codes and dataset are available at https://github.com/stevejobws/iGRLDTI/. Bo-Wei Zhao, Xiao-Rui Su 0001, Pengwei Hu 0001, Zhu-Hong You, Lun Hu |
Bioinform. | 5 |
| 2023 | GKLOMLI: a link prediction model for inferring miRNA-lncRNA interactions by using Gaussian kernel-based method on network profile and linear optimization algorithmabstractBACKGROUND: The limited knowledge of miRNA-lncRNA interactions is considered as an obstruction of revealing the regulatory mechanism. Accumulating evidence on Human diseases indicates that the modulation of gene expression has a great relationship with the interactions between miRNAs and lncRNAs. However, such interaction validation via crosslinking-immunoprecipitation and high-throughput sequencing (CLIP-seq) experiments that inevitably costs too much money and time but with unsatisfactory results. Therefore, more and more computational prediction tools have been developed to offer many reliable candidates for a better design of further bio-experiments. METHODS: In this work, we proposed a novel link prediction model based on Gaussian kernel-based method and linear optimization algorithm for inferring miRNA-lncRNA interactions (GKLOMLI). Given an observed miRNA-lncRNA interaction network, the Gaussian kernel-based method was employed to output two similarity matrixes of miRNAs and lncRNAs. Based on the integrated matrix combined with similarity matrixes and the observed interaction network, a linear optimization-based link prediction model was trained for inferring miRNA-lncRNA interactions. RESULTS: To evaluate the performance of our proposed method, k-fold cross-validation (CV) and leave-one-out CV were implemented, in which each CV experiment was carried out 100 times on a training set generated randomly. The high area under the curves (AUCs) at 0.8623 ± 0.0027 (2-fold CV), 0.9053 ± 0.0017 (5-fold CV), 0.9151 ± 0.0013 (10-fold CV), and 0.9236 (LOO-CV), illustrated the precision and reliability of our proposed method. CONCLUSION: GKLOMLI with high performance is anticipated to be used to reveal underlying interactions between miRNA and their target lncRNAs, and deciphers the potential mechanisms of the complex diseases. Leon Wong, Lei Wang 0121, Zhu-Hong You, Chang-an Yuan 0001, Mei-Yuan Cao |
BMC Bioinform. | 3 |
| 2023 | PDA-PRGCN: identification of Piwi-interacting RNA-disease associations through subgraph projection and residual scaling-based feature augmentationabstractBACKGROUND: Emerging evidences show that Piwi-interacting RNAs (piRNAs) play a pivotal role in numerous complex human diseases. Identifying potential piRNA-disease associations (PDAs) is crucial for understanding disease pathogenesis at molecular level. Compared to the biological wet experiments, the computational methods provide a cost-effective strategy. However, few computational methods have been developed so far. RESULTS: Here, we proposed an end-to-end model, referred to as PDA-PRGCN (PDA prediction using subgraph Projection and Residual scaling-based feature augmentation through Graph Convolutional Network). Specifically, starting with the known piRNA-disease associations represented as a graph, we applied subgraph projection to construct piRNA-piRNA and disease-disease subgraphs for the first time, followed by a residual scaling-based feature augmentation algorithm for node initial representation. Then, we adopted graph convolutional network (GCN) to learn and identify potential PDAs as a link prediction task on the constructed heterogeneous graph. Comprehensive experiments, including the performance comparison of individual components in PDA-PRGCN, indicated the significant improvement of integrating subgraph projection, node feature augmentation and dual-loss mechanism into GCN for PDA prediction. Compared with state-of-the-art approaches, PDA-PRGCN gave more accurate and robust predictions. Finally, the case studies further corroborated that PDA-PRGCN can reliably detect PDAs. CONCLUSION: PDA-PRGCN provides a powerful method for PDA prediction, which can also serve as a screening tool for studies of complex diseases. Ping Zhang 0027, Weicheng Sun, Dengguo Wei, Jinsheng Xu, Zhu-Hong You, Bo-Wei Zhao, Li Li 0057 |
BMC Bioinform. | 6 |
| 2023 | In silico prediction methods of self-interacting proteins: an empirical and academic survey
Zhu-Hong You, Qinhu Zhang, Zhen-Hao Guo, Siguo Wang |
Frontiers Comput. Sci. | 2 |
| 2023 | Knowledge graph embedding for profiling the interaction between transcription factors and their target genesabstractInteractions between transcription factor and target gene form the main part of gene regulation network in human, which are still complicating factors in biological research. Specifically, for nearly half of those interactions recorded in established database, their interaction types are yet to be confirmed. Although several computational methods exist to predict gene interactions and their type, there is still no method available to predict them solely based on topology information. To this end, we proposed here a graph-based prediction model called KGE-TGI and trained in a multi-task learning manner on a knowledge graph that we specially constructed for this problem. The KGE-TGI model relies on topology information rather than being driven by gene expression data. In this paper, we formulate the task of predicting interaction types of transcript factor and target genes as a multi-label classification problem for link types on a heterogeneous graph, coupled with solving another link prediction problem that is inherently related. We constructed a ground truth dataset as benchmark and evaluated the proposed method on it. As a result of the 5-fold cross experiments, the proposed method achieved average AUC values of 0.9654 and 0.9339 in the tasks of link prediction and link type classification, respectively. In addition, the results of a series of comparison experiments also prove that the introduction of knowledge information significantly benefits to the prediction and that our methodology achieve state-of-the-art performance in this problem. Yang-Han Wu, Jianqiang Li 0001, Zhu-Hong You, Pengwei Hu 0001, Lun Hu, Victor C. M. Leung, Zhihua Du |
PLoS Comput. Biol. | 4 |
| 2023 | Predicting MiRNA-Disease Associations by Graph Representation Learning Based on Jumping Knowledge NetworksabstractGrowing studies have shown that miRNAs are inextricably linked with many human diseases, and a great deal of effort has been spent on identifying their potential associations. Compared with traditional experimental methods, computational approaches have achieved promising results. In this article, we propose a graph representation learning method to predict miRNA-disease associations. Specifically, we first integrate the verified miRNA-disease associations with the similarity information of miRNA and disease to construct a miRNA-disease heterogeneous graph. Then, we apply a graph attention network to aggregate the neighbor information of nodes in each layer, and then feed the representation of the hidden layer into the structure-aware jumping knowledge network to obtain the global features of nodes. The output features of miRNAs and diseases are then concatenated and fed into a fully connected layer to score the potential associations. Through five-fold cross-validation, the average AUC, accuracy and precision values of our model are 93.30%, 85.18% and 88.90%, respectively. In addition, for three case studies of the esophageal tumor, lymphoma and prostate tumor, 46, 45 and 45 of the top 50 miRNAs predicted by our model were confirmed by relevant databases. Overall, our method could provide a reliable alternative for miRNA-disease association prediction. Zhengwei Li 0001, Chang-an Yuan 0001, Pengyong Han, Zhu-Hong You, Lei Wang 0121 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2023 | Combining K Nearest Neighbor With Nonnegative Matrix Factorization for Predicting Circrna-Disease AssociationsabstractAccumulating evidences show that circular RNAs (circRNAs) play an important role in regulating gene expression, and involve in many complex human diseases. Identifying associations of circRNA with disease helps to understand the pathogenesis, treatment and diagnosis of complex diseases. Since inferring circRNA-disease associations by biological experiments is costly and time-consuming, there is an urgently need to develop a computational model to identify the association between them. In this paper, we proposed a novel method named KNN-NMF, which combines K nearest neighbors with nonnegative matrix factorization to infer associations between circRNA and disease (KNN-NMF). Frist, we compute the Gaussian Interaction Profile (GIP) kernel similarity of circRNA and disease, the semantic similarity of disease, respectively. Then, the circRNA-disease new interaction profiles are established using weight K nearest neighbors to reduce the false negative association impact on prediction performance. Finally, Nonnegative Matrix Factorization is implemented to predict associations of circRNA with disease. The experiment results indicate that the prediction performance of KNN-NMF outperforms the competing methods under five-fold cross-validation. Moreover, case studies of two common diseases further show that KNN-NMF can identify potential circRNA-disease associations effectively. Meineng Wang, Xue-Jun Xie, Zhu-Hong You, Leon Wong, Liping Li 0003 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | Predicting Mirna-Disease Associations Based on Neighbor Selection Graph Attention NetworksabstractNumerous experiments have shown that the occurrence of complex human diseases is often accompanied by abnormal expression of microRNA (miRNA). Identifying the associations between miRNAs and diseases is of great significance in the development of clinical medicine. However, traditional experimental methods are often time-consuming and inefficient. To this end, we proposed a deep learning method based on neighbor selection graph attention networks for predicting miRNA-disease associations (NSAMDA). Specifically, we firstly fused miRNA sequence similarity information and miRNA integrated similarity information to enrich miRNA feature information. Secondly, we used the fused miRNA feature information and disease integrated similarity information to construct a miRNA-disease heterogeneous graph. Thirdly, we introduced a neighbor selection method based on graph attention networks to select k-most important neighbors for aggregation. Finally, we used the inner product decoder to score miRNA-disease pairs. The results of five-fold cross-validation show that the mean AUC of NSAMDA is 93.69% on HMDD v2.0 dataset. In addition, case studies on the esophageal neoplasm, lung neoplasm and lymphoma were carried out to further confirm the effectiveness of the NSAMDA model. The results showed that the NSAMDA method achieves satisfactory performance on predicting miRNA-disease associations and is superior to the most advanced model. Zhengwei Li 0001, Zhu-Hong You, Ru Nie, Tangbo Zhong |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | MGRCDA: Metagraph Recommendation Method for Predicting CircRNA-Disease AssociationabstractClinical evidence began to accumulate, suggesting that circRNAs can be novel therapeutic targets for various diseases and play a critical role in human health. However, limited by the complex mechanism of circRNA, it is difficult to quickly and large-scale explore the relationship between disease and circRNA in the wet-lab experiment. In this work, we design a new computational model MGRCDA on account of the metagraph recommendation theory to predict the potential circRNA-disease associations. Specifically, we first regard the circRNA-disease association prediction problem as the system recommendation problem, and design a series of metagraphs according to the heterogeneous biological networks; then extract the semantic information of the disease and the Gaussian interaction profile kernel (GIPK) similarity of circRNA and disease as network attributes; finally, the iterative search of the metagraph recommendation algorithm is used to calculate the scores of the circRNA-disease pair. On the gold standard dataset circR2Disease, MGRCDA achieved a prediction accuracy of 92.49% with an area under the ROC curve of 0.9298, which is significantly higher than other state-of-the-art models. Furthermore, among the top 30 disease-related circRNAs recommended by the model, 25 have been verified by the latest published literature. The experimental results prove that MGRCDA is feasible and efficient, and it can recommend reliable candidates to further wet-lab experiment and reduce the scope of the experiment. Lei Wang 0121, Zhu-Hong You, De-Shuang Huang, Jianqiang Li 0001 |
IEEE Trans. Cybern. | 2 |
| 2023 | PPAEDTI: Personalized Propagation Auto-Encoder Model for Predicting Drug-Target InteractionsabstractIdentifying protein targets for drugs establishes an indispensable knowledge foundation for drug repurposing and drug development. Though expensive and time-consuming, vitro trials are widely employed to discover drug targets, and the existing relevant computational algorithms still cannot satisfy the demand for real application in drug R&D with regards to the prediction accuracy and performance efficiency, which are urgently needed to be improved. To this end, we propose here the PPAEDTI model, which uses the graph personalized propagation technique to predict drug-target interactions from the known interaction network. To evaluate the prediction performance, six benchmark datasets were used for testing with some state-of-the-art methods compared. As a result, using the 5-fold cross-validation, the proposed PPAEDTI model achieves average AUCs>90% on 5 collected datasets. We also manually checked the top-20 prediction list for 2 proteins (hsa:775 and hsa:779) and a kind of drug (D00618), and successfully confirmed 18, 17, and 20 items from the public datasets, respectively. The experimental results indicate that, given known drug-target interactions, the PPAEDTI model can provide accurate predictions for the new ones, which is anticipated to serve as a useful tool for pharmacology research. Using the proposed model that was trained with the collected datasets, we have built a computational platform that is accessible at http://120.77.11.78/PPAEDTI/ and corresponding codes and datasets are also released. Yue-Chao Li, Zhu-Hong You, Lei Wang 0121, Leon Wong, Lun Hu, Pengwei Hu 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Predicting Drug-Target Interactions Over Heterogeneous Information NetworkabstractIdentifying Drug-Target Interactions (DTIs) is a critical step in studying pathogenesis and drug development. Due to the fact that conventional experimental methods usually suffer from high costs and low efficiency, various computational methods have been proposed to detect potential DTIs by extracting features from the biological information of drugs and their target proteins. Though effective, most of them fall short of considering the topological structure of the DTI network, which provides a global view to discover novel DTIs. In this paper, a network-based computational method, namely LG-DTI, is proposed to accurately predict DTIs over a heterogeneous information network. For drugs and target proteins, LG-DTI first learns not only their local representations from drug molecular structures and protein sequences, but also their global representations by using a semi-supervised heterogeneous network embedding method. These two kinds of representations consist of the final representations of drugs and target proteins, which are then incorporated into a Random Forest classifier to complete the task of DTI prediction. The performance of LG-DTI has been evaluated on two independent datasets and also compared with several state-of-the-art methods. Experimental results show the superior performance of LG-DTI. Moreover, our case study indicates that LG-DTI can be a valuable tool for identifying novel DTIs. Xiao-Rui Su 0001, Pengwei Hu 0001, Zhu-Hong You, Lun Hu |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Biomedical Knowledge Graph Embedding With Capsule Network for Multi-Label Drug-Drug Interaction PredictionabstractDrug-drug interaction (DDI) plays an important role in drug development and administration. Most of existing network-based computation models regard the DDI prediction as a binary classification problem and generate negative DDI samples randomly, but the binary classification is not in line with the real problem since there are dozens of types of DDI and randomly generating negative samples may introduce false-negative samples since the non-observed facts can be either false or just missing. To address the above limitations, we propose a new framework called KG2ECapsule that explicitly models the multi-relational DDI data based on biomedical knowledge graphs in an end-to-end fashion. It first generates high-quality negative samples based on the average number of tail entities and head entities for each relation to reduce false-negative samples to some extent. KG2ECapsule then refines the representations of entities by recursively propagating the embeddings from the attention-based receptive fields of entities. Empirical results on three biomedical knowledge graphs of different scales show that KG2ECapsule outperforms the state-of-the-art methods consistently in multi-label DDI prediction task and further studies verify the efficacy of both probability-based sampling strategy and non-linear transformation for modeling multi-relational data. Xiao-Rui Su 0001, Zhu-Hong You, De-Shuang Huang, Lei Wang 0121, Leon Wong, Bo-Wei Zhao |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | Predicting circRNA-disease associations using similarity assessing graph convolution from multi-source information networksabstractCircular RNA (circRNA), a novel endogenous noncoding RNA molecule with a closed-loop structure, can be used as a biomarker for many complex human diseases. Determining the relationship between circRNAs and diseases helps us to understand the diagnosis, treatment, and pathogenesis of complex diseases, which plays a critical role in clinical research. Nevertheless, the discovery of new circRNA-disease associations by wet-lab methods is not only time-consuming and costly but also randomized and blinded, which is also limited to small-scale studies. Thus, there is an urgent need to establish efficient and reliable computational methods to infer potential circRNA-disease associations on a large scale to effectively reduce costs and save time, and avoid high false-positive rates. In this paper, we propose a novel computational method for predicting circRNA-disease association based on the Similarity Assessing Graph Convolution Network (SAGCN) algorithm, which combines the multi-source similarity network constructed by circRNA and disease. Firstly, we fuse the multi-source similarity information of circRNAs and diseases and construct the multi-source similarity network respectively. Then we use the SAGCN algorithm to extract the hidden feature representations of circRNAs and diseases efficiently and objectively in the way of measuring the similarity between different nodes in the network. Finally, the obtained high-level features of circRNAs and diseases are fed to the multilayer perceptron (MLP) classifier for accurate prediction. Using the 5-fold cross-validation method, the AUC scores of the four SAGCN algorithms, on the benchmark circR2Disease dataset are 93.30%, 92.98%, 92.22% and 91.94%, respectively. Furthermore, case studies further validated that the proposed model was supported by biological experiments, and 25 of the top 30 circRNA-disease associations with the highest scores were confirmed by recent literature. Based on these reliable results, it can be anticipated that the proposed model can be used as an effective computational tool to predict circRNA-disease associations and can provide the most promising candidates for biological experiments. Yang Li 0111, Xue-Gang Hu, Pei-Pei Li 0001, Lei Wang 0121, Zhu-Hong You |
BIBM | 5 |
| 2022 | Predicting Drug-Disease Associations via Meta-path Representation Learning based on Heterogeneous Information Net works
Menglong Zhang, Bo-Wei Zhao, Lun Hu, Zhu-Hong You |
ICIC (2) | 4 |
| 2022 | MRLDTI: A Meta-path-Based Representation Learning Model for Drug-Target Interaction Prediction
Bo-Wei Zhao, Lun Hu, Pengwei Hu 0001, Zhu-Hong You, Xiao-Rui Su 0001, Dongxu Li 0002, Ping Zhang 0027 |
ICIC (2) | 4 |
| 2022 | Identification of miRNA-lncRNA Underlying Interactions Through Representation for Multiplex Heterogeneous Network
Ji-Ren Zhou, Zhu-Hong You, Xuequn Shang 0001, Rui Niu, Yue Yun |
ICIC (2) | 2 |
| 2022 | GraphTGI: an attention-based graph embedding model for predicting TF-target gene interactionsabstractMOTIVATION: Interaction between transcription factor (TF) and its target genes establishes the knowledge foundation for biological researches in transcriptional regulation, the number of which is, however, still limited by biological techniques. Existing computational methods relevant to the prediction of TF-target interactions are mostly proposed for predicting binding sites, rather than directly predicting the interactions. To this end, we propose here a graph attention-based autoencoder model to predict TF-target gene interactions using the information of the known TF-target gene interaction network combined with two sequential and chemical gene characters, considering that the unobserved interactions between transcription factors and target genes can be predicted by learning the pattern of the known ones. To the best of our knowledge, the proposed model is the first attempt to solve this problem by learning patterns from the known TF-target gene interaction network. RESULTS: In this paper, we formulate the prediction task of TF-target gene interactions as a link prediction problem on a complex knowledge graph and propose a deep learning model called GraphTGI, which is composed of a graph attention-based encoder and a bilinear decoder. We evaluated the prediction performance of the proposed method on a real dataset, and the experimental results show that the proposed model yields outstanding performance with an average AUC value of 0.8864 +/- 0.0057 in the 5-fold cross-validation. It is anticipated that the GraphTGI model can effectively and efficiently predict TF-target gene interactions on a large scale. AVAILABILITY: Python code and the datasets used in our studies are made available at https://github.com/YanghanWu/GraphTGI. Zhihua Du, Yang-Han Wu, Jie Chen 0027, Gui-Qing Pan, Lun Hu, Zhu-Hong You, Jianqiang Li 0001 |
Briefings Bioinform. | 7 |
| 2022 | A novel circRNA-miRNA association prediction model based on structural deep neural network embeddingabstractA large amount of clinical evidence began to mount, showing that circular ribonucleic acids (RNAs; circRNAs) perform a very important function in complex diseases by participating in transcription and translation regulation of microRNA (miRNA) target genes. However, with strict high-throughput techniques based on traditional biological experiments and the conditions and environment, the association between circRNA and miRNA can be discovered to be labor-intensive, expensive, time-consuming, and inefficient. In this paper, we proposed a novel computational model based on Word2vec, Structural Deep Network Embedding (SDNE), Convolutional Neural Network and Deep Neural Network, which predicts the potential circRNA-miRNA associations, called Word2vec, SDNE, Convolutional Neural Network and Deep Neural Network (WSCD). Specifically, the WSCD model extracts attribute feature and behaviour feature by word embedding and graph embedding algorithm, respectively, and ultimately feed them into a feature fusion model constructed by combining Convolutional Neural Network and Deep Neural Network to deduce potential circRNA-miRNA interactions. The proposed method is proved on dataset and obtained a prediction accuracy and an area under the receiver operating characteristic curve of 81.61% and 0.8898, respectively, which is shown to have much higher accuracy than the state-of-the-art models and classifier models in prediction. In addition, 23 miRNA-related circular RNAs (circRNAs) from the top 30 were confirmed in relevant experiences. In these works, all results represent that WSCD would be a helpful supplementary reliable method for predicting potential miRNA-circRNA associations compared to wet laboratory experiments. Lu-Xiang Guo, Zhu-Hong You, Lei Wang 0121, Bo-Wei Zhao, Zhong-Hao Ren, Jie Pan 0007 |
Briefings Bioinform. | 2 |
| 2022 | MNMDCDA: prediction of circRNA-disease associations by learning mixed neighborhood information from multiple distancesabstractEmerging evidence suggests that circular RNA (circRNA) is an important regulator of a variety of pathological processes and serves as a promising biomarker for many complex human diseases. Nevertheless, there are relatively few known circRNA-disease associations, and uncovering new circRNA-disease associations by wet-lab methods is time consuming and costly. Considering the limitations of existing computational methods, we propose a novel approach named MNMDCDA, which combines high-order graph convolutional networks (high-order GCNs) and deep neural networks to infer associations between circRNAs and diseases. Firstly, we computed different biological attribute information of circRNA and disease separately and used them to construct multiple multi-source similarity networks. Then, we used the high-order GCN algorithm to learn feature embedding representations with high-order mixed neighborhood information of circRNA and disease from the constructed multi-source similarity networks, respectively. Finally, the deep neural network classifier was implemented to predict associations of circRNAs with diseases. The MNMDCDA model obtained AUC scores of 95.16%, 94.53%, 89.80% and 91.83% on four benchmark datasets, i.e., CircR2Disease, CircAtlas v2.0, Circ2Disease and CircRNADisease, respectively, using the 5-fold cross-validation approach. Furthermore, 25 of the top 30 circRNA-disease pairs with the best scores of MNMDCDA in the case study were validated by recent literature. Numerous experimental results indicate that MNMDCDA can be used as an effective computational tool to predict circRNA-disease associations and can provide the most promising candidates for biological experiments. Yang Li 0111, Xue-Gang Hu, Lei Wang 0121, Pei-Pei Li 0001, Zhu-Hong You |
Briefings Bioinform. | 5 |
| 2022 | A biomedical knowledge graph-based method for drug-drug interactions prediction through combining local and global features with deep neural networksabstractDrug-drug interactions (DDIs) prediction is a challenging task in drug development and clinical application. Due to the extremely large complete set of all possible DDIs, computer-aided DDIs prediction methods are getting lots of attention in the pharmaceutical industry and academia. However, most existing computational methods only use single perspective information and few of them conduct the task based on the biomedical knowledge graph (BKG), which can provide more detailed and comprehensive drug lateral side information flow. To this end, a deep learning framework, namely DeepLGF, is proposed to fully exploit BKG fusing local-global information to improve the performance of DDIs prediction. More specifically, DeepLGF first obtains chemical local information on drug sequence semantics through a natural language processing algorithm. Then a model of BFGNN based on graph neural network is proposed to extract biological local information on drug through learning embedding vector from different biological functional spaces. The global feature information is extracted from the BKG by our knowledge graph embedding method. In DeepLGF, for fusing local-global features well, we designed four aggregating methods to explore the most suitable ones. Finally, the advanced fusing feature vectors are fed into deep neural network to train and predict. To evaluate the prediction performance of DeepLGF, we tested our method in three prediction tasks and compared it with state-of-the-art models. In addition, case studies of three cancer-related and COVID-19-related drugs further demonstrated DeepLGF's superior ability for potential DDIs prediction. The webserver of the DeepLGF predictor is freely available at http://120.77.11.78/DeepLGF/. Zhong-Hao Ren, Zhu-Hong You, Liping Li 0003, Yongjian Guan, Lu-Xiang Guo, Jie Pan 0007 |
Briefings Bioinform. | 2 |
| 2022 | A deep learning method for repurposing antiviral drugs against new viruses via multi-view nonnegative matrix factorization and its application to SARS-CoV-2abstractThe outbreak of COVID-19 caused by SARS-coronavirus (CoV)-2 has made millions of deaths since 2019. Although a variety of computational methods have been proposed to repurpose drugs for treating SARS-CoV-2 infections, it is still a challenging task for new viruses, as there are no verified virus-drug associations (VDAs) between them and existing drugs. To efficiently solve the cold-start problem posed by new viruses, a novel constrained multi-view nonnegative matrix factorization (CMNMF) model is designed by jointly utilizing multiple sources of biological information. With the CMNMF model, the similarities of drugs and viruses can be preserved from their own perspectives when they are projected onto a unified latent feature space. Based on the CMNMF model, we propose a deep learning method, namely VDA-DLCMNMF, for repurposing drugs against new viruses. VDA-DLCMNMF first initializes the node representations of drugs and viruses with their corresponding latent feature vectors to avoid a random initialization and then applies graph convolutional network to optimize their representations. Given an arbitrary drug, its probability of being associated with a new virus is computed according to their representations. To evaluate the performance of VDA-DLCMNMF, we have conducted a series of experiments on three VDA datasets created for SARS-CoV-2. Experimental results demonstrate that the promising prediction accuracy of VDA-DLCMNMF. Moreover, incorporating the CMNMF model into deep learning gains new insight into the drug repurposing for SARS-CoV-2, as the results of molecular docking experiments reveal that four antiviral drugs identified by VDA-DLCMNMF have the potential ability to treat SARS-CoV-2 infections. Xiao-Rui Su 0001, Lun Hu, Zhu-Hong You, Pengwei Hu 0001, Lei Wang 0121, Bo-Wei Zhao |
Briefings Bioinform. | 3 |
| 2022 | Attention-based Knowledge Graph Representation Learning for Predicting Drug-drug InteractionsabstractDrug-drug interactions (DDIs) are known as the main cause of life-threatening adverse events, and their identification is a key task in drug development. Existing computational algorithms mainly solve this problem by using advanced representation learning techniques. Though effective, few of them are capable of performing their tasks on biomedical knowledge graphs (KGs) that provide more detailed information about drug attributes and drug-related triple facts. In this work, an attention-based KG representation learning framework, namely DDKG, is proposed to fully utilize the information of KGs for improved performance of DDI prediction. In particular, DDKG first initializes the representations of drugs with their embeddings derived from drug attributes with an encoder-decoder layer, and then learns the representations of drugs by recursively propagating and aggregating first-order neighboring information along top-ranked network paths determined by neighboring node embeddings and triple facts. Last, DDKG estimates the probability of being interacting for pairwise drugs with their representations in an end-to-end manner. To evaluate the effectiveness of DDKG, extensive experiments have been conducted on two practical datasets with different sizes, and the results demonstrate that DDKG is superior to state-of-the-art algorithms on the DDI prediction task in terms of different evaluation metrics across all datasets. Xiao-Rui Su 0001, Lun Hu, Zhu-Hong You, Pengwei Hu 0001, Bo-Wei Zhao |
Briefings Bioinform. | 3 |
| 2022 | A machine learning framework based on multi-source feature fusion for circRNA-disease association predictionabstractCircular RNAs (circRNAs) are involved in the regulatory mechanisms of multiple complex diseases, and the identification of their associations is critical to the diagnosis and treatment of diseases. In recent years, many computational methods have been designed to predict circRNA-disease associations. However, most of the existing methods rely on single correlation data. Here, we propose a machine learning framework for circRNA-disease association prediction, called MLCDA, which effectively fuses multiple sources of heterogeneous information including circRNA sequences and disease ontology. Comprehensive evaluation in the gold standard dataset showed that MLCDA can successfully capture the complex relationships between circRNAs and diseases and accurately predict their potential associations. In addition, the results of case studies on real data show that MLCDA significantly outperforms other existing methods. MLCDA can serve as a useful tool for circRNA-disease association prediction, providing mechanistic insights for disease research and thus facilitating the progress of disease treatment. Lei Wang 0121, Leon Wong, Zhengwei Li 0001, Xiao-Rui Su 0001, Bo-Wei Zhao, Zhu-Hong You |
Briefings Bioinform. | 7 |
| 2022 | Graph representation learning in bioinformatics: trends, methods and applicationsabstractGraph is a natural data structure for describing complex systems, which contains a set of objects and relationships. Ubiquitous real-life biomedical problems can be modeled as graph analytics tasks. Machine learning, especially deep learning, succeeds in vast bioinformatics scenarios with data represented in Euclidean domain. However, rich relational information between biological elements is retained in the non-Euclidean biomedical graphs, which is not learning friendly to classic machine learning methods. Graph representation learning aims to embed graph into a low-dimensional space while preserving graph topology and node properties. It bridges biomedical graphs and modern machine learning methods and has recently raised widespread interest in both machine learning and bioinformatics communities. In this work, we summarize the advances of graph representation learning and its representative applications in bioinformatics. To provide a comprehensive and structured analysis and perspective, we first categorize and analyze both graph embedding methods (homogeneous graph embedding, heterogeneous graph embedding, attribute graph embedding) and graph neural networks. Furthermore, we summarize their representative applications from molecular level to genomics, pharmaceutical and healthcare systems level. Moreover, we provide open resource platforms and libraries for implementing these graph representation learning methods and discuss the challenges and opportunities of graph representation learning in bioinformatics. This work provides a comprehensive survey of emerging graph representation learning algorithms and their applications in bioinformatics. It is anticipated that it could bring valuable insights for researchers to contribute their knowledge to graph representation learning and future-oriented bioinformatics studies. Zhu-Hong You, De-Shuang Huang, Chee Keong Kwoh 0001 |
Briefings Bioinform. | 2 |
| 2022 | iGRLCDA: identifying circRNA-disease association based on graph representation learningabstractWhile the technologies of ribonucleic acid-sequence (RNA-seq) and transcript assembly analysis have continued to improve, a novel topology of RNA transcript was uncovered in the last decade and is called circular RNA (circRNA). Recently, researchers have revealed that they compete with messenger RNA (mRNA) and long noncoding for combining with microRNA in gene regulation. Therefore, circRNA was assumed to be associated with complex disease and discovering the relationship between them would contribute to medical research. However, the work of identifying the association between circRNA and disease in vitro takes a long time and usually without direction. During these years, more and more associations were verified by experiments. Hence, we proposed a computational method named identifying circRNA-disease association based on graph representation learning (iGRLCDA) for the prediction of the potential association of circRNA and disease, which utilized a deep learning model of graph convolution network (GCN) and graph factorization (GF). In detail, iGRLCDA first derived the hidden feature of known associations between circRNA and disease using the Gaussian interaction profile (GIP) kernel combined with disease semantic information to form a numeric descriptor. After that, it further used the deep learning model of GCN and GF to extract hidden features from the descriptor. Finally, the random forest classifier is introduced to identify the potential circRNA-disease association. The five-fold cross-validation of iGRLCDA shows strong competitiveness in comparison with other excellent prediction models at the gold standard data and achieved an average area under the receiver operating characteristic curve of 0.9289 and an area under the precision-recall curve of 0.9377. On reviewing the prediction results from the relevant literature, 22 of the top 30 predicted circRNA-disease associations were noted in recent published papers. These exceptional results make us believe that iGRLCDA can provide reliable circRNA-disease associations for medical research and reduce the blindness of wet-lab experiments. Lei Wang 0121, Zhu-Hong You, Lun Hu, Bo-Wei Zhao, Zhengwei Li 0001, Yang-Ming Li |
Briefings Bioinform. | 3 |
| 2022 | HINGRL: predicting drug-disease associations with graph representation learning on heterogeneous information networksabstractIdentifying new indications for drugs plays an essential role at many phases of drug research and development. Computational methods are regarded as an effective way to associate drugs with new indications. However, most of them complete their tasks by constructing a variety of heterogeneous networks without considering the biological knowledge of drugs and diseases, which are believed to be useful for improving the accuracy of drug repositioning. To this end, a novel heterogeneous information network (HIN) based model, namely HINGRL, is proposed to precisely identify new indications for drugs based on graph representation learning techniques. More specifically, HINGRL first constructs a HIN by integrating drug-disease, drug-protein and protein-disease biological networks with the biological knowledge of drugs and diseases. Then, different representation strategies are applied to learn the features of nodes in the HIN from the topological and biological perspectives. Finally, HINGRL adopts a Random Forest classifier to predict unknown drug-disease associations based on the integrated features of drugs and diseases obtained in the previous step. Experimental results demonstrate that HINGRL achieves the best performance on two real datasets when compared with state-of-the-art models. Besides, our case studies indicate that the simultaneous consideration of network topology and biological knowledge of drugs and diseases allows HINGRL to precisely predict drug-disease associations from a more comprehensive perspective. The promising performance of HINGRL also reveals that the utilization of rich heterogeneous information provides an alternative view for HINGRL to identify novel drug-disease associations especially for new diseases. Bo-Wei Zhao, Lun Hu, Zhu-Hong You, Lei Wang 0121, Xiao-Rui Su 0001 |
Briefings Bioinform. | 3 |
| 2022 | Line graph attention networks for predicting disease-associated Piwi-interacting RNAsabstractPIWI proteins and Piwi-Interacting RNAs (piRNAs) are commonly detected in human cancers, especially in germline and somatic tissues, and correlate with poorer clinical outcomes, suggesting that they play a functional role in cancer. As the problem of combinatorial explosions between ncRNA and disease exposes gradually, new bioinformatics methods for large-scale identification and prioritization of potential associations are therefore of interest. However, in the real world, the network of interactions between molecules is enormously intricate and noisy, which poses a problem for efficient graph mining. Line graphs can extend many heterogeneous networks to replace dichotomous networks. In this study, we present a new graph neural network framework, line graph attention networks (LGAT). And we apply it to predict PiRNA disease association (GAPDA). In the experiment, GAPDA performs excellently in 5-fold cross-validation with an AUC of 0.9038. Not only that, it still has superior performance compared with methods based on collaborative filtering and attribute features. The experimental results show that GAPDA ensures the prospect of the graph neural network on such problems and can be an excellent supplement for future biomedical research. Kai Zheng 0020, Xin-Lu Zhang, Lei Wang 0121, Zhu-Hong You, Zhaohui Zhan |
Briefings Bioinform. | 4 |
| 2022 | Predicting miRNA-disease associations based on graph random propagation network and attention networkabstractNumerous experiments have demonstrated that abnormal expression of microRNAs (miRNAs) in organisms is often accompanied by the emergence of specific diseases. The research of miRNAs can promote the prevention and drug research of specific diseases. However, there are still many undiscovered links between miRNAs and diseases, which greatly limits the research of miRNAs. Therefore, for exploring the unknown miRNA-disease associations, we combine the graph random propagation network based on DropFeature with attention network to propose a novel deep learning model to predict the miRNA-disease associations (GRPAMDA). Specifically, we firstly construct the miRNA-disease heterogeneous graph based on miRNA-disease association information. Secondly, we adopt DropFeature to randomly delete the features of nodes in the graph and then perform propagation operations to enhance the features of miRNA and disease nodes. Thirdly, we employ the attention mechanism to fuse the features of random propagation by aggregating the enhanced neighbor features of miRNA and disease nodes. Finally, miRNA-disease association scores are generated by a fully connected layer. The average area under the curve of GRPAMDA model based on 5-fold cross-validation is 93.46% on HMDD v2.0. Case studies of esophageal tumors, lymphomas and prostate tumors show that 48, 47 and 46 of the top 50 miRNAs associated with these diseases are confirmed by dbDEMC and miR2Disease database, respectively. In short, the GRPAMDA model can be used as a valuable method to study miRNA-disease associations. Tangbo Zhong, Zhengwei Li 0001, Zhu-Hong You, Ru Nie |
Briefings Bioinform. | 3 |
| 2022 | Robust and accurate prediction of self-interacting proteins from protein sequence information by exploiting weighted sparse representation based classifierabstractBACKGROUND: Self-interacting proteins (SIPs), two or more copies of the protein that can interact with each other expressed by one gene, play a central role in the regulation of most living cells and cellular functions. Although numerous SIPs data can be provided by using high-throughput experimental techniques, there are still several shortcomings such as in time-consuming, costly, inefficient, and inherently high in false-positive rates, for the experimental identification of SIPs even nowadays. Therefore, it is more and more significant how to develop efficient and accurate automatic approaches as a supplement of experimental methods for assisting and accelerating the study of predicting SIPs from protein sequence information. RESULTS: In this paper, we present a novel framework, termed GLCM-WSRC (gray level co-occurrence matrix-weighted sparse representation based classification), for predicting SIPs automatically based on protein evolutionary information from protein primary sequences. More specifically, we firstly convert the protein sequence into Position Specific Scoring Matrix (PSSM) containing protein sequence evolutionary information, exploiting the Position Specific Iterated BLAST (PSI-BLAST) tool. Secondly, using an efficient feature extraction approach, i.e., GLCM, we extract abstract salient and invariant feature vectors from the PSSM, and then perform a pre-processing operation, the adaptive synthetic (ADASYN) technique, to balance the SIPs dataset to generate new feature vectors for classification. Finally, we employ an efficient and reliable WSRC model to identify SIPs according to the known information of self-interacting and non-interacting proteins. CONCLUSIONS: Extensive experimental results show that the proposed approach exhibits high prediction performance with 98.10% accuracy on the yeast dataset, and 91.51% accuracy on the human dataset, which further reveals that the proposed model could be a useful tool for large-scale self-interacting protein prediction and other bioinformatics tasks detection in the future. Yang Li 0111, Xuegang Hu, Zhu-Hong You, Liping Li 0003, Pei-Pei Li 0001 |
BMC Bioinform. | 3 |
| 2022 | Multi-view heterogeneous molecular network representation learning for protein-protein interaction predictionabstractBACKGROUND: Protein-protein interaction (PPI) plays an important role in regulating cells and signals. Despite the ongoing efforts of the bioassay group, continued incomplete data limits our ability to understand the molecular roots of human disease. Therefore, it is urgent to develop a computational method to predict PPIs from the perspective of molecular system. METHODS: In this paper, a highly efficient computational model, MTV-PPI, is proposed for PPI prediction based on a heterogeneous molecular network by learning inter-view protein sequences and intra-view interactions between molecules simultaneously. On the one hand, the inter-view feature is extracted from the protein sequence by k-mer method. On the other hand, we use a popular embedding method LINE to encode the heterogeneous molecular network to obtain the intra-view feature. Thus, the protein representation used in MTV-PPI is constructed by the aggregation of its inter-view feature and intra-view feature. Finally, random forest is integrated to predict potential PPIs. RESULTS: To prove the effectiveness of MTV-PPI, we conduct extensive experiments on a collected heterogeneous molecular network with the accuracy of 86.55%, sensitivity of 82.49%, precision of 89.79%, AUC of 0.9301 and AUPR of 0.9308. Further comparison experiments are performed with various protein representations and classifiers to indicate the effectiveness of MTV-PPI in predicting PPIs based on a complex network. CONCLUSION: The achieved experimental results illustrate that MTV-PPI is a promising tool for PPI prediction, which may provide a new perspective for the future interactions prediction researches based on heterogeneous molecular network. Xiao-Rui Su 0001, Lun Hu, Zhu-Hong You, Pengwei Hu 0001, Bo-Wei Zhao |
BMC Bioinform. | 3 |
| 2022 | Identifying Protein Complexes From Protein-Protein Interaction Networks Based on Fuzzy Clustering and GO Semantic InformationabstractProtein complexes are of great significance to provide valuable insights into the mechanisms of biological processes of proteins. A variety of computational algorithms have thus been proposed to identify protein complexes in a protein-protein interaction network. However, few of them can perform their tasks by taking into account both network topology and protein attribute information in a unified fuzzy-based clustering framework. Since proteins in the same complex are similar in terms of their attribute information and the consideration of fuzzy clustering can also make it possible for us to identify overlapping complexes, we target to propose such a novel fuzzy-based clustering framework, namely FCAN-PCI, for an improved identification accuracy. To do so, the semantic similarity between the attribute information of proteins is calculated and we then integrate it into a well-established fuzzy clustering model together with the network topology. After that, a momentum method is adopted to accelerate the clustering procedure. FCAN-PCI finally applies a heuristical search strategy to identify overlapping protein complexes. A series of extensive experiments have been conducted to evaluate the performance of FCAN-PCI by comparing it with state-of-the-art identification algorithms and the results demonstrate the promising performance of FCAN-PCI. Xiangyu Pan, Lun Hu, Pengwei Hu 0001, Zhu-Hong You |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2022 | NSECDA: Natural Semantic Enhancement for CircRNA-Disease Association PredictionabstractIncreasing evidence suggest that circRNA, as one of the most promising emerging biomarkers, has a very close relationship with diseases. Exploring the relationship between circRNA and diseases can provide novel perspective for diseases diagnosis and pathogenesis. The existing circRNA-disease association (CDA) prediction models, however, generally treat the data attributes equally, do not pay special attention to the attributes with more significant influence, and do not make full use of the correlation and symbiosis between attributes to dig into the latent semantic information of the data. Therefore, in response to the above problems, this paper proposes a natural semantic enhancement method NSECDA to predict CDA. In practical terms, we first recognize the circRNA sequence as a biological language, and analyze its natural semantic properties through the natural language understanding theory; then integrate it with disease attributes, circRNA and disease Gaussian Interaction Profile (GIP) kernel attributes, and use Graph Attention Network (GAT) to focus on the influential attributes, so as to mine the deeply hidden features; finally, the Rotation Forest (RoF) classifier was used to accurately determine CDA. In the gold standard data set CircR2Disease, NSECDA achieved 92.49% accuracy with 0.9225 AUC score. In comparison with the non-natural semantic enhancement model and other classifier models, NSECDA also shows competitive performance. Additionally, 25 of the CDA pairs with unknown associations in the top 30 prediction scores of NSECDA have been proven by newly reported studies. These achievements suggest that NSECDA is an effective model to predict CDA, which can provide credible candidate for subsequent wet experiments, thus significantly reducing the scope of investigations. Lei Wang 0121, Leon Wong, Zhu-Hong You, De-Shuang Huang, Xiao-Rui Su 0001, Bo-Wei Zhao |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | Predicting miRNA-Disease Associations via a New MeSH Headings Representation of Diseases and eXtreme Gradient Boosting
Zhu-Hong You, Lei Wang 0121, Leon Wong, Xiao-Rui Su 0001, Bo-Wei Zhao |
ICIC (3) | 2 |
| 2021 | Computational Prediction of Protein-Protein Interactions in Plants Using Only Sequence Information
Jie Pan 0007, Liping Li 0003, Zhu-Hong You, Zhong-Hao Ren, Yongjian Guan |
ICIC (1) | 4 |
| 2021 | Protein-Protein Interaction Prediction by Integrating Sequence Information and Heterogeneous Network Representation
Xiao-Rui Su 0001, Zhu-Hong You, Zhen-Hao Guo |
ICIC (3) | 2 |
| 2021 | Detection of Drug-Drug Interactions Through Knowledge Graph Integrating Multi-attention with Capsule Network
Xiao-Rui Su 0001, Zhu-Hong You, Bo-Wei Zhao |
ICIC (3) | 2 |
| 2021 | Weighted Nonnegative Matrix Factorization Based on Multi-source Fusion Information for Predicting CircRNA-Disease Associations
Meineng Wang, Xue-Jun Xie, Zhu-Hong You, Leon Wong, Liping Li 0003 |
ICIC (3) | 3 |
| 2021 | CNNEMS: Using Convolutional Neural Networks to Predict Drug-Target Interactions by Combining Protein Evolution and Molecular Structures Information
Zhu-Hong You, Lei Wang 0121, Peng-Peng Chen |
ICIC (3) | 2 |
| 2021 | A Multi-graph Deep Learning Model for Predicting Drug-Disease Associations
Bo-Wei Zhao, Zhu-Hong You, Lun Hu, Leon Wong, Ping Zhang 0027 |
ICIC (3) | 2 |
| 2021 | MeSHHeading2vec: a new method for representing MeSH headings as vectors based on graph embedding algorithmabstractEffectively representing Medical Subject Headings (MeSH) headings (terms) such as disease and drug as discriminative vectors could greatly improve the performance of downstream computational prediction models. However, these terms are often abstract and difficult to quantify. In this paper, we converted the MeSH tree structure into a relationship network and applied several graph embedding algorithms on it to represent these terms. Specifically, the relationship network consisting of nodes (MeSH headings) and edges (relationships), which can be constructed by the tree num. Then, five graph embedding algorithms including DeepWalk, LINE, SDNE, LAP and HOPE were implemented on the relationship network to represent MeSH headings as vectors. In order to evaluate the performance of the proposed methods, we carried out the node classification and relationship prediction tasks. The results show that the MeSH headings characterized by graph embedding algorithms can not only be treated as an independent carrier for representation, but also can be utilized as additional information to enhance the representation ability of vectors. Thus, it can serve as an input and continue to play a significant role in any computational models related to disease, drug, microbe, etc. Besides, our method holds great hope to inspire relevant researchers to study the representation of terms in this network perspective. Zhen-Hao Guo, Zhu-Hong You, De-Shuang Huang, Kai Zheng 0020 |
Briefings Bioinform. | 2 |
| 2021 | A survey on computational models for predicting protein-protein interactionsabstractProteins interact with each other to play critical roles in many biological processes in cells. Although promising, laboratory experiments usually suffer from the disadvantages of being time-consuming and labor-intensive. The results obtained are often not robust and considerably uncertain. Due recently to advances in high-throughput technologies, a large amount of proteomics data has been collected and this presents a significant opportunity and also a challenge to develop computational models to predict protein-protein interactions (PPIs) based on these data. In this paper, we present a comprehensive survey of the recent efforts that have been made towards the development of effective computational models for PPI prediction. The survey introduces the algorithms that can be used to learn computational models for predicting PPIs, and it classifies these models into different categories. To understand their relative merits, the paper discusses different validation schemes and metrics to evaluate the prediction performance. Biological databases that are commonly used in different experiments for performance comparison are also described and their use in a series of extensive experiments to compare different prediction models are discussed. Finally, we present some open issues in PPI prediction for future work. We explain how the performance of PPI prediction can be improved if these issues are effectively tackled. Lun Hu, Pengwei Hu 0001, Zhu-Hong You |
Briefings Bioinform. | 5 |
| 2021 | Predicting microRNA-disease associations from lncRNA-microRNA interactions via Multiview Multitask LearningabstractMOTIVATION: Identifying microRNAs that are associated with different diseases as biomarkers is a problem of great medical significance. Existing computational methods for uncovering such microRNA-diseases associations (MDAs) are mostly developed under the assumption that similar microRNAs tend to associate with similar diseases. Since such an assumption is not always valid, these methods may not always be applicable to all kinds of MDAs. Considering that the relationship between long noncoding RNA (lncRNA) and different diseases and the co-regulation relationships between the biological functions of lncRNA and microRNA have been established, we propose here a multiview multitask method to make use of the known lncRNA-microRNA interaction to predict MDAs on a large scale. The investigation is performed in the absence of complete information of microRNAs and any similarity measurement for it and to the best knowledge, the work represents the first ever attempt to discover MDAs based on lncRNA-microRNA interactions. RESULTS: In this paper, we propose to develop a deep learning model called MVMTMDA that can create a multiview representation of microRNAs. The model is trained based on an end-to-end multitasking approach to machine learning so that, based on it, missing data in the side information can be determined automatically. Experimental results show that the proposed model yields an average area under ROC curve of 0.8410+/-0.018, 0.8512+/-0.012 and 0.8521+/-0.008 when k is set to 2, 5 and 10, respectively. In addition, we also propose here a statistical approach to predicting lncRNA-disease associations based on these associations and the MDA discovered using MVMTMDA. AVAILABILITY: Python code and the datasets used in our studies are made available at https://github.com/yahuang1991polyu/MVMTMDA/. Keith C. C. Chan, Zhu-Hong You, Pengwei Hu 0001, Lei Wang 0121, Zhi-an Huang |
Briefings Bioinform. | 3 |
| 2021 | A graph auto-encoder model for miRNA-disease associations predictionabstractEmerging evidence indicates that the abnormal expression of miRNAs involves in the evolution and progression of various human complex diseases. Identifying disease-related miRNAs as new biomarkers can promote the development of disease pathology and clinical medicine. However, designing biological experiments to validate disease-related miRNAs is usually time-consuming and expensive. Therefore, it is urgent to design effective computational methods for predicting potential miRNA-disease associations. Inspired by the great progress of graph neural networks in link prediction, we propose a novel graph auto-encoder model, named GAEMDA, to identify the potential miRNA-disease associations in an end-to-end manner. More specifically, the GAEMDA model applies a graph neural networks-based encoder, which contains aggregator function and multi-layer perceptron for aggregating nodes' neighborhood information, to generate the low-dimensional embeddings of miRNA and disease nodes and realize the effective fusion of heterogeneous information. Then, the embeddings of miRNA and disease nodes are fed into a bilinear decoder to identify the potential links between miRNA and disease nodes. The experimental results indicate that GAEMDA achieves the average area under the curve of $93.56\pm 0.44\%$ under 5-fold cross-validation. Besides, we further carried out case studies on colon neoplasms, esophageal neoplasms and kidney neoplasms. As a result, 48 of the top 50 predicted miRNAs associated with these diseases are confirmed by the database of differentially expressed miRNAs in human cancers and microRNA deregulation in human disease database, respectively. The satisfactory prediction performance suggests that GAEMDA model could serve as a reliable tool to guide the following researches on the regulatory role of miRNAs. Besides, the source codes are available at https://github.com/chimianbuhetang/GAEMDA. Zhengwei Li 0001, Jiashu Li, Ru Nie, Zhu-Hong You, Wenzheng Bao |
Briefings Bioinform. | 4 |
| 2021 | SGANRDA: semi-supervised generative adversarial networks for predicting circRNA-disease associationsabstractEmerging research shows that circular RNA (circRNA) plays a crucial role in the diagnosis, occurrence and prognosis of complex human diseases. Compared with traditional biological experiments, the computational method of fusing multi-source biological data to identify the association between circRNA and disease can effectively reduce cost and save time. Considering the limitations of existing computational models, we propose a semi-supervised generative adversarial network (GAN) model SGANRDA for predicting circRNA-disease association. This model first fused the natural language features of the circRNA sequence and the features of disease semantics, circRNA and disease Gaussian interaction profile kernel, and then used all circRNA-disease pairs to pre-train the GAN network, and fine-tune the network parameters through labeled samples. Finally, the extreme learning machine classifier is employed to obtain the prediction result. Compared with the previous supervision model, SGANRDA innovatively introduced circRNA sequences and utilized all the information of circRNA-disease pairs during the pre-training process. This step can increase the information content of the feature to some extent and reduce the impact of too few known associations on the model performance. SGANRDA obtained AUC scores of 0.9411 and 0.9223 in leave-one-out cross-validation and 5-fold cross-validation, respectively. Prediction results on the benchmark dataset show that SGANRDA outperforms other existing models. In addition, 25 of the top 30 circRNA-disease pairs with the highest scores of SGANRDA in case studies were verified by recent literature. These experimental results demonstrate that SGANRDA is a useful model to predict the circRNA-disease association and can provide reliable candidates for biological experiments. Lei Wang 0121, Zhu-Hong You, Xi Zhou 0007 |
Briefings Bioinform. | 3 |
| 2021 | HiSCF: leveraging higher-order structures for clustering analysis in biological networksabstractMOTIVATION: Clustering analysis in a biological network is to group biological entities into functional modules, thus providing valuable insight into the understanding of complex biological systems. Existing clustering techniques make use of lower-order connectivity patterns at the level of individual biological entities and their connections, but few of them can take into account of higher-order connectivity patterns at the level of small network motifs. RESULTS: Here, we present a novel clustering framework, namely HiSCF, to identify functional modules based on the higher-order structure information available in a biological network. Taking advantage of higher-order Markov stochastic process, HiSCF is able to perform the clustering analysis by exploiting a variety of network motifs. When compared with several state-of-the-art clustering models, HiSCF yields the best performance for two practical clustering applications, i.e. protein complex identification and gene co-expression module detection, in terms of accuracy. The promising performance of HiSCF demonstrates that the consideration of higher-order network motifs gains new insight into the analysis of biological networks, such as the identification of overlapping protein complexes and the inference of new signaling pathways, and also reveals the rich higher-order organizational structures presented in biological networks. AVAILABILITY AND IMPLEMENTATION: HiSCF is available at https://github.com/allenv5/HiSCF. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lun Hu, Jun Zhang 0003, Xiangyu Pan, Hong Yan 0001, Zhu-Hong You |
Bioinform. | 5 |
| 2021 | A learning-based method to predict LncRNA-disease associations by combining CNN and ELMabstractBACKGROUND: lncRNAs play a critical role in numerous biological processes and life activities, especially diseases. Considering that traditional wet experiments for identifying uncovered lncRNA-disease associations is limited in terms of time consumption and labor cost. It is imperative to construct reliable and efficient computational models as addition for practice. Deep learning technologies have been proved to make impressive contributions in many areas, but the feasibility of it in bioinformatics has not been adequately verified. RESULTS: In this paper, a machine learning-based model called LDACE was proposed to predict potential lncRNA-disease associations by combining Extreme Learning Machine (ELM) and Convolutional Neural Network (CNN). Specifically, the representation vectors are constructed by integrating multiple types of biology information including functional similarity and semantic similarity. Then, CNN is applied to mine both local and global features. Finally, ELM is chosen to carry out the prediction task to detect the potential lncRNA-disease associations. The proposed method achieved remarkable Area Under Receiver Operating Characteristic Curve of 0.9086 in Leave-one-out cross-validation and 0.8994 in fivefold cross-validation, respectively. In addition, 2 kinds of case studies based on lung cancer and endometrial cancer indicate the robustness and efficiency of LDACE even in a real environment. CONCLUSIONS: Substantial results demonstrated that the proposed model is expected to be an auxiliary tool to guide and assist biomedical research, and the close integration of deep learning and biology big data will provide life sciences with novel insights. Zhen-Hao Guo, Zhu-Hong You, Meineng Wang |
BMC Bioinform. | 3 |
| 2021 | In silico drug repositioning using deep learning and comprehensive similarity measuresabstractBACKGROUND: Drug repositioning, meanings finding new uses for existing drugs, which can accelerate the processing of new drugs research and development. Various computational methods have been presented to predict novel drug-disease associations for drug repositioning based on similarity measures among drugs and diseases. However, there are some known associations between drugs and diseases that previous studies not utilized. METHODS: In this work, we develop a deep gated recurrent units model to predict potential drug-disease interactions using comprehensive similarity measures and Gaussian interaction profile kernel. More specifically, the similarity measure is used to exploit discriminative feature for drugs based on their chemical fingerprints. Meanwhile, the Gaussian interactions profile kernel is employed to obtain efficient feature of diseases based on known disease-disease associations. Then, a deep gated recurrent units model is developed to predict potential drug-disease interactions. RESULTS: The performance of the proposed model is evaluated on two benchmark datasets under tenfold cross-validation. And to further verify the predictive ability, case studies for predicting new potential indications of drugs were carried out. CONCLUSION: The experimental results proved the proposed model is a useful tool for predicting new indications for drugs or new treatments for diseases, and can accelerate drug repositioning and related drug research and discovery. Zhu-Hong You, Lei Wang 0065, Xiao-Rui Su 0001, Xi Zhou 0007, Tonghai Jiang |
BMC Bioinform. | 2 |
| 2021 | A computational approach for predicting drug-target interactions from protein sequence and drug substructure fingerprint informationabstractIdentification of drug–target interactions (DTIs) is critical for discovering potential target protein candidates for new drugs. However, traditional experimental methods have limitations in discovering DTIs. They are time-consuming, tedious, and expensive, and often suffer from high false-positive rates and false-negative rates. Therefore, using computational methods to predict DTIs has received extensive attention from many researchers in recent years. To address this issue, in this paper, an effective prediction model is presented which is based on the information of drug molecular structure data and protein sequence data. It performs prediction with the following procedures. First, we transform the sequences of each target into a position-specific scoring matrix (PSSM), such that the features can retain biological evolutionary information. We then use a feature vector of molecular substructure fingerprints to describe the chemical structure information of the drug compounds. Second, the Legendre moments algorithm is used to extract new features from the PSSM. Finally, a classification algorithm called rotation forest is used to perform prediction, we tested its prediction performance on four golden standard data sets: enzymes, G-protein-coupled receptors, ion channels, and nuclear receptors. As a result, the proposed method achieves average accuracies of 0.9026, 0.8260, 0.8703, and 0.7444 on these four data sets using five-fold cross-validation. We also compare the proposed method with the support vector machine and other existing approaches. The proposed model is proved to be superior to comparative methods, showing that it is feasible, effective, and robust for predicting potential DTI. Yang Li 0111, Xiaozhang Liu, Zhu-Hong You, Liping Li 0003, Jian-Xin Guo, Zheng Wang 0065 |
Int. J. Intell. Syst. | 3 |
| 2021 | LDGRNMF: LncRNA-disease associations prediction based on graph regularized non-negative matrix factorization
Meineng Wang, Zhu-Hong You, Lei Wang 0121, Liping Li 0003, Kai Zheng 0020 |
Neurocomputing | 2 |
| 2021 | Multi-Neighborhood Learning for Global Alignment in Biological NetworksabstractThe global alignment of biological networks (GABN) aims to find an optimal alignment between proteins across species, such that both the biological structures and the topological structures of the proteins are maximally conserved. The research on GABN has attracted great attention due to its applications on species evolution, orthology detection and genetic analyses. Most of the existing methods for GABN are difficult to obtain a good tradeoff between the conservation of the biological structures and topological structures. In this paper, we propose a multi-neighborhood learning method for solving GABN (called as CLMNA). CLMNA first models GABN as an optimization of a weighted similarity which evaluates the conserved biological and topological similarities of an alignment, and then it combines a first-proximity, second-proximity and individual-aware proximity learning algorithm to solve the modeled problem. Finally, systematic experiments on 10 pairs of biological networks across 5 species show the superiority of CLMNA over the state-of-the-art network alignment algorithms. They also validate the effectiveness of CLMNA as a refinement method on improving the performance of the compared algorithms. Lijia Ma, Shiqiang Wang 0003, Qiuzhen Lin, Jianqiang Li 0001, Zhu-Hong You, Jiaxiang Huang, Maoguo Gong |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2021 | Learning Representation of Molecules in Association Network for Predicting Intermolecular AssociationsabstractA key aim of post-genomic biomedical research is to systematically understand molecules and their interactions in human cells. Multiple biomolecules coordinate to sustain life activities, and interactions between various biomolecules are interconnected. However, existing studies usually only focusing on associations between two or very limited types of molecules. In this study, we propose a network representation learning based computational framework MAN-SDNE to predict any intermolecular associations. More specifically, we constructed a large-scale molecular association network of multiple biomolecules in human by integrating associations among long non-coding RNA, microRNA, protein, drug, and disease, containing 6,528 molecular nodes, 9 kind of,105,546 associations. And then, the feature of each node is represented by its network proximity and attribute features. Furthermore, these features are used to train Random Forest classifier to predict intermolecular associations. MAN-SDNE achieves a remarkable performance with an AUC of 0.9552 and an AUPR of 0.9338 under five-fold cross-validation. To indicate the ability to predict specific types of interactions, a case study for predicting lncRNA-protein interactions using MAN-SDNE is also executed. Experimental results demonstrate this work offers a systematic insight for understanding the synergistic associations between molecules and complex diseases and provides a network-based computational tool to systematically explore intermolecular interactions. Zhu-Hong You, Zhen-Hao Guo, De-Shuang Huang, Keith C. C. Chan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | MISSIM: An Incremental Learning-Based Model With Applications to the Prediction of miRNA-Disease AssociationabstractIn the past few years, the prediction models have shown remarkable performance in most biological correlation prediction tasks. These tasks traditionally use a fixed dataset, and the model, once trained, is deployed as is. These models often encounter training issues such as sensitivity to hyperparameter tuning and "catastrophic forgetting" when adding new data. However, with the development of biomedicine and the accumulation of biological data, new predictive models are required to face the challenge of adapting to change. To this end, we propose a computational approach based on Broad learning system (BLS) to predict potential disease-associated miRNAs that retain the ability to distinguish prior training associations when new data need to be adapted. In particular, we are introducing incremental learning to the field of biological association prediction for the first time and proposed a new method for quantifying sequence similarity. In the performance evaluation, the AUC in the 5-fold cross-validation was 0.9400 +/- 0.0041. To better assess the effectiveness of MISSIM, we compared it with various classifiers and former prediction models. Its performance is superior to the previous method. Besides, the case study on identifying miRNAs associated with breast neoplasms, lung neoplasms and esophageal neoplasms show that 34, 36 and 35 out of the top 40 associations predicted by MISSIM are confirmed by recent biomedical resources. These results provide ample convincing evidence of this approach have potential value and prospect in promoting biomedical research productivity. Kai Zheng 0020, Zhu-Hong You, Lei Wang 0121, Ji-Ren Zhou, Haitao Zeng |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | IMS-CDA: Prediction of CircRNA-Disease Associations From the Integration of Multisource Similarity Information With Deep Stacked Autoencoder ModelabstractEmerging evidence indicates that circular RNA (circRNA) has been an indispensable role in the pathogenesis of human complex diseases and many critical biological processes. Using circRNA as a molecular marker or therapeutic target opens up a new avenue for our treatment and detection of human complex diseases. The traditional biological experiments, however, are usually limited to small scale and are time consuming, so the development of an effective and feasible computational-based approach for predicting circRNA-disease associations is increasingly favored. In this study, we propose a new computational-based method, called IMS-CDA, to predict potential circRNA-disease associations based on multisource biological information. More specifically, IMS-CDA combines the information from the disease semantic similarity, the Jaccard and Gaussian interaction profile kernel similarity of disease and circRNA, and extracts the hidden features using the stacked autoencoder (SAE) algorithm of deep learning. After training in the rotation forest (RF) classifier, IMS-CDA achieves 88.08% area under the ROC curve with 88.36% accuracy at the sensitivity of 91.38% on the CIRCR2Disease dataset. Compared with the state-of-the-art support vector machine and K -nearest neighbor models and different descriptor models, IMS-CDA achieves the best overall performance. In the case studies, eight of the top 15 circRNA-disease associations with the highest prediction score were confirmed by recent literature. These results indicated that IMS-CDA has an outstanding ability to predict new circRNA-disease associations and can provide reliable candidates for biological experiments. Lei Wang 0121, Zhu-Hong You, Jianqiang Li 0001 |
IEEE Trans. Cybern. | 2 |
| 2020 | Prediction of LncRNA-Disease Associations Based on Network Representation LearningabstractMassive observations have indicated that long noncoding RNAs (lncRNAs) are crucial in a number of biological processes and associated with various human diseases. Developing an efficient calculation model to predict the associations between lncRNA and diseases is not only beneficial to disease diagnosis, treatment, prognosis and potential drug targets in drug discovery, but also avoid the waste of human and material resources brought by biological experiments. In this paper, we proposed a novel prediction of lncRNA-disease associations based on complex and comprehensive molecular associations network (MAN), which integrated nine kinds of interactions among five molecules, including lncRNA, miRNA, disease, drug and protein. Network embedding Node2vec method was applied to extract behavior feature from MAN to generate a low-dimension vector containing nodes and edges information. After implementing 5-fold cross validation, the proposed method yielded good prediction performance with an average Accuracy of 91.91%, Sensitivity of 94.05%, Specificity of 89.76%, Precision of 90.21%, MCC value of 83.91%, AUC value of 0.9746 and AUPR of 0.9693. Comparative experiment indicates the behavior feature extracted by Node2vec is more representative than attribute features of lncRNA adopted 3-mer and diseases extracted by semantic similarity. Moreover, breast cancer, colon cancer and lung cancer are explored in case study. As a results, more than half of top 5 interactions are successfully confirmed for each disease by other datasets. Based on these reliable results, it is anticipated that proposed model is feasible and effective to predict lncRNA-disease associations at a global molecules level, which is a new respective for future biomedical researches. Xiao-Rui Su 0001, Zhu-Hong You |
BIBM | 2 |
| 2020 | Predicting Drug-Target Interactions by Node2vec Node Embedding in Molecular Associations Network
Zhu-Hong You, Zhen-Hao Guo, Gong-Xu Luo |
ICIC (2) | 2 |
| 2020 | Inferring Drug-miRNA Associations by Integrating Drug SMILES and MiRNA Sequence Information
Zhen-Hao Guo, Zhu-Hong You, Liping Li 0003 |
ICIC (2) | 2 |
| 2020 | Identification of Autistic Risk Genes Using Developmental Brain Gene Expression Data
Zhi-an Huang, Zhu-Hong You, Shanwen Zhang, Wenzhun Huang |
ICIC (2) | 3 |
| 2020 | A MapReduce-Based Parallel Random Forest Approach for Predicting Large-Scale Protein-Protein Interactions
Zhu-Hong You, Ji-Ren Zhou, Pengwei Hu 0001 |
ICIC (3) | 2 |
| 2020 | A Highly Efficient Biomolecular Network Representation Model for Predicting Drug-Disease Associations
Hanjing Jiang, Zhu-Hong You, Lun Hu, Zhen-Hao Guo, Leon Wong |
ICIC (3) | 2 |
| 2020 | A Network Embedding-Based Method for Predicting miRNA-Disease Associations by Integrating Multiple Information
Zhu-Hong You, Zhengwei Li 0001, Ji-Ren Zhou, Pengwei Hu 0001 |
ICIC (3) | 2 |
| 2020 | Predicting Protein-Protein Interactions from Protein Sequence Information Using Dual-Tree Complex Wavelet Transform
Jie Pan 0007, Zhu-Hong You, Liping Li 0003, Xinke Zhan |
ICIC (2) | 2 |
| 2020 | Prediction of lncRNA-Disease Associations from Heterogeneous Information Network Based on DeepWalk Embedding Model
Xiao-Yu Song, Ze-Yang Qiu, Zhu-Hong You, Li-Ting Jin, Xiao-Bei Feng, Lin Zhu 0008 |
ICIC (3) | 4 |
| 2020 | A Novel Computational Approach for Predicting Drug-Target Interactions via Network Representation Learning
Xiao-Rui Su 0001, Zhu-Hong You, Ji-Ren Zhou, Xiao Li 0007 |
ICIC (2) | 2 |
| 2020 | Embracing Disease Progression with a Learning System for Real World Evidence Discovery
Zefang Tang, Lun Hu, Xu Min, Jing Mei, Kenney Ng, Shaochun Li, Pengwei Hu 0001, Zhu-Hong You |
ICIC (2) | 9 |
| 2020 | WGMFDDA: A Novel Weighted-Based Graph Regularized Matrix Factorization for Predicting Drug-Disease Associations
Meineng Wang, Zhu-Hong You, Liping Li 0003, Xue-Jun Xie |
ICIC (3) | 2 |
| 2020 | GCNSP: A Novel Prediction Method of Self-Interacting Proteins Based on Graph Convolutional Networks
Lei Wang 0121, Zhu-Hong You, Kai Zheng 0020, Zhengwei Li 0001 |
ICIC (2) | 2 |
| 2020 | A Gaussian Kernel Similarity-Based Linear Optimization Model for Predicting miRNA-lncRNA Interactions
Leon Wong, Zhu-Hong You, Xi Zhou 0007, Mei-Yuan Cao |
ICIC (2) | 2 |
| 2020 | DTIFS: A Novel Computational Approach for Predicting Drug-Target Interactions from Drug Structure and Protein Sequence
Zhu-Hong You, Lei Wang 0121, Liping Li 0003, Kai Zheng 0020, Meineng Wang |
ICIC (2) | 2 |
| 2020 | A Unified Deep Biological Sequence Representation Learning with Pretrained Encoder-Decoder Model
Zhu-Hong You, Xiao-Rui Su 0001, De-Shuang Huang, Zhen-Hao Guo |
ICIC (2) | 2 |
| 2020 | Predicting Protein-Protein Interactions from Protein Sequence Using Locality Preserving Projections and Rotation Forest
Xinke Zhan, Zhu-Hong You, Jie Pan 0007 |
ICIC (2) | 2 |
| 2020 | A Novel Computational Method for Predicting LncRNA-Disease Associations from Heterogeneous Information Network with SDNE Embedding Model
Ping Zhang 0027, Bo-Wei Zhao, Leon Wong, Zhu-Hong You, Zhen-Hao Guo |
ICIC (2) | 4 |
| 2020 | Predicting LncRNA-miRNA Interactions via Network Embedding with Integrated Structure and Attribute Information
Bo-Wei Zhao, Ping Zhang 0027, Zhu-Hong You, Ji-Ren Zhou, Xiao Li 0007 |
ICIC (2) | 3 |
| 2020 | Predicting Human Disease-Associated piRNAs Based on Multi-source Information and Random Forest
Kai Zheng 0020, Zhu-Hong You, Lei Wang 0121 |
ICIC (2) | 2 |
| 2020 | Inferring Disease-Associated Piwi-Interacting RNAs via Graph Attention Networks
Kai Zheng 0020, Zhu-Hong You, Lei Wang 0121, Leon Wong |
ICIC (2) | 2 |
| 2020 | Prediction of lncRNA-miRNA Interactions via an Embedding Learning Graph Factorize Over Heterogeneous Information Network
Ji-Ren Zhou, Zhu-Hong You, Xi Zhou 0007 |
ICIC (2) | 2 |
| 2020 | Graph convolution for predicting associations between miRNA and drug resistanceabstractMOTIVATION: MicroRNA (miRNA) therapeutics is becoming increasingly important. However, aberrant expression of miRNAs is known to cause drug resistance and can become an obstacle for miRNA-based therapeutics. At present, little is known about associations between miRNA and drug resistance and there is no computational tool available for predicting such association relationship. Since it is known that miRNAs can regulate genes that encode specific proteins that are keys for drug efficacy, we propose here a computational approach, called GCMDR, for finding a three-layer latent factor model that can be used to predict miRNA-drug resistance associations. RESULTS: In this paper, we discuss how the problem of predicting such associations can be formulated as a link prediction problem involving a bipartite attributed graph. GCMDR makes use of the technique of graph convolution to build a latent factor model, which can effectively utilize information of high-dimensional attributes of miRNA/drug in an end-to-end learning scheme. In addition, GCMDR also learns graph embedding features for miRNAs and drugs. We leveraged the data from multiple databases storing miRNA expression profile, drug substructure fingerprints, gene ontology and disease ontology. The test for performance shows that the GCMDR prediction model can achieve AUCs of 0.9301 ± 0.0005, 0.9359 ± 0.0006 and 0.9369 ± 0.0003 based on 2-fold, 5-fold and 10-fold cross validation, respectively. Using this model, we show that the associations between miRNA and drug resistance can be reliably predicted by properly introducing useful side information like miRNA expression profile and drug structure fingerprints. AVAILABILITY AND IMPLEMENTATION: Python codes and dataset are available at https://github.com/yahuang1991polyu/GCMDR/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Pengwei Hu 0001, Keith C. C. Chan, Zhu-Hong You |
Bioinform. | 4 |
| 2020 | An efficient approach based on multi-sources information to predict circRNA-disease associations using deep convolutional neural networkabstractMOTIVATION: Emerging evidence indicates that circular RNA (circRNA) plays a crucial role in human disease. Using circRNA as biomarker gives rise to a new perspective regarding our diagnosing of diseases and understanding of disease pathogenesis. However, detection of circRNA-disease associations by biological experiments alone is often blind, limited to small scale, high cost and time consuming. Therefore, there is an urgent need for reliable computational methods to rapidly infer the potential circRNA-disease associations on a large scale and to provide the most promising candidates for biological experiments. RESULTS: In this article, we propose an efficient computational method based on multi-source information combined with deep convolutional neural network (CNN) to predict circRNA-disease associations. The method first fuses multi-source information including disease semantic similarity, disease Gaussian interaction profile kernel similarity and circRNA Gaussian interaction profile kernel similarity, and then extracts its hidden deep feature through the CNN and finally sends them to the extreme learning machine classifier for prediction. The 5-fold cross-validation results show that the proposed method achieves 87.21% prediction accuracy with 88.50% sensitivity at the area under the curve of 86.67% on the CIRCR2Disease dataset. In comparison with the state-of-the-art SVM classifier and other feature extraction methods on the same dataset, the proposed model achieves the best results. In addition, we also obtained experimental support for prediction results by searching published literature. As a result, 7 of the top 15 circRNA-disease pairs with the highest scores were confirmed by literature. These results demonstrate that the proposed model is a suitable method for predicting circRNA-disease associations and can provide reliable candidates for biological experiments. AVAILABILITY AND IMPLEMENTATION: The source code and datasets explored in this work are available at https://github.com/look0012/circRNA-Disease-association. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lei Wang 0121, Zhu-Hong You, De-Shuang Huang, Keith C. C. Chan |
Bioinform. | 2 |
| 2020 | NEMPD: a network embedding-based method for predicting miRNA-disease associations by preserving behavior and attribute informationabstractBACKGROUND: As an important non-coding RNA, microRNA (miRNA) plays a significant role in a series of life processes and is closely associated with a variety of Human diseases. Hence, identification of potential miRNA-disease associations can make great contributions to the research and treatment of Human diseases. However, to our knowledge, many existing computational methods only utilize the single type of known association information between miRNAs and diseases to predict their potential associations, without focusing on their interactions or associations with other types of molecules. RESULTS: In this paper, we propose a network embedding-based method for predicting miRNA-disease associations by preserving behavior and attribute information. Firstly, a heterogeneous network is constructed by integrating known associations among miRNA, protein and disease, and the network representation method Learning Graph Representations with Global Structural Information (GraRep) is implemented to learn the behavior information of miRNAs and diseases in the network. Then, the behavior information of miRNAs and diseases is combined with the attribute information of them to represent miRNA-disease association pairs. Finally, the prediction model is established based on the Random Forest algorithm. Under the five-fold cross validation, the proposed NEMPD model obtained average 85.41% prediction accuracy with 80.96% sensitivity at the AUC of 91.58%. Furthermore, the performance of NEMPD is also validated by the case studies. Among the top 50 predicted disease-related miRNAs, 48 (breast neoplasms), 47 (colon neoplasms), 47 (lung neoplasms) were confirmed by two other databases. CONCLUSIONS: The proposed NEMPD model has a good performance in predicting the potential associations between miRNAs and diseases, and has great potency in the field of miRNA-disease association prediction in the future. Zhu-Hong You, Leon Wong |
BMC Bioinform. | 2 |
| 2020 | RPI-SE: a stacking ensemble learning framework for ncRNA-protein interactions prediction using sequence informationabstractBACKGROUND: The interactions between non-coding RNAs (ncRNA) and proteins play an essential role in many biological processes. Several high-throughput experimental methods have been applied to detect ncRNA-protein interactions. However, these methods are time-consuming and expensive. Accurate and efficient computational methods can assist and accelerate the study of ncRNA-protein interactions. RESULTS: In this work, we develop a stacking ensemble computational framework, RPI-SE, for effectively predicting ncRNA-protein interactions. More specifically, to fully exploit protein and RNA sequence feature, Position Weight Matrix combined with Legendre Moments is applied to obtain protein evolutionary information. Meanwhile, k-mer sparse matrix is employed to extract efficient feature of ncRNA sequences. Finally, an ensemble learning framework integrated different types of base classifier is developed to predict ncRNA-protein interactions using these discriminative features. The accuracy and robustness of RPI-SE was evaluated on three benchmark data sets under five-fold cross-validation and compared with other state-of-the-art methods. CONCLUSIONS: The results demonstrate that RPI-SE is competent for ncRNA-protein interactions prediction task with high accuracy and robustness. It's anticipated that this work can provide a computational prediction tool to advance ncRNA-protein interactions related biomedical research. Zhu-Hong You, Meineng Wang, Zhen-Hao Guo, Ji-Ren Zhou |
BMC Bioinform. | 2 |
| 2020 | A survey of current trends in computational predictions of protein-protein interactions
Zhu-Hong You, Liping Li 0003 |
Frontiers Comput. Sci. | 2 |
| 2020 | GCNCDA: A new method for predicting circRNA-disease associations based on Graph Convolutional Network AlgorithmabstractNumerous evidences indicate that Circular RNAs (circRNAs) are widely involved in the occurrence and development of diseases. Identifying the association between circRNAs and diseases plays a crucial role in exploring the pathogenesis of complex diseases and improving the diagnosis and treatment of diseases. However, due to the complex mechanisms between circRNAs and diseases, it is expensive and time-consuming to discover the new circRNA-disease associations by biological experiment. Therefore, there is increasingly urgent need for utilizing the computational methods to predict novel circRNA-disease associations. In this study, we propose a computational method called GCNCDA based on the deep learning Fast learning with Graph Convolutional Networks (FastGCN) algorithm to predict the potential disease-associated circRNAs. Specifically, the method first forms the unified descriptor by fusing disease semantic similarity information, disease and circRNA Gaussian Interaction Profile (GIP) kernel similarity information based on known circRNA-disease associations. The FastGCN algorithm is then used to objectively extract the high-level features contained in the fusion descriptor. Finally, the new circRNA-disease associations are accurately predicted by the Forest by Penalizing Attributes (Forest PA) classifier. The 5-fold cross-validation experiment of GCNCDA achieved 91.2% accuracy with 92.78% sensitivity at the AUC of 90.90% on circR2Disease benchmark dataset. In comparison with different classifier models, feature extraction models and other state-of-the-art methods, GCNCDA shows strong competitiveness. Furthermore, we conducted case study experiments on diseases including breast cancer, glioma and colorectal cancer. The results showed that 16, 15 and 17 of the top 20 candidate circRNAs with the highest prediction scores were respectively confirmed by relevant literature and databases. These results suggest that GCNCDA can effectively predict potential circRNA-disease associations and provide highly credible candidates for biological experiments. Lei Wang 0121, Zhu-Hong You, Yang-Ming Li, Kai Zheng 0020 |
PLoS Comput. Biol. | 2 |
| 2020 | iCDA-CGR: Identification of circRNA-disease associations based on Chaos Game RepresentationabstractFound in recent research, tumor cell invasion, proliferation, or other biological processes are controlled by circular RNA. Understanding the association between circRNAs and diseases is an important way to explore the pathogenesis of complex diseases and promote disease-targeted therapy. Most methods, such as k-mer and PSSM, based on the analysis of high-throughput expression data have the tendency to think functionally similar nucleic acid lack direct linear homology regardless of positional information and only quantify nonlinear sequence relationships. However, in many complex diseases, the sequence nonlinear relationship between the pathogenic nucleic acid and ordinary nucleic acid is not much different. Therefore, the analysis of positional information expression can help to predict the complex associations between circRNA and disease. To fill up this gap, we propose a new method, named iCDA-CGR, to predict the circRNA-disease associations. In particular, we introduce circRNA sequence information and quantifies the sequence nonlinear relationship of circRNA by Chaos Game Representation (CGR) technology based on the biological sequence position information for the first time in the circRNA-disease prediction model. In the cross-validation experiment, our method achieved 0.8533 AUC, which was significantly higher than other existing methods. In the validation of independent data sets including circ2Disease, circRNADisease and CRDD, the prediction accuracy of iCDA-CGR reached 95.18%, 90.64% and 95.89%. Moreover, in the case studies, 19 of the top 30 circRNA-disease associations predicted by iCDA-CGR on circRDisease dataset were confirmed by newly published literature. These results demonstrated that iCDA-CGR has outstanding robustness and stability, and can provide highly credible candidates for biological experiments. Kai Zheng 0020, Zhu-Hong You, Jianqiang Li 0001, Lei Wang 0121, Zhen-Hao Guo |
PLoS Comput. Biol. | 2 |
| 2020 | Learning Multimodal Networks From Heterogeneous Data for Prediction of lncRNA-miRNA InteractionsabstractLong noncoding RNAs (lncRNAs) is an important class of non-protein coding RNAs. They have recently been found to potentially be able to act as a regulatory molecule in some important biological processes. MicroRNAs (miRNAs) have been confirmed to be closely related to the regulation of various human diseases. Recent studies have suggested that lncRNAs could interact with miRNAs to modulate their regulatory roles. Hence, predicting lncRNA-miRNA interactions are biologically significant due to their potential roles in determining the effectiveness of diagnostic biomarkers and therapeutic targets for various human diseases. For the details of the mechanisms to be better understood, it would be useful if some computational approaches are developed to allow for such investigations. As diverse heterogeneous datasets for describing lncRNA and miRNA have been made available, it becomes more feasible for us to develop a model to describe potential interactions between lncRNAs and miRNAs. In this work, we present a novel computational approach called LMNLMI for such purpose. LMNLMI works in several phases. First, it learns patterns from expression, sequences and functional data. Based on the patterns, it then constructs several networks including an expression-similarity network, a functional-similarity network, and a sequence-similarity network. Based on a measure of similarities between these networks, LMNLMI computes an interaction score for each pair of lncRNA and miRNA in the database. The novelty of LMNLMI lies in the use of a network fusion technique to combine the patterns inherent in multiple similarity networks and a matrix completion technique in predicting interaction relationships. Using a set of real data, we show that LMNLMI can be a very effective approach for the accurate prediction of lncRNA-miRNA interactions. Pengwei Hu 0001, Keith C. C. Chan, Zhu-Hong You |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2020 | Incorporating the Coevolving Information of Substrates in Predicting HIV-1 Protease Cleavage SitesabstractHuman immunodeficiency virus 1 (HIV-1) protease (PR) plays a crucial role in the maturation of the virus. The study of substrate specificity of HIV-1 PR as a new endeavor strives to increase our ability to understand how HIV-1 PR recognizes its various cleavage sites. To predict HIV-1 PR cleavage sites, most of the existing approaches have been developed solely based on the homogeneity of substrate sequence information with supervised classification techniques. Although efficient, these approaches are found to be restricted to the ability of explaining their results and probably provide few insights into the mechanisms by which HIV-1 PR cleaves the substrates in a site-specific manner. In this work, a coevolutionary pattern-based prediction model for HIV-1 PR cleavage sites, namely EvoCleave, is proposed by integrating the coevolving information obtained from substrate sequences with a linear SVM classifier. The experiment results showed that EvoCleave yielded a very promising performance in terms of ROC analysis and f-measure. We also prospectively assessed the biological significance of coevolutionary patterns by applying them to study three fundamental issues of HIV-1 PR cleavage site. The analysis results demonstrated that the coevolutionary patterns offered valuable insights into the understanding of substrate specificity of HIV-1 PR. Lun Hu, Pengwei Hu 0001, Xin Luo 0001, Zhu-Hong You |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2020 | Using Weighted Extreme Learning Machine Combined With Scale-Invariant Feature Transform to Predict Protein-Protein Interactions From Protein Evolutionary InformationabstractProtein-Protein Interactions (PPIs) play an irreplaceable role in biological activities of organisms. Although many high-throughput methods are used to identify PPIs from different kinds of organisms, they have some shortcomings, such as high cost and time-consuming. To solve the above problems, computational methods are developed to predict PPIs. Thus, in this paper, we present a method to predict PPIs using protein sequences. First, protein sequences are transformed into Position Weight Matrix (PWM), in which Scale-Invariant Feature Transform (SIFT) algorithm is used to extract features. Then Principal Component Analysis (PCA) is applied to reduce the dimension of features. At last, Weighted Extreme Learning Machine (WELM) classifier is employed to predict PPIs and a series of evaluation results are obtained. In our method, since SIFT and WELM are used to extract features and classify respectively, we called the proposed method SIFT-WELM. When applying the proposed method on three well-known PPIs datasets of Yeast, Human and Helicobacter.pylori, the average accuracies of our method using five-fold cross validation are obtained as high as 94.83, 97.60 and 83.64 percent, respectively. In order to evaluate the proposed approach properly, we compare it with Support Vector Machine (SVM) classifier and other recent-developed methods in different aspects. Moreover, the training time of our method is greatly shortened, which is obviously superior to the previous methods, such as SVM, ACC, PCVMZM and so on. Jianqiang Li 0001, Zhu-Hong You, Zhuangzhuang Chen, Qiuzhen Lin |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2020 | Combining High Speed ELM Learning with a Deep Convolutional Neural Network Feature Encoding for Predicting Protein-RNA InteractionsabstractEmerging evidence has shown that RNA plays a crucial role in many cellular processes, and their biological functions are primarily achieved by binding with a variety of proteins. High-throughput biological experiments provide a lot of valuable information for the initial identification of RNA-protein interactions (RPIs), but with the increasing complexity of RPIs networks, this method gradually falls into expensive and time-consuming situations. Therefore, there is an urgent need for high speed and reliable methods to predict RNA-protein interactions. In this study, we propose a computational method for predicting the RNA-protein interactions using sequence information. The deep learning convolution neural network (CNN) algorithm is utilized to mine the hidden high-level discriminative features from the RNA and protein sequences and feed it into the extreme learning machine (ELM) classifier. The experimental results with 5-fold cross-validation indicate that the proposed method achieves superior performance on benchmark datasets (RPI1807, RPI2241, and RPI369) with the accuracy of 98.83, 90.83, and 85.63 percent, respectively. We further evaluate the performance of the proposed model by comparing it with the state-of-the-art SVM classifier and other existing methods on the same benchmark data set. In addition, we predicted the independent NPInter v2.0 data set using the model trained on RPI369. The experimental results show that our model can serve as a useful tool for predicting RNA-protein interactions. Lei Wang 0121, Zhu-Hong You, De-Shuang Huang, Fengfeng Zhou |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2020 | Privacy-Preserving Global Structural Balance Computation in Signed NetworksabstractThe studies on signed networks have received a great attention due to their capabilities on presenting conflicting relationships, which reflect the potential conflicts and tensions of complex systems. To further understand those conflicts and tensions, many methods have been proposed for computing the global structural balance (GSB) of signed networks, which aim to discover the most balanced state of the networks with the least number of unbalanced links. However, most of them request full access to all information (structures, signs, and balance states) of links, which are usually sensitive and private. In this article, we propose a privacy-preserving GSB computation (PGSBC) framework, which aims to compute the GSB while preserving the privacy of the networks. The PGSBC first protects the sensitive information (structures, signs, and balance states) of links by using encryption techniques (the homomorphic cryptosystem and the random disturbances) and then computes the GSB of the signed networks on the encrypted structures. In the PGSBC, a balance-aware energy function is adopted to evaluate the balance degree of a clustering, while a fast two-level greedy algorithm (called as HM-Louvain) is presented to discover the most balanced clustering of signed networks. Simulation results on 11 LFR benchmark networks and 10 real signed networks show that the proposed framework can effectively compute the GSB of the networks while preserving the privacy of links’ sensitive information. Lijia Ma, Xiaopeng Huang, Jianqiang Li 0001, Qiuzhen Lin, Zhu-Hong You, Maoguo Gong, Victor C. M. Leung |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2019 | Predicting circRNA-disease associations using deep generative adversarial network based on multi-source fusion informationabstractCircular RNA (circRNA) is a kind of novel discovered non-coding RNA molecule with a closed loop structure, which plays a critical regulatory role in human diseases. Identifying the association between circRNAs and diseases has important potential value for the diagnosis and treatment of complex human diseases. Although biological experiments can more accurately identify the association between circRNAs and diseases, they are usually blind and limited by small scale and high cost. Therefore, there is an urgent need for efficient and feasible computational methods to predict the potential circRNA-disease associations on a large scale, so as to provide the most promising candidate for biological experiments. In this paper, we propose a novel computational method based on the deep Generative Adversarial Network (GAN) algorithm combined with the multi-source similarity information to predict the circRNA-disease associations. Firstly, we fuse the multi-source information of disease semantic similarity, disease and circRNA Gaussian interaction profile kernel similarity, and then use GAN to extract the hidden features of fusion information objectively and effectively in the way of confrontation learning, and finally send them to Logistic Model Tree (LMT) classifier for accurate prediction. The 5-fold cross-validation experiment of the proposed model achieved 89.2% accuracy with 89.4% precision at the AUC of 90.6% on the CIRCR2Disease dataset. Compared with the state-of-the-art SVM classifier and other feature extraction methods, the proposed model shows strong competitiveness. In addition, the predicted results of this model are supported by the biological experiments, and 9 of the top 15 circRNA-disease associations with the highest scores were confirmed by recently published literature. These promising results indicate that the proposed model is an effective tool for predicting circRNA-disease associations and can provide reliable candidates for biological experiments. Lei Wang 0121, Zhu-Hong You, Liping Li 0003, Kai Zheng 0020 |
BIBM | 2 |
| 2019 | Combining LSTM Network Model and Wavelet Transform for Predicting Self-interacting Proteins
Zhu-Hong You, Liping Li 0003, Zhen-Hao Guo, Pengwei Hu 0001, Hanjing Jiang |
ICIC (1) | 2 |
| 2019 | Combining High Speed ELM with a CNN Feature Encoding to Predict LncRNA-Disease Associations
Zhen-Hao Guo, Zhu-Hong You, Liping Li 0003 |
ICIC (2) | 2 |
| 2019 | Learning from Deep Representations of Multiple Networks for Predicting Drug-Target Interactions
Pengwei Hu 0001, Zhu-Hong You, Shaochun Li, Keith C. C. Chan, Henry Leung 0001, Lun Hu |
ICIC (2) | 3 |
| 2019 | Precise Prediction of Pathogenic Microorganisms Using 16S rRNA Gene Sequences
Zhi-an Huang, Zhu-Hong You, Pengwei Hu 0001, Liping Li 0003, Zhengwei Li 0001, Lei Wang 0121 |
ICIC (2) | 3 |
| 2019 | Predicting of Drug-Disease Associations via Sparse Auto-Encoder-Based Rotation Forest
Hanjing Jiang, Zhu-Hong You, Kai Zheng 0020 |
ICIC (3) | 2 |
| 2019 | LRMDA: Using Logistic Regression and Random Walk with Restart for MiRNA-Disease Association Prediction
Zhengwei Li 0001, Ru Nie, Zhu-Hong You |
ICIC (2) | 3 |
| 2019 | Combining Evolutionary Information and Sparse Bayesian Probability Model to Accurately Predict Self-interacting Proteins
Zhu-Hong You, Zhen-Hao Guo, Kai Zheng 0020 |
ICIC (2) | 2 |
| 2019 | A Gated Recurrent Unit Model for Drug Repositioning by Combining Comprehensive Similarity Measures and Gaussian Interaction Profile Kernel
Zhu-Hong You, Liping Li 0003, Lun Hu, Leon Wong |
ICIC (2) | 3 |
| 2019 | In Silico Identification of Anticancer Peptides with Stacking Heterogeneous Ensemble Learning Model and Sequence Information
Zhu-Hong You, Zhen-Hao Guo |
ICIC (2) | 2 |
| 2019 | An Efficient LightGBM Model to Predict Protein Self-interacting Using Chebyshev Moments and Bi-gram
Zhaohui Zhan, Zhu-Hong You, Yong Zhou 0003, Kai Zheng 0020, Zhengwei Li 0001 |
ICIC (2) | 2 |
| 2019 | MISSIM: Improved miRNA-Disease Association Prediction Model Based on Chaos Game Representation and Broad Learning System
Kai Zheng 0020, Zhu-Hong You, Lei Wang 0121, Hanjing Jiang |
ICIC (3) | 2 |
| 2019 | MicroRNAs and complex diseases: from experimental results to computational modelsabstractCircular RNAs (circRNAs) are a class of single-stranded, covalently closed RNA molecules with a variety of biological functions. Studies have shown that circRNAs are involved in a variety of biological processes and play an important role in the development of various complex diseases, so the identification of circRNA-disease associations would contribute to the diagnosis and treatment of diseases. In this review, we summarize the discovery, classifications and functions of circRNAs and introduce four important diseases associated with circRNAs. Then, we list some significant and publicly accessible databases containing comprehensive annotation resources of circRNAs and experimentally validated circRNA-disease associations. Next, we introduce some state-of-the-art computational models for predicting novel circRNA-disease associations and divide them into two categories, namely network algorithm-based and machine learning-based models. Subsequently, several evaluation methods of prediction performance of these computational models are summarized. Finally, we analyze the advantages and disadvantages of different types of computational models and provide some suggestions to promote the development of circRNA-disease association identification from the perspective of the construction of new computational models and the accumulation of circRNA-related data. Xing Chen 0001, Di Xie, Qi Zhao 0010, Zhu-Hong You |
Briefings Bioinform. | 4 |
| 2019 | Using discriminative vector machine model with 2DPCA to predict interactions among proteinsabstractBACKGROUND: The interactions among proteins act as crucial roles in most cellular processes. Despite enormous effort put for identifying protein-protein interactions (PPIs) from a large number of organisms, existing firsthand biological experimental methods are high cost, low efficiency, and high false-positive rate. The application of in silico methods opens new doors for predicting interactions among proteins, and has been attracted a great deal of attention in the last decades. RESULTS: Here we present a novelty computational model with the adoption of our proposed Discriminative Vector Machine (DVM) model and a 2-Dimensional Principal Component Analysis (2DPCA) descriptor to identify candidate PPIs only based on protein sequences. To be more specific, a 2DPCA descriptor is employed to capture discriminative feature information from Position-Specific Scoring Matrix (PSSM) of amino acid sequences by the tool of PSI-BLAST. Then, a robust and powerful DVM classifier is employed to infer PPIs. When applied on both gold benchmark datasets of Yeast and H. pylori, our model obtained mean prediction accuracies as high as of 97.06 and 92.89%, respectively, which demonstrates a noticeable improvement than some state-of-the-art methods. Moreover, we constructed Support Vector Machines (SVM) based predictive model and made comparison it with our model on Human benchmark dataset. In addition, to further demonstrate the predictive reliability of our proposed method, we also carried out extensive experiments for identifying cross-species PPIs on five other species datasets. CONCLUSIONS: All the experimental results indicate that our method is very effective for identifying potential PPIs and could serve as a practical approach to aid bioexperiment in proteomics research. Zhengwei Li 0001, Ru Nie, Zhu-Hong You, Chen Cao 0002, Jiashu Li |
BMC Bioinform. | 3 |
| 2019 | Plant disease leaf image segmentation based on superpixel clustering and EM algorithm
Shanwen Zhang, Zhu-Hong You, Xiaowei Wu 0003 |
Neural Comput. Appl. | 2 |
| 2019 | LMTRDA: Using logistic model tree to predict MiRNA-disease associations by fusing multi-source information of sequences and similaritiesabstractEmerging evidence has shown microRNAs (miRNAs) play an important role in human disease research. Identifying potential association among them is significant for the development of pathology, diagnose and therapy. However, only a tiny portion of all miRNA-disease pairs in the current datasets are experimentally validated. This prompts the development of high-precision computational methods to predict real interaction pairs. In this paper, we propose a new model of Logistic Model Tree for predicting miRNA-Disease Association (LMTRDA) by fusing multi-source information including miRNA sequences, miRNA functional similarity, disease semantic similarity, and known miRNA-disease associations. In particular, we introduce miRNA sequence information and extract its features using natural language processing technique for the first time in the miRNA-disease prediction model. In the cross-validation experiment, LMTRDA obtained 90.51% prediction accuracy with 92.55% sensitivity at the AUC of 90.54% on the HMDD V3.0 dataset. To further evaluate the performance of LMTRDA, we compared it with different classifier and feature descriptor models. In addition, we also validate the predictive ability of LMTRDA in human diseases including Breast Neoplasms, Breast Neoplasms and Lymphoma. As a result, 28, 27 and 26 out of the top 30 miRNAs associated with these diseases were verified by experiments in different kinds of case studies. These experimental results demonstrate that LMTRDA is a reliable model for predicting the association among miRNAs and diseases. Lei Wang 0121, Zhu-Hong You, Xing Chen 0001, Yang-Ming Li, Ya-Nan Dong, Liping Li 0003, Kai Zheng 0020 |
PLoS Comput. Biol. | 2 |
| 2019 | An Efficient Ensemble Learning Approach for Predicting Protein-Protein Interactions by Integrating Protein Primary Sequence and Evolutionary InformationabstractProtein-protein interactions (PPIs) perform a very important function in a number of cellular processes, including signal transduction, post-translational modifications, apoptosis, and cell growth. Deregulation of PPIs will lead to many diseases, including pernicious anemia or cancers. Although a large number of high-throughput techniques are designed to generate PPIs data, they are generally expensive, inefficient, and labor-intensive. Hence, there is an urgent need for developing a computational method to accurately and rapidly detect PPIs. In this article, we proposed a highly efficient method to detect PPIs by integrating a new protein sequence sub-stitution matrix feature representation and ensemble weighted sparse representation model classifier. The proposed method is demonstrated on Saccharomyces cerevisiae dataset and achieved 99.26 percent prediction accuracy with 98.53 percent sensitivity at precision of 100 percent, which is shown to have much higher predictive accuracy than the state-of-the-art methods. Extensive contrast experiments are performed with the benchmark data set from Human and Helicobacter pylori that our proposed method can achieve outstanding better success rates than other existing approaches in this problem. Experiment results illustrate that our proposed method presents an economical approach for computational building of PPI networks, which can be a helpful supplementary method for future proteomics researches. Zhu-Hong You, Wenzhun Huang, Shanwen Zhang, Liping Li 0003 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2019 | An Efficient Attribute-Based Encryption Scheme With Policy Update and File Update in Cloud ComputingabstractRecently, more and more users and enterprises have entrusted data storage and platform construction to proxy cloud service provider (PCSP) through cloud technology. Under this background, the attribute-based encryption (ABE) mechanism is an alternative to fill the drawbacks of the traditional encryption through flexible fine-grained access policy and collusion prevention. However, there exist some security issues when the access policy and file need to be updated in practical applications. And the ABE has the problems of excessive computation and storage costs. In this article, an efficient ciphertext-policy ABE scheme with policy update and file update is proposed in cloud computing. The ciphertext components generated by first encryption can be shared when the policy update and file update happens. It reduces the storage and communication costs of the client, and the computational cost of the PCSP. Moreover, the proposed scheme is proved to be secure under the assumption of decision q-parallel bilinear Diffie–Hellman exponent (BDHE). Finally, experimental simulation shows that the proposed scheme is highly efficient in terms of policy update and file update. Jianqiang Li 0001, Shulan Wang, Haiyan Wang 0009, Huihui Wang 0001, Jianyong Chen, Zhu-Hong You |
IEEE Trans. Ind. Informatics | 8 |
| 2019 | Protein-Protein Interactions Prediction via Multimodal Deep Polynomial Network and Regularized Extreme Learning MachineabstractPredicting the protein-protein interactions (PPIs) has played an important role in many applications. Hence, a novel computational method for PPIs prediction is highly desirable. PPIs endow with protein amino acid mutation rate and two physicochemical properties of protein (e.g., hydrophobicity and hydrophilicity). Deep polynomial network (DPN) is well-suited to integrate these modalities since it can represent any function on a finite sample dataset via the supervised deep learning algorithm. We propose a multimodal DPN (MDPN) algorithm to effectively integrate these modalities to enhance prediction performance. MDPN consists of a two-stage DPN, the first stage feeds multiple protein features into DPN encoding to obtain high-level feature representation while the second stage fuses and learns features by cascading three types of high-level features in the DPN encoding. We employ a regularized extreme learning machine to predict PPIs. The proposed method is tested on the public dataset of H. pylori, Human, and Yeast and achieves average accuracies of 97.87%, 99.90%, and 98.11%, respectively. The proposed method also achieves good accuracies on other datasets. Furthermore, we test our method on three kinds of PPI networks and obtain superior prediction results. Haijun Lei, Yuting Wen, Zhu-Hong You, Ahmed El-Azab, Ee-Leng Tan, Bai Ying Lei |
IEEE J. Biomed. Health Informatics | 3 |
| 2018 | Learning Latent Patterns in Molecular Data for Explainable Drug Side Effects Prediction
Pengwei Hu 0001, Zhu-Hong You, Tiantian He 0001, Shaochun Li, Shuhang Gu, Keith C. C. Chan |
BIBM | 2 |
| 2018 | RP-FIRF: Prediction of Self-interacting Proteins Using Random Projection Classifier Combining with Finite Impulse Response Filter
Zhu-Hong You, Liping Li 0003, Xiao Li 0007 |
ICIC (2) | 2 |
| 2018 | Discovering an Integrated Network in Heterogeneous Data for Predicting lncRNA-miRNA Interactions
Pengwei Hu 0001, Keith C. C. Chan, Zhu-Hong You |
ICIC (1) | 4 |
| 2018 | Using Weighted Extreme Learning Machine Combined with Scale-Invariant Feature Transform to Predict Protein-Protein Interactions from Protein Evolutionary Information
Jianqiang Li 0001, Zhu-Hong You, Zhuangzhuang Chen, Qiuzhen Lin |
ICIC (1) | 3 |
| 2018 | Efficient Framework for Predicting ncRNA-Protein Interactions Based on Sequence Information by Deep Learning
Zhaohui Zhan, Zhu-Hong You, Yong Zhou 0003, Liping Li 0003, Zhengwei Li 0001 |
ICIC (2) | 2 |
| 2018 | A novel approach based on KATZ measure to predict associations of human microbiota with non-infectious diseasesabstractBioinformatics (2017) 33 (5): 733–739. DOI: https://doi.org/10.1093/bioinformatics/btw715 The publisher wishes to inform readers that the footnote for the † symbol was erroneously removed from the paper as first published. The paper has now been corrected online to include the footnote: ‘†The authors wish it to be known that, in their opinion, the first two authors should be regarded as Joint First Authors’. Xing Chen 0001, Zhu-Hong You, Guiying Yan, Xuesong Wang 0001 |
Bioinform. | 3 |
| 2018 | BNPMDA: Bipartite Network Projection for MiRNA-Disease Association predictionabstractMotivation: A large number of resources have been devoted to exploring the associations between microRNAs (miRNAs) and diseases in the recent years. However, the experimental methods are expensive and time-consuming. Therefore, the computational methods to predict potential miRNA-disease associations have been paid increasing attention. Results: In this paper, we proposed a novel computational model of Bipartite Network Projection for MiRNA-Disease Association prediction (BNPMDA) based on the known miRNA-disease associations, integrated miRNA similarity and integrated disease similarity. We firstly described the preference degree of a miRNA for its related disease and the preference degree of a disease for its related miRNA with the bias ratings. We constructed bias ratings for miRNAs and diseases by using agglomerative hierarchical clustering according to the three types of networks. Then, we implemented the bipartite network recommendation algorithm to predict the potential miRNA-disease associations by assigning transfer weights to resource allocation links between miRNAs and diseases based on the bias ratings. BNPMDA had been shown to improve the prediction accuracy in comparison with previous models according to the area under the receiver operating characteristics (ROC) curve (AUC) results of three typical cross validations. As a result, the AUCs of Global LOOCV, Local LOOCV and 5-fold cross validation obtained by implementing BNPMDA were 0.9028, 0.8380 and 0.8980 ± 0.0013, respectively. We further implemented two types of case studies on several important human complex diseases to confirm the effectiveness of BNPMDA. In conclusion, BNPMDA could effectively predict the potential miRNA-disease associations at a high accuracy level. Availability and implementation: BNPMDA is available via http://www.escience.cn/system/file?fileId=99559. Supplementary information: Supplementary data are available at Bioinformatics online. Xing Chen 0001, Di Xie, Lei Wang 0121, Qi Zhao 0010, Zhu-Hong You, Hongsheng Liu 0001 |
Bioinform. | 5 |
| 2018 | Constructing prediction models from expression profiles for large scale lncRNA-miRNA interaction profilingabstractMotivation: The interaction of miRNA and lncRNA is known to be important for gene regulations. However, not many computational approaches have been developed to analyze known interactions and predict the unknown ones. Given that there are now more evidences that suggest that lncRNA-miRNA interactions are closely related to their relative expression levels in the form of a titration mechanism, we analyzed the patterns in large-scale expression profiles of known lncRNA-miRNA interactions. From these uncovered patterns, we noticed that lncRNAs tend to interact collaboratively with miRNAs of similar expression profiles, and vice versa. Results: By representing known interaction between lncRNA and miRNA as a bipartite graph, we propose here a technique, called EPLMI, to construct a prediction model from such a graph. EPLMI performs its tasks based on the assumption that lncRNAs that are highly similar to each other tend to have similar interaction or non-interaction patterns with miRNAs and vice versa. The effectiveness of the prediction model so constructed has been evaluated using the latest dataset of lncRNA-miRNA interactions. The results show that the prediction model can achieve AUCs of 0.8522 and 0.8447 ± 0.0017 based on leave-one-out cross validation and 5-fold cross validation. Using this model, we show that lncRNA-miRNA interactions can be reliably predicted. We also show that we can use it to select the most likely lncRNA targets that specific miRNAs would interact with. We believe that the prediction models discovered by EPLMI can yield great insights for further research on ceRNA regulation network. To the best of our knowledge, EPLMI is the first technique that is developed for large-scale lncRNA-miRNA interaction profiling. Availability and implementation: Matlab codes and dataset are available at https://github.com/yahuang1991polyu/EPLMI/. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Keith C. C. Chan, Zhu-Hong You |
Bioinform. | 3 |
| 2018 | DroidDet: Effective and robust detection of android malware using static analysis along with rotation forest model
Zhu-Hong You, Zexuan Zhu 0001, Wei-Lei Shi, Xing Chen 0001 |
Neurocomputing | 2 |
| 2018 | HEMD: a highly efficient random forest-based malware detection framework for Android
Tonghai Jiang, Bo Ma 0004, Zhu-Hong You, Wei-Lei Shi |
Neural Comput. Appl. | 4 |
| 2018 | An improved efficient rotation forest algorithm to predict the interactions among proteins
Lei Wang 0121, Zhu-Hong You, Shixiong Xia, Xing Chen 0001, Yong Zhou 0003, Feng Liu 0039 |
Soft Comput. | 2 |
| 2018 | Incorporation of Efficient Second-Order Solvers Into Latent Factor Models for Accurate Prediction of Missing QoS DataabstractGenerating highly accurate predictions for missing quality-of-service (QoS) data is an important issue. Latent factor (LF)-based QoS-predictors have proven to be effective in dealing with it. However, they are based on first-order solvers that cannot well address their target problem that is inherently bilinear and nonconvex, thereby leaving a significant opportunity for accuracy improvement. This paper proposes to incorporate an efficient second-order solver into them to raise their accuracy. To do so, we adopt the principle of Hessian-free optimization and successfully avoid the direct manipulation of a Hessian matrix, by employing the efficiently obtainable product between its Gauss-Newton approximation and an arbitrary vector. Thus, the second-order information is innovatively integrated into them. Experimental results on two industrial QoS datasets indicate that compared with the state-of-the-art predictors, the newly proposed one achieves significantly higher prediction accuracy at the expense of affordable computational burden. Hence, it is especially suitable for industrial applications requiring high prediction accuracy of unknown QoS data. Xin Luo 0001, MengChu Zhou, Shuai Li 0002, Yunni Xia, Zhu-Hong You, Qingsheng Zhu, Hareton K. N. Leung |
IEEE Trans. Cybern. | 5 |
| 2017 | Computational Methods for the Prediction of Drug-Target Interactions from Drug Fingerprints and Protein Sequences by Stacked Auto-Encoder Deep Neural Network
Lei Wang 0121, Zhu-Hong You, Xing Chen 0001, Shixiong Xia, Feng Liu 0039, Yong Zhou 0003 |
ISBRA | 2 |
| 2017 | Long non-coding RNAs and complex diseases: from experimental results to computational modelsabstractLncRNAs have attracted lots of attentions from researchers worldwide in recent decades. With the rapid advances in both experimental technology and computational prediction algorithm, thousands of lncRNA have been identified in eukaryotic organisms ranging from nematodes to humans in the past few years. More and more research evidences have indicated that lncRNAs are involved in almost the whole life cycle of cells through different mechanisms and play important roles in many critical biological processes. Therefore, it is not surprising that the mutations and dysregulations of lncRNAs would contribute to the development of various human complex diseases. In this review, we first made a brief introduction about the functions of lncRNAs, five important lncRNA-related diseases, five critical disease-related lncRNAs and some important publicly available lncRNA-related databases about sequence, expression, function, etc. Nowadays, only a limited number of lncRNAs have been experimentally reported to be related to human diseases. Therefore, analyzing available lncRNA-disease associations and predicting potential human lncRNA-disease associations have become important tasks of bioinformatics, which would benefit human complex diseases mechanism understanding at lncRNA level, disease biomarker detection and disease diagnosis, treatment, prognosis and prevention. Furthermore, we introduced some state-of-the-art computational models, which could be effectively used to identify disease-related lncRNAs on a large scale and select the most promising disease-related lncRNAs for experimental validation. We also analyzed the limitations of these models and discussed the future directions of developing computational models for lncRNA research. Xing Chen 0001, Chenggang Yan 0001, Xu Zhang 0028, Zhu-Hong You |
Briefings Bioinform. | 4 |
| 2017 | A novel approach based on KATZ measure to predict associations of human microbiota with non-infectious diseasesabstractMotivation: Accumulating clinical observations have indicated that microbes living in the human body are closely associated with a wide range of human noninfectious diseases, which provides promising insights into the complex disease mechanism understanding. Predicting microbe-disease associations could not only boost human disease diagnostic and prognostic, but also improve the new drug development. However, little efforts have been attempted to understand and predict human microbe-disease associations on a large scale until now. Results: In this work, we constructed a microbe-human disease association network and further developed a novel computational model of KATZ measure for Human Microbe-Disease Association prediction (KATZHMDA) based on the assumption that functionally similar microbes tend to have similar interaction and non-interaction patterns with noninfectious diseases, and vice versa. To our knowledge, KATZHMDA is the first tool for microbe-disease association prediction. The reliable prediction performance could be attributed to the use of KATZ measurement, and the introduction of Gaussian interaction profile kernel similarity for microbes and diseases. LOOCV and k-fold cross validation were implemented to evaluate the effectiveness of this novel computational model based on known microbe-disease associations obtained from HMDAD database. As a result, KATZHMDA achieved reliable performance with average AUCs of 0.8130 ± 0.0054, 0.8301 ± 0.0033 and 0.8382 in 2-fold and 5-fold cross validation and LOOCV framework, respectively. It is anticipated that KATZHMDA could be used to obtain more novel microbes associated with important noninfectious human diseases and therefore benefit drug discovery and human medical improvement. Availability and Implementation: Matlab codes and dataset explored in this work are available at http://dwz.cn/4oX5mS . Contacts: [email protected] or [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Xing Chen 0001, Zhu-Hong You, Guiying Yan, Xuesong Wang 0001 |
Bioinform. | 3 |
| 2017 | An improved sequence-based prediction protocol for protein-protein interactions using amino acids substitution matrix and rotation forest ensemble classifiers
Zhu-Hong You, Xiao Li 0007, Keith C. C. Chan |
Neurocomputing | 1 |
| 2017 | PBMDA: A novel and effective path-based computational model for miRNA-disease association predictionabstractIn the recent few years, an increasing number of studies have shown that microRNAs (miRNAs) play critical roles in many fundamental and important biological processes. As one of pathogenetic factors, the molecular mechanisms underlying human complex diseases still have not been completely understood from the perspective of miRNA. Predicting potential miRNA-disease associations makes important contributions to understanding the pathogenesis of diseases, developing new drugs, and formulating individualized diagnosis and treatment for diverse human complex diseases. Instead of only depending on expensive and time-consuming biological experiments, computational prediction models are effective by predicting potential miRNA-disease associations, prioritizing candidate miRNAs for the investigated diseases, and selecting those miRNAs with higher association probabilities for further experimental validation. In this study, Path-Based MiRNA-Disease Association (PBMDA) prediction model was proposed by integrating known human miRNA-disease associations, miRNA functional similarity, disease semantic similarity, and Gaussian interaction profile kernel similarity for miRNAs and diseases. This model constructed a heterogeneous graph consisting of three interlinked sub-graphs and further adopted depth-first search algorithm to infer potential miRNA-disease associations. As a result, PBMDA achieved reliable performance in the frameworks of both local and global LOOCV (AUCs of 0.8341 and 0.9169, respectively) and 5-fold cross validation (average AUC of 0.9172). In the cases studies of three important human diseases, 88% (Esophageal Neoplasms), 88% (Kidney Neoplasms) and 90% (Colon Neoplasms) of top-50 predicted miRNAs have been manually confirmed by previous experimental reports from literatures. Through the comparison performance between PBMDA and other previous models in case studies, the reliable performance also demonstrates that PBMDA could serve as a powerful computational tool to accelerate the identification of disease-miRNA associations. Zhu-Hong You, Zhi-an Huang, Zexuan Zhu 0001, Guiying Yan, Zhengwei Li 0001, Zhenkun Wen, Xing Chen 0001 |
PLoS Comput. Biol. | 1 |
| 2017 | PSPEL: In Silico Prediction of Self-Interacting Proteins from Amino Acids Sequences Using Ensemble LearningabstractSelf interacting proteins (SIPs) play an important role in various aspects of the structural and functional organization of the cell. Detecting SIPs is one of the most important issues in current molecular biology. Although a large number of SIPs data has been generated by experimental methods, wet laboratory approaches are both time-consuming and costly. In addition, they yield high false negative and positive rates. Thus, there is a great need for in silico methods to predict SIPs accurately and efficiently. In this study, a new sequence-based method is proposed to predict SIPs. The evolutionary information contained in Position-Specific Scoring Matrix (PSSM) is extracted from of protein with known sequence. Then, features are fed to an ensemble classifier to distinguish the self-interacting and non-self-interacting proteins. When performed on Saccharomyces cerevisiae and Human SIPs data sets, the proposed method can achieve high accuracies of 86.86 and 91.30 percent, respectively. Our method also shows a good performance when compared with the SVM classifier and previous methods. Consequently, the proposed method can be considered to be a novel promising tool to predict SIPs. Jianqiang Li 0001, Zhu-Hong You, Xiao Li 0007, Zhong Ming 0001, Xing Chen 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2017 | Identifying Spurious Interactions in the Protein-Protein Interaction Networks Using Local Similarity Preserving EmbeddingabstractIn recent years, a remarkable amount of protein-protein interaction (PPI) data are being available owing to the advance made in experimental high-throughput technologies. However, the experimentally detected PPI data usually contain a large amount of spurious links, which could contaminate the analysis of the biological significance of protein links and lead to incorrect biological discoveries, thereby posing new challenges to both computational and biological scientists. In this paper, we develop a new embedding algorithm called local similarity preserving embedding (LSPE) to rank the interaction possibility of protein links. By going beyond limitations of current geometric embedding methods for network denoising and emphasizing the local information of PPI networks, LSPE can avoid the unstableness of previous methods. We demonstrate experimental results on benchmark PPI networks and show that LSPE was the overall leader, outperforming the state-of-the-art methods in topological false links elimination problems. Lin Zhu 0008, Suping Deng, Zhu-Hong You, De-Shuang Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2017 | Highly Efficient Framework for Predicting Interactions Between ProteinsabstractProtein-protein interactions (PPIs) play a central role in many biological processes. Although a large amount of human PPI data has been generated by high-throughput experimental techniques, they are very limited compared to the estimated 130 000 protein interactions in humans. Hence, automatic methods for human PPI-detection are highly desired. This work proposes a novel framework, i.e., Low-rank approximation-kernel Extreme Learning Machine (LELM), for detecting human PPI from a protein's primary sequences automatically. It has three main steps: 1) mapping each protein sequence into a matrix built on all kinds of adjacent amino acids; 2) applying the low-rank approximation model to the obtained matrix to solve its lowest rank representation, which reflects its true subspace structures; and 3) utilizing a powerful kernel extreme learning machine to predict the probability for PPI based on this lowest rank representation. Experimental results on a large-scale human PPI dataset demonstrate that the proposed LELM has significant advantages in accuracy and efficiency over the state-of-art approaches. Hence, this work establishes a new and effective way for the automatic detection of PPI. Zhu-Hong You, MengChu Zhou, Xin Luo 0001, Shuai Li 0002 |
IEEE Trans. Cybern. | 1 |
| 2016 | Large-scale prediction of drug-target interactions from deep representationsabstractIdentifying drug-target interactions (DTIs) is a major challenge in drug development. Traditionally, similarity-based methods use drug and target similarity matrices to infer the potential drug-target interactions. But these techniques do not handle biochemical data directly. While recent feature-based methods reveal simple patterns of physicochemical properties, efficient method to study large interactive features and precisely predict interactions is still missing. Deep learning has been found to be an appropriate tool for converting high-dimensional features to low-dimensional representations. These deep representations generated from drug-protein pair can serve as training examples for the interaction predictor. In this paper, we propose a promising approach called multi-scale features deep representations inferring interactions (MFDR). We extract the large-scale chemical structure and protein sequence descriptors so as to machine learning model predict if certain human target protein can interact with a specific drug. MFDR use Auto-Encoders as building blocks of deep network for reconstruct drug and protein features to low-dimensional new representations. Then, we make use of support vector machine to infer the potential drug-target interaction from deep representations. The experiment result shows that a deep neural network with Stacked Auto-Encoders exactly output interactive representations for the DTIs prediction task. MFDR is able to predict large-scale drug-target interactions with high accuracy and achieves results better than other feature-based approaches. Peng-Wei, Keith C. C. Chan, Zhu-Hong You |
IJCNN | 3 |
| 2016 | Sequence-based prediction of protein-protein interactions using weighted sparse representation model combined with global encodingabstractBACKGROUND: Proteins are the important molecules which participate in virtually every aspect of cellular function within an organism in pairs. Although high-throughput technologies have generated considerable protein-protein interactions (PPIs) data for various species, the processes of experimental methods are both time-consuming and expensive. In addition, they are usually associated with high rates of both false positive and false negative results. Accordingly, a number of computational approaches have been developed to effectively and accurately predict protein interactions. However, most of these methods typically perform worse when other biological data sources (e.g., protein structure information, protein domains, or gene neighborhoods information) are not available. Therefore, it is very urgent to develop effective computational methods for prediction of PPIs solely using protein sequence information. RESULTS: In this study, we present a novel computational model combining weighted sparse representation based classifier (WSRC) and global encoding (GE) of amino acid sequence. Two kinds of protein descriptors, composition and transition, are extracted for representing each protein sequence. On the basis of such a feature representation, novel weighted sparse representation based classifier is introduced to predict protein interaction class. When the proposed method was evaluated with the PPIs data of S. cerevisiae, Human and H. pylori, it achieved high prediction accuracies of 96.82, 97.66 and 92.83 % respectively. Extensive experiments were performed for cross-species PPIs prediction and the prediction accuracies were also very promising. CONCLUSIONS: To further evaluate the performance of the proposed method, we then compared its performance with the method based on support vector machine (SVM). The results show that the proposed method achieved a significant improvement. Thus, the proposed method is a very efficient method to predict PPIs and may be a useful supplementary tool for future proteomics studies. Zhu-Hong You, Xing Chen 0001, Keith C. C. Chan, Xin Luo 0001 |
BMC Bioinform. | 2 |
| 2016 | Construction of reliable protein-protein interaction networks using weighted sparse representation based classifier with pseudo substitution matrix representation features
Zhu-Hong You, Xiao Li 0007, Xing Chen 0001, Pengwei Hu 0001, Shuai Li 0002, Xin Luo 0001 |
Neurocomputing | 2 |
| 2016 | An Incremental-and-Static-Combined Scheme for Matrix-Factorization-Based Collaborative FilteringabstractCollaborative filtering (CF)-based recommenders are achieved by matrix factorization (MF) to obtain high prediction accuracy and scalability. Most current MF-based models, however, are static ones that cannot adapt to incremental user feedbacks. This work aims to develop a general, incremental- and-static-combined scheme for MF-based CF to obtain highly accurate and computationally affordable incremental recommenders. With it, a recommender is designed to consist of two components, i.e., a static one built on static rating data, and an incremental one built on a sub-matrix related to rating-variations only. Highly reliable predictions are thus generated by fusing their results. The experiments on large industrial datasets show that desired accuracy and acceptable computational complexity are achieved by the resulting recommender with the proposed scheme. Xin Luo 0001, MengChu Zhou, Hareton K. N. Leung, Yunni Xia, Qingsheng Zhu, Zhu-Hong You, Shuai Li 0002 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2016 | Inverse-Free Extreme Learning Machine With Optimal Information UpdatingabstractThe extreme learning machine (ELM) has drawn insensitive research attentions due to its effectiveness in solving many machine learning problems. However, the matrix inversion operation involved in the algorithm is computational prohibitive and limits the wide applications of ELM in many scenarios. To overcome this problem, in this paper, we propose an inverse-free ELM to incrementally increase the number of hidden nodes, and update the connection weights progressively and optimally. Theoretical analysis proves the monotonic decrease of the training error with the proposed updating procedure and also proves the optimality in every updating step. Extensive numerical experiments show the effectiveness and accuracy of the proposed algorithm. Shuai Li 0002, Zhu-Hong You, Xin Luo 0001, Zhong-Qiu Zhao |
IEEE Trans. Cybern. | 2 |
| 2016 | A Nonnegative Latent Factor Model for Large-Scale Sparse Matrices in Recommender Systems via Alternating Direction MethodabstractNonnegative matrix factorization (NMF)-based models possess fine representativeness of a target matrix, which is critically important in collaborative filtering (CF)-based recommender systems. However, current NMF-based CF recommenders suffer from the problem of high computational and storage complexity, as well as slow convergence rate, which prevents them from industrial usage in context of big data. To address these issues, this paper proposes an alternating direction method (ADM)-based nonnegative latent factor (ANLF) model. The main idea is to implement the ADM-based optimization with regard to each single feature, to obtain high convergence rate as well as low complexity. Both computational and storage costs of ANLF are linear with the size of given data in the target matrix, which ensures high efficiency when dealing with extremely sparse matrices usually seen in CF problems. As demonstrated by the experiments on large, real data sets, ANLF also ensures fast convergence and high prediction accuracy, as well as the maintenance of nonnegativity constraints. Moreover, it is simple and easy to implement for real applications of learning systems. Xin Luo 0001, MengChu Zhou, Shuai Li 0002, Zhu-Hong You, Yunni Xia, Qingsheng Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2015 | Predicting Protein-Protein Interactions from Amino Acid Sequences Using SaE-ELM Combined with Continuous Wavelet Descriptor and PseAA Composition
Zhu-Hong You, Jianqiang Li 0001, Leon Wong, Shubin Cai |
ICIC (2) | 2 |
| 2015 | Detection of Protein-Protein Interactions from Amino Acid Sequences Using a Rotation Forest Model with a Novel PR-LPQ Descriptor
Leon Wong, Zhu-Hong You, Shuai Li 0002 |
ICIC (3) | 2 |
| 2015 | Improving network topology-based protein interactome mapping via collaborative filtering
Xin Luo 0001, Zhong Ming 0001, Zhu-Hong You, Shuai Li 0002, Yunni Xia, Hareton K. N. Leung |
Knowl. Based Syst. | 3 |
| 2015 | An Efficient Second-Order Approach to Factorize Sparse Matrices in Recommender SystemsabstractRecommender systems are an important kind of learning systems, which can be achieved by latent-factor (LF)-based collaborative filtering (CF) with high efficiency and scalability. LF-based CF models rely on an optimization process with respect to some desired latent features; however, most of them employ first-order optimization algorithms, e.g., gradient decent schemes, to conduct their optimization task, thereby failing in discovering patterns reflected by higher order information. This work proposes to build a new LF-based CF model via second-order optimization to achieve higher accuracy. We first investigate a Hessian-free optimization framework, and employ its principle to avoid direct usage of the Hessian matrix by computing its product with an arbitrary vector. We then propose the Hessian-free optimization-based LF model, which is able to extract latent factors from the given incomplete matrices via a second-order optimization process. Compared with LF models based on first-order optimization algorithms, experimental results on two industrial datasets show that the proposed one can offer higher prediction accuracy with reasonable computational efficiency. Hence, it is a promising model for implementing high-performance recommenders. Xin Luo 0001, MengChu Zhou, Shuai Li 0002, Yunni Xia, Zhu-Hong You, Qingsheng Zhu, Hareton K. N. Leung |
IEEE Trans. Ind. Informatics | 5 |
| 2014 | Using Chou's amphiphilic Pseudo-Amino Acid Composition and Extreme Learning Machine for prediction of Protein-protein interactionsabstractProtein-protein interactions (PPIs) play crucial roles in the execution of various cellular processes. Almost every cellular process relies on transient or permanent physical bindings of proteins. Unfortunately, the experimental methods for identifying PPIs are both time-consuming and expensive. Therefore, it is important to develop computational approaches for predicting PPIs. In this study, a novel approach is presented to predict PPIs using only the information of protein sequences. This method is developed based on learning algorithm-Extreme Learning Machine (ELM) combined with the concept of Chous Pseudo-Amino Acid Composition (PseAAC) composition. PseAAC is a combination of a set of discrete sequence correlation factors and the 20 components of the conventional amino acid composition, so this method can observe a remarkable improvement in prediction quality. ELM classifier is selected as prediction engine, which is a kind of accurate and fast-learning innovative classification method based on the random generation of the input-to-hidden-units weights followed by the resolution of the linear equations to obtain the hidden-to-output weights. When performed on the PPIs data of Saccharomyces cerevisiae, the proposed method achieved 79.66% prediction accuracy with 79.16% sensitivity at the precision of 79.96%. Extensive experiments are performed to compare our method with state-of-the-art techniques Support Vector Machine (SVM). Achieved results show that the proposed approach is very promising for predicting PPIs, and it can be a helpful supplement for PPIs prediction. Qiao-Ying Huang, Zhu-Hong You, Shuai Li 0002, Zexuan Zhu 0001 |
IJCNN | 2 |
| 2014 | Identifying Spurious Interactions in the Protein-Protein Interaction Networks Using Local Similarity Preserving Embedding
Lin Zhu 0008, Zhu-Hong You, De-Shuang Huang |
ISBRA | 2 |
| 2014 | Prediction of protein-protein interactions from amino acid sequences using a novel multi-scale continuous and discontinuous feature setabstractBACKGROUND: Identifying protein-protein interactions (PPIs) is essential for elucidating protein functions and understanding the molecular mechanisms inside the cell. However, the experimental methods for detecting PPIs are both time-consuming and expensive. Therefore, computational prediction of protein interactions are becoming increasingly popular, which can provide an inexpensive way of predicting the most likely set of interactions at the entire proteome scale, and can be used to complement experimental approaches. Although much progress has already been achieved in this direction, the problem is still far from being solved and new approaches are still required to overcome the limitations of the current prediction models. RESULTS: In this work, a sequence-based approach is developed by combining a novel Multi-scale Continuous and Discontinuous (MCD) feature representation and Support Vector Machine (SVM). The MCD representation gives adequate consideration to the interactions between sequentially distant but spatially close amino acid residues, thus it can sufficiently capture multiple overlapping continuous and discontinuous binding patterns within a protein sequence. An effective feature selection method mRMR was employed to construct an optimized and more discriminative feature set by excluding redundant features. Finally, a prediction model is trained and tested based on SVM algorithm to predict the interaction probability of protein pairs. CONCLUSIONS: When performed on the yeast PPIs data set, the proposed approach achieved 91.36% prediction accuracy with 91.94% precision at the sensitivity of 90.67%. Extensive experiments are conducted to compare our method with the existing sequence-based method. Experimental results show that the performance of our predictor is better than several other state-of-the-art predictors, whose average prediction accuracy is 84.91%, sensitivity is 83.24%, and precision is 86.12%. Achieved results show that the proposed approach is very promising for predicting PPI, so it can be a useful supplementary tool for future proteomics studies. The source code and the datasets are freely available at http://csse.szu.edu.cn/staff/youzh/MCDPPI.zip for academic use. Zhu-Hong You, Lin Zhu 0008, Chun-Hou Zheng 0001, Suping Deng, Zhen Ji |
BMC Bioinform. | 1 |
| 2014 | Orthogonal locally discriminant spline embedding for plant leaf recognition
Ying-Ke Lei, Ji-Wei Zou, Tianbao Dong, Zhu-Hong You, Yihua Hu 0001 |
Comput. Vis. Image Underst. | 4 |
| 2014 | Predicting dynamic deformation of retaining structure by LSSVR-based time series method
Zhiwei Ji, Bing Wang 0004, Suping Deng, Zhu-Hong You |
Neurocomputing | 4 |
| 2014 | A MapReduce based parallel SVM for large-scale predicting protein-protein interactions
Zhu-Hong You, Jian-Zhong Yu, Lin Zhu 0008, Shuai Li 0002, Zhenkun Wen |
Neurocomputing | 1 |
| 2013 | Research on Signaling Pathways Reconstruction by Integrating High Content RNAi Screening and Functional Gene Network
Zhu-Hong You, Zhong Ming 0001, Liping Li 0003, Qiao-Ying Huang |
ICIC (2) | 1 |
| 2013 | A SVM-Based System for Predicting Protein-Protein Interactions Using a Novel Representation of Protein Sequences
Zhu-Hong You, Zhong Ming 0001, Suping Deng, Zexuan Zhu 0001 |
ICIC (1) | 1 |
| 2013 | Prediction of protein-protein interactions from amino acid sequences with ensemble extreme learning machines and principal component analysisabstractBACKGROUND: Protein-protein interactions (PPIs) play crucial roles in the execution of various cellular processes and form the basis of biological mechanisms. Although large amount of PPIs data for different species has been generated by high-throughput experimental techniques, current PPI pairs obtained with experimental methods cover only a fraction of the complete PPI networks, and further, the experimental methods for identifying PPIs are both time-consuming and expensive. Hence, it is urgent and challenging to develop automated computational methods to efficiently and accurately predict PPIs. RESULTS: We present here a novel hierarchical PCA-EELM (principal component analysis-ensemble extreme learning machine) model to predict protein-protein interactions only using the information of protein sequences. In the proposed method, 11188 protein pairs retrieved from the DIP database were encoded into feature vectors by using four kinds of protein sequences information. Focusing on dimension reduction, an effective feature extraction method PCA was then employed to construct the most discriminative new feature set. Finally, multiple extreme learning machines were trained and then aggregated into a consensus classifier by majority voting. The ensembling of extreme learning machine removes the dependence of results on initial random weights and improves the prediction performance. CONCLUSIONS: When performed on the PPI data of Saccharomyces cerevisiae, the proposed method achieved 87.00% prediction accuracy with 86.15% sensitivity at the precision of 87.59%. Extensive experiments are performed to compare our method with state-of-the-art techniques Support Vector Machine (SVM). Experimental results demonstrate that proposed PCA-EELM outperforms the SVM method by 5-fold cross-validation. Besides, PCA-EELM performs faster than PCA-SVM based method. Consequently, the proposed approach can be considered as a new promising and powerful tools for predicting PPI with excellent performance and less time. Zhu-Hong You, Ying-Ke Lei, Lin Zhu 0008, Junfeng Xia, Bing Wang 0004 |
BMC Bioinform. | 1 |
| 2013 | Increasing the reliability of protein-protein interaction networks via non-convex semantic embedding
Lin Zhu 0008, Zhu-Hong You, De-Shuang Huang |
Neurocomputing | 2 |
| 2013 | Increasing reliability of protein interactome by fast manifold embedding
Ying-Ke Lei, Zhu-Hong You, Tianbao Dong, Yun-Xiao Jiang, Junan Yang |
Pattern Recognit. Lett. | 2 |
| 2012 | Assessing and predicting protein interactions by combining manifold embedding with multiple information integrationabstractBACKGROUND: Protein-protein interactions (PPIs) play crucial roles in virtually every aspect of cellular function within an organism. Over the last decade, the development of novel high-throughput techniques has resulted in enormous amounts of data and provided valuable resources for studying protein interactions. However, these high-throughput protein interaction data are often associated with high false positive and false negative rates. It is therefore highly desirable to develop scalable methods to identify these errors from the computational perspective. RESULTS: We have developed a robust computational technique for assessing the reliability of interactions and predicting new interactions by combining manifold embedding with multiple information integration. Validation of the proposed method was performed with extensive experiments on densely-connected and sparse PPI networks of yeast respectively. Results demonstrate that the interactions ranked top by our method have high functional homogeneity and localization coherence. CONCLUSIONS: Our proposed method achieves better performances than the existing methods no matter assessing or predicting protein interactions. Furthermore, our method is general enough to work over a variety of PPI networks irrespectively of densely-connected or sparse PPI network. Therefore, the proposed algorithm is a much more promising method to detect both false positive and false negative interactions in PPI networks. Ying-Ke Lei, Zhu-Hong You, Zhen Ji, Lin Zhu 0008, De-Shuang Huang |
BMC Bioinform. | 2 |
| 2010 | Increasing Reliability of Protein Interactome by Combining Heterogeneous Data Sources with Weighted Network Topological Metrics
Zhu-Hong You, Liping Li 0003, Sanfeng Chen, Shu-Lin Wang |
ICIC (1) | 1 |
| 2010 | Comparison of DNA Truncated Barcodes and Full-Barcodes for Species Identification
Zhu-Hong You |
ICIC (2) | 2 |
| 2010 | Using manifold embedding for assessing and predicting protein interactions from high-throughput experimental dataabstractMOTIVATION: High-throughput protein interaction data, with ever-increasing volume, are becoming the foundation of many biological discoveries, and thus high-quality protein-protein interaction (PPI) maps are critical for a deeper understanding of cellular processes. However, the unreliability and paucity of current available PPI data are key obstacles to the subsequent quantitative studies. It is therefore highly desirable to develop an approach to deal with these issues from the computational perspective. Most previous works for assessing and predicting protein interactions either need supporting evidences from multiple information resources or are severely impacted by the sparseness of PPI networks. RESULTS: We developed a robust manifold embedding technique for assessing the reliability of interactions and predicting new interactions, which purely utilizes the topological information of PPI networks and can work on a sparse input protein interactome without requiring additional information types. After transforming a given PPI network into a low-dimensional metric space using manifold embedding based on isometric feature mapping (ISOMAP), the problem of assessing and predicting protein interactions is recasted into the form of measuring similarity between points of its metric space. Then a reliability index, a likelihood indicating the interaction of two proteins, is assigned to each protein pair in the PPI networks based on the similarity between the points in the embedded space. Validation of the proposed method is performed with extensive experiments on densely connected and sparse PPI network of yeast, respectively. Results demonstrate that the interactions ranked top by our method have high-functional homogeneity and localization coherence, especially our method is very efficient for large sparse PPI network with which the traditional algorithms fail. Therefore, the proposed algorithm is a much more promising method to detect both false positive and false negative interactions in PPI networks. AVAILABILITY: MATLAB code implementing the algorithm is available from the web site http://home.ustc.edu.cn/∼yzh33108/Manifold.htm. Zhu-Hong You, Ying-Ke Lei, Jie Gui, De-Shuang Huang, Xiaobo Zhou 0001 |
Bioinform. | 1 |
| 2010 | A semi-supervised learning approach to predict synthetic genetic interactions by combining functional and topological properties of functional gene networkabstractBACKGROUND: Genetic interaction profiles are highly informative and helpful for understanding the functional linkages between genes, and therefore have been extensively exploited for annotating gene functions and dissecting specific pathway structures. However, our understanding is rather limited to the relationship between double concurrent perturbation and various higher level phenotypic changes, e.g. those in cells, tissues or organs. Modifier screens, such as synthetic genetic arrays (SGA) can help us to understand the phenotype caused by combined gene mutations. Unfortunately, exhaustive tests on all possible combined mutations in any genome are vulnerable to combinatorial explosion and are infeasible either technically or financially. Therefore, an accurate computational approach to predict genetic interaction is highly desirable, and such methods have the potential of alleviating the bottleneck on experiment design. RESULTS: In this work, we introduce a computational systems biology approach for the accurate prediction of pairwise synthetic genetic interactions (SGI). First, a high-coverage and high-precision functional gene network (FGN) is constructed by integrating protein-protein interaction (PPI), protein complex and gene expression data; then, a graph-based semi-supervised learning (SSL) classifier is utilized to identify SGI, where the topological properties of protein pairs in weighted FGN is used as input features of the classifier. We compare the proposed SSL method with the state-of-the-art supervised classifier, the support vector machines (SVM), on a benchmark dataset in S. cerevisiae to validate our method's ability to distinguish synthetic genetic interactions from non-interaction gene pairs. Experimental results show that the proposed method can accurately predict genetic interactions in S. cerevisiae (with a sensitivity of 92% and specificity of 91%). Noticeably, the SSL method is more efficient than SVM, especially for very small training sets and large test sets. CONCLUSIONS: We developed a graph-based SSL classifier for predicting the SGI. The classifier employs topological properties of weighted FGN as input features and simultaneously employs information induced from labelled and unlabelled data. Our analysis indicates that the topological properties of weighted FGN can be employed to accurately predict SGI. Also, the graph-based SSL method outperforms the traditional standard supervised approach, especially when used with small training sets. The proposed method can alleviate experimental burden of exhaustive test and provide a useful guide for the biologist in narrowing down the candidate gene pairs with SGI. The data and source code implementing the method are available from the website: http://home.ustc.edu.cn/~yzh33108/GeneticInterPred.htm. Zhu-Hong You, Zheng Yin, Kyungsook Han, De-Shuang Huang, Xiaobo Zhou 0001 |
BMC Bioinform. | 1 |
| 2009 | Integration of Genomic and Proteomic Data to Predict Synthetic Genetic Interactions Using Semi-supervised Learning
Zhu-Hong You, Shanwen Zhang, Liping Li 0003 |
ICIC (2) | 1 |
| 2008 | A Novel Hybrid Method of Gene Selection and Its Application on Tumor Classification
Zhu-Hong You, Shulin Wang, Jie Gui, Shanwen Zhang |
ICIC (2) | 1 |
| 2008 | An improvement on learning with local and global consistencyabstractA modified version for semi-supervised learning algorithm with local and global consistency was proposed in this paper. The new method adds the label information, and adopts the geodesic distance rather than Euclidean distance as the measure of the difference between two data points when conducting calculation. In addition we add class prior knowledge. It was found that the effect of class prior knowledge was different between under high label rate and low label rate. The experimental results show that the changes attain the satisfying classification performance better than the original algorithms. Jie Gui, De-Shuang Huang, Zhu-Hong You |
ICPR | 3 |