Ping Xuan

dblp:06/4918 · DBLP profile ↗
← Back
56ranked-venue papers
26as first author
46since 2021 · last 2026
0000-0001-5328-691XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 46 · 24 first-author · 38 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Multi-granularity transformer contrastive learning and feature reconstruction for prediction of disease-related miRNAs
Ping Xuan, Zhicheng Guo, Tiangang Zhang
BMC Bioinform.1
2026 Inferring drug-related microbes through multi-perspective node feature distribution encoding and multi-scale hypergraph learning
Fengjiao Sun, Sentao Chen, Hui Cui 0002, Ping Xuan, Tiangang Zhang
Eng. Appl. Artif. Intell.5
2026 Topology-enhanced hypergraph learning and adaptive multi-graph transformer for prediction of drug-related side effects
Ping Xuan, Xidong Yang, Sentao Chen, Hui Cui 0002, Zelong Xu, Qiangguo Jin, Tiangang Zhang
Expert Syst. Appl.1
2026 Open-Set Domain Adaptation by Joint Distribution Alignment and Unknown Risk Minimization
Lisheng Wen, Sentao Chen, Lin Zheng 0003, Ping Xuan
Pattern Recognit.4
2026 A Multi-Scale Neighbor Topology Guided Transformer and Kolmogorov-Arnold Network Enhanced Feature Learning Model for Disease-Related circRNA Prediction
abstract
As circular non-coding RNA (circRNA) is closely associated with various human diseases, identifying disease-related circRNAs can provide a deeper understanding of the mechanisms underlying disease pathogenesis. Advanced circRNA-disease association prediction methods mainly focus on graph learning techniques such as graph convolutional networks. However, these methods do not fully encode the multiscale neighbor topologies of each node, and the dependencies among the pairwise attributes. We propose a multi-scale neighbor topology-guided transformer with Kolmogorov-Arnold network (KAN) enhanced feature learning for circRNA and disease association prediction, termed MKCD. First, MKCD incorporates an adaptive multiscale neighbor topology embedding construction strategy (AMNE), which generates neighbor topologies covering varying scopes of neighbors by random walks. Second, we design a dynamic multi-scale neighbor topology-guided transformer (DMTT) that leverages the multi-scale neighbor topologies to guide the learning of relationships among circRNA, miRNA, and disease nodes. The multi-scale neighbor topology is dynamically evolved, providing adaptive guidance to the transformer's learning process. Third, we establish a feature-gated network (FGN) to evaluate the importance of topological features and the original node attributes. Finally, we propose an adaptive joint convolutional neural networks and KAN learning strategy (ACK) to learn the global and local dependencies of pairwise features. Comprehensive comparison experiments show that MKCD outperforms six state-of-the-art methods, improving AUC and AUPR by at least 14.1% and 7.6%, respectively. Ablation experiments further validate the effectiveness of AMNE, DMTT, FGN and ACK innovations. Case studies on three diseases further validate the application value of our method in discovering reliable circRNA candidates for the diseases.
Ping Xuan, Hui Cui 0002, Zelong Xu, Toshiya Nakaguchi, Tiangang Zhang
IEEE J. Biomed. Health Informatics1
2026 Open Set Domain Adaptation via Known Joint Distribution Matching and Unknown Classification Risk Reformulation
abstract
Open set domain adaptation (OSDA) is an important problem in machine learning and computer vision. In OSDA, one is given a labeled dataset from a source domain (source joint distribution) and an unlabeled dataset from a target domain (target joint distribution), where the target domain contains not only the known classes presented in the source domain but also the unknown class. The goal of OSDA is to train a neural network with minimal target classification risk. From the statistical learning perspective, there are two fundamental challenges in this problem: (1) the source-target joint distribution difference regarding the known classes and (2) the target classification risk estimation regarding the unknown class. Although prior works have proposed various sophisticated solutions to the problem and achieved inspiring experimental results, they do not fully resolve these two challenges. In this article, we introduce a principled approach named known joint distribution matching and unknown classification risk reformulation (KMUR). KMUR tackles the first challenge by matching the source joint distribution to the target known joint distribution such that the distribution difference can be reduced and addresses the second challenge by reformulating the target unknown classification risk such that the reformulated risk can be estimated on the unlabeled target and source data. To be specific, we exploit cross entropy as the classification loss and triangular discrimination (TD) distance as the joint distribution matching loss. Since the TD distance needs to be estimated from data, we develop an innovative technique named least squares TD estimation (LSTDE), which casts the estimation into least squares classification. To achieve the OSDA goal, we train the network to minimize the estimations of target classification risk and TD distance. Experiments on benchmark and real-world datasets confirm the effectiveness of our approach. The introductory video and PyTorch code are available on GitHub (https://github.com/sentaochen/Known-Joint-Distribution-Matching-and-Unknown-Classification-Risk-Reformulation). One can also visit https://github.com/sentaochen for more source code on domain adaptation (DA), multi-source DA, partial DA, and domain generalization approaches.
Sentao Chen, Ping Xuan, Lifang He 0001
IEEE Trans. Neural Networks Learn. Syst.2
2025 Structure-sensitive transformer and multi-view graph contrastive learning enhanced prediction of drug-related microbes
abstract
BACKGROUND: The human microbiome plays a crucial role in regulating the efficacy and toxicity of drugs as well as in developing the drugs. Therefore, predicting the drug-related microbes is beneficial for analyzing the functional mechanisms of drugs. Recently, the graph learning based methods demonstrated their advantages in extracting the node features from the biological heterogeneous graphs. However, the previous methods failed to completely preserve the intrinsic structures of biological data and did not fully utilize the topological and positional information for predicting the drug-microbe associations. RESULTS: We propose a new prediction model, structure-sensitive transformer and multi-view graph contrastive learning for microbe-drug association prediction (SMMDA), to encode and integrate the topological structures, semantics, and multiple-view embedding features of the drugs and microbes. Considering the sparsity of the original features of drugs and microbes, the learnable data augmentation strategy is designed to learn their global representations. Since similar drugs are more likely to associate with the similar microbes, a structure-sensitive transformer is proposed to integrate the topology structures composed of drugs (microbes) to form the multi-view embedding features. We design two contrastive learning strategies to exploit the complementary semantics across multiple views. As the embedding features from multiple views have various semantics, we design view-level attention to adaptively integrate these features. CONCLUSIONS: The extensive experimental results show that SMMDA outperforms several state-of-the-art methods for predicting the drug-related candidate microbes. The ablation studies show the effectiveness of the major innovations which include the learnable data augmentation, structure-sensitive transformer-based node feature learning, and multi-view contrastive learning. The case studies on three drugs also demonstrate SMMDA's capability in retrieving the potential microbe candidates for the drugs.
Ping Xuan, Hui Cui 0002, Tiangang Zhang
BMC Bioinform.1
2025 Robust subspace structure discovery for cell type identification in scRNA-seq data
abstract
Single-cell RNA sequencing (scRNA-seq) technology has transformed gene expression studies by enabling analysis at the individual cell level, offering unprecedented insights into cellular heterogeneity. A key challenge in scRNA-seq data analysis is cell type identification, which requires grouping cells with similar gene expression profiles using unsupervised clustering methods. However, the high dimensionality, inherent noise, and significant sparsity of scRNA-seq data present substantial obstacles to accurately determining relationships among cell samples. To address these challenges, we propose a novel deep subspace clustering approach for cell type identification that captures a more reliable subspace structure from scRNA-seq data. Our method leverages a robust self-representation learning framework to effectively characterize and learn the underlying cluster structure. This framework is optimized through an integrated strategy combining a structure-guided approach with an optimal transport algorithm, enhancing the robustness of the subspace clustering process. By mitigating the effects of noise and sparsity in scRNA-seq data, this approach enables more accurate cell clustering. Experimental results on 18 real scRNA-seq datasets demonstrate that our method outperforms several state-of-the-art clustering approaches tailored for scRNA-seq data, excelling in both accuracy and interpretability.
Xianyong Zhou, Xindian Wei, Cheng Liu 0001, Ping Xuan, Si Wu 0002, Hau-San Wong
BMC Bioinform.5
2025 PKDF-Net: Anticancer peptide prediction via a prior-knowledge-aware dual-path feature-entangled network
Qiangguo Jin, Ankang Wu, Leyi Wei, Hui Cui 0002, Ping Xuan, Xikang Feng, Ran Su
Eng. Appl. Artif. Intell.5
2025 Iterative pseudo-labeling based adaptive copy-paste supervision for semi-supervised tumor segmentation
Qiangguo Jin, Hui Cui 0002, Changming Sun, Yimiao He, Ping Xuan, Cong Cong 0001, Leyi Wei, Ran Su
Knowl. Based Syst.6
2025 Fourier Transform and Kolmogorov-Arnold Network Enhanced Relation-Aware Generative and Adversarial Network for miRNA-Disease Association Prediction
abstract
Identifying the miRNAs that are associated with diseases can assist to explore the pathogenesis of diseases. Traditional prediction methods primarily focus on integrating multi-sourced data related to miRNAs and diseases within euclidean space to infer potential candidate disease-related miRNAs. Research indicates that miRNAs belonging to the same family, residing in the same cluster, or sharing more common target proteins are more likely to be involved into similar disease processes. However, existing approaches have not fully integrated the family, cluster, and common target protein attributes of miRNAs, nor explored the low-frequency smoothness and high-frequency local fluctuation features of miRNA and disease nodes in the frequency domain. To overcome these issues, we propose a Fourier transform and Kolmogorov-Arnold network enhanced relation-aware generative adversarial network (FKRGAN) model. FKRGAN incorporates a Fourier transform enhanced dual-space feature learning (FDFL) strategy, which helps learn the topological features of miRNA and disease nodes in euclidean space, as well as their low-frequency and high-frequency characteristics in the frequency domain. Furthermore, a feature-level attention mechanism is designed to determine the significance of features learned from dual-space representations and the family, cluster, and target protein features derived from homogeneous graphs, thereby facilitating the adaptive fusion of these features. We develop a relation-aware generative adversarial network with Kolmogorov-Arnold networks (KRGAN), which enhances feature learning for each miRNA and disease node by employing generative and adversarial strategies. The generator, composed of multi-layer Kolmogorov-Arnold networks (KAN), fuses multiple connection relationships between miRNA and disease nodes during the generation process, effectively capturing nonlinear dependencies among node features to produce relation-aware node feature representations. Comparative experiments on public datasets demonstrate that our method outperforms eight state-of-the-art prediction methods. Case studies on three diseases further show FKRGAN's ability to identify candidate miRNA-disease associations.
Chunxu Xu, Ping Xuan, Mengxia Wang, Tiangang Zhang
IEEE Trans. Comput. Biol. Bioinform.3
2025 Joint Distribution Weighted Alignment for Multi-Source Domain Adaptation via Kernel Relative Entropy Estimation
abstract
The objective of Multi-Source Domain Adaptation (MSDA) is to train a neural network on labeled data from multiple joint source distributions (source domains) and unlabeled data from a joint target distribution (target domain), and use the trained network to estimate the target data labels. The challenge in this MSDA problem is that the multiple joint source distributions are relevant but distinct from the joint target distribution. To address this challenge, we propose a Joint Distribution Weighted Alignment (JDWA) approach to align a weighted joint source distribution to the joint target distribution under the relative entropy. Specifically, the weighted joint source distribution is defined as the weighted sum of the multiple joint source distributions, and is parameterized by the relevance weights. Since the relative entropy is unknown in practice, we propose a Kernel Relative Entropy Estimation (KREE) method to estimate it from data. Our KREE method first reformulates relative entropy as the negative of the minimal value of a functional, then exploits a function from the Reproducing Kernel Hilbert Space (RKHS) as the functional's input, and finally solves the resultant convex problem with a global optimal solution. We also incorporate entropy regularization to enhance the network's performance. Together, we minimize cross entropy, relative entropy, and entropy to learn both the relevance weights and the neural network. Experimental results on benchmark image classification datasets demonstrate that our JDWA approach performs better than the comparison methods. Pytorch code of our approach will be released upon the paper's publication.
Sentao Chen, Ping Xuan, Zhifeng Hao 0004
IEEE Trans. Multim.2
2024 TSEML: A task-specific embedding-based method for few-shot classification of cancer molecular subtypes
abstract
Molecular subtyping of cancer is recognized as a critical and challenging upstream task for personalized therapy. Existing deep learning methods have achieved significant performance in this domain when abundant data samples are available. However, the acquisition of densely labeled samples for cancer molecular subtypes remains a significant challenge for conventional data-intensive deep learning approaches. In this work, we focus on the few-shot molecular subtype prediction problem in heterogeneous and small cancer datasets, aiming to enhance precise diagnosis and personalized treatment. We first construct a new few-shot dataset for cancer molecular subtype classification and auxiliary cancer classification, named TCGA Few-Shot, from existing publicly available datasets. To effectively leverage the relevant knowledge from both tasks, we introduce a task-specific embedding-based meta-learning framework (TSEML). TSEML leverages the synergistic strengths of a model-agnostic meta-learning (MAML) approach and a prototypical network (ProtoNet) to capture diverse and fine-grained features. Comparative experiments conducted on the TCGA FewShot dataset demonstrate that our TSEML framework achieves superior performance in addressing the problem of few-shot molecular subtype classification.
Ran Su, Hui Cui 0002, Ping Xuan, Chengyan Fang, Xikang Feng, Qiangguo Jin
BIBM4
2024 MSKI-Net: Towards modality-specific knowledge interaction for glioma survival prediction
abstract
Gliomas hold a prominent position in neurooncology due to their high malignancy and poor survival rates. Accurately predicting the prognosis and survival risk of glioma patients is crucial for clinical treatment. Recent advances in survival prediction methods emphasize the importance of integrating complementary information from diverse modalities while neglecting the significant modality gap between pathological images and genomic data. To address this issue, we propose a modality-specific knowledge interaction network (MSKI-Net), which integrates whole slide images (WSI), RNA-Seq gene expression data, and copy number variation (CNV) data for glioma survival analysis. The MSKI-Net consists of a modality-specific feature enhancement (MSFE) module, a modality-interactive cross-attention (MICA) module, and a modality-specific knowledge-guided representation learning (MSKR) module. The three modules collaborate by complementing modality-specific features with modality-agnostic knowledge to improve the learning capability of MSKI-Net. Furthermore, we construct a dataset named TCGAmm, which combines WSI, RNA-Seq, and CNV data from The Cancer Genome Atlas (TCGA) to address the issue of data scarcity. Extensive experiments demonstrate that MSKI-Net achieves superior performance in predicting the survival risk of glioma cancer.
Ran Su, Hui Cui 0002, Ping Xuan, Xikang Feng, Leyi Wei, Qiangguo Jin
BIBM4
2024 Location Embedding Based Pairwise Distance Learning for Fine-Grained Diagnosis of Urinary Stones
Qiangguo Jin, Jiapeng Huang, Changming Sun, Hui Cui 0002, Ping Xuan, Ran Su, Leyi Wei, Yu-Jie Wu, Chia-An Wu, Henry Been-Lirn Duh, Yueh-Hsun Lu
MICCAI (11)5
2024 A Multi-information Dual-Layer Cross-Attention Model for Esophageal Fistula Prognosis
Jianqiao Zhang 0003, Hao Xiong 0001, Qiangguo Jin, Tian Feng 0001, Jiquan Ma, Ping Xuan, Peng Cheng 0002, Zhiyu Ning, ChangYang Li, Hui Cui 0002
MICCAI (5)6
2024 DrugMGR: a deep bioactive molecule binding method to identify compounds targeting proteins
abstract
MOTIVATION: Understanding the intermolecular interactions of ligand-target pairs is key to guiding the optimization of drug research on cancers, which can greatly mitigate overburden workloads for wet labs. Several improved computational methods have been introduced and exhibit promising performance for these identification tasks, but some pitfalls restrict their practical applications: (i) first, existing methods do not sufficiently consider how multigranular molecule representations influence interaction patterns between proteins and compounds; and (ii) second, existing methods seldom explicitly model the binding sites when an interaction occurs to enable better prediction and interpretation, which may lead to unexpected obstacles to biological researchers. RESULTS: To address these issues, we here present DrugMGR, a deep multigranular drug representation model capable of predicting binding affinities and regions for each ligand-target pair. We conduct consistent experiments on three benchmark datasets using existing methods and introduce a new specific dataset to better validate the prediction of binding sites. For practical application, target-specific compound identification tasks are also carried out to validate the capability of real-world compound screen. Moreover, the visualization of some practical interaction scenarios provides interpretable insights from the results of the predictions. The proposed DrugMGR achieves excellent overall performance in these datasets, exhibiting its advantages and merits against state-of-the-art methods. Thus, the downstream task of DrugMGR can be fine-tuned for identifying the potential compounds that target proteins for clinical treatment. AVAILABILITY AND IMPLEMENTATION: https://github.com/lixiaokun2020/DrugMGR.
Xiaokun Li, Qiang Yang 0015, Weihe Dong, Gongning Luo, Wei Wang 0169, Suyu Dong, Kuanquan Wang, Ping Xuan, Xianyu Zhang 0004, Xin Gao 0001
Bioinform.9
2024 Multi-scale topology and position feature learning and relationship-aware graph reasoning for prediction of drug-related microbes
abstract
MOTIVATION: The human microbiome may impact the effectiveness of drugs by modulating their activities and toxicities. Predicting candidate microbes for drugs can facilitate the exploration of the therapeutic effects of drugs. Most recent methods concentrate on constructing of the prediction models based on graph reasoning. They fail to sufficiently exploit the topology and position information, the heterogeneity of multiple types of nodes and connections, and the long-distance correlations among nodes in microbe-drug heterogeneous graph. RESULTS: We propose a new microbe-drug association prediction model, NGMDA, to encode the position and topological features of microbe (drug) nodes, and fuse the different types of features from neighbors and the whole heterogeneous graph. First, we formulate the position and topology features of microbe (drug) nodes by t-step random walks, and the features reveal the topological neighborhoods at multiple scales and the position of each node. Second, as the features of nodes are high-dimensional and sparse, we designed an embedding enhancement strategy based on supervised fully connected autoencoders to form the embeddings with representative features and the more discriminative node distributions. Third, we propose an adaptive neighbor feature fusion module, which fuses features of neighbors by the constructed position- and topology-sensitive heterogeneous graph neural networks. A novel self-attention mechanism is developed to estimate the importance of the position and topology of each neighbor to a target node. Finally, a heterogeneous graph feature fusion module is constructed to learn the long-distance correlations among the nodes in the whole heterogeneous graph by a relationship-aware graph transformer. Relationship-aware graph transformer contains the strategy for encoding the connection relationship types among the nodes, which is helpful for integrating the diverse semantics of these connections. The extensive comparison experimental results demonstrate NGMDA's superior performance over five state-of-the-art prediction methods. The ablation experiment shows the contributions of the multi-scale topology and position feature learning, the embedding enhancement strategy, the neighbor feature fusion, and the heterogeneous graph feature fusion. Case studies over three drugs further indicate that NGMDA has ability in discovering the potential drug-related microbes. AVAILABILITY AND IMPLEMENTATION: Source codes and Supplementary Material are available at https://github.com/pingxuan-hlju/NGMDA.
Ping Xuan, Hui Cui 0002, Shuai Wang 0038, Toshiya Nakaguchi, Tiangang Zhang
Bioinform.1
2024 Dynamic category-sensitive hypergraph inferring and homo-heterogeneous neighbor feature learning for drug-related microbe prediction
abstract
MOTIVATION: The microbes in human body play a crucial role in influencing the functions of drugs, as they can regulate the activities and toxicities of drugs. Most recent methods for predicting drug-microbe associations are based on graph learning. However, the relationships among multiple drugs and microbes are complex, diverse, and heterogeneous. Existing methods often fail to fully model the relationships. In addition, the attributes of drug-microbe pairs exhibit long-distance spatial correlations, which previous methods have not integrated effectively. RESULTS: We propose a new prediction method named DHDMP which is designed to encode the relationships among multiple drugs and microbes and integrate the attributes of various neighbor nodes along with the pairwise long-distance correlations. First, we construct a hypergraph with dynamic topology, where each hyperedge represents a specific relationship among multiple drug nodes and microbe nodes. Considering the heterogeneity of node attributes across different categories, we developed a node category-sensitive hypergraph convolution network to encode these diverse relationships. Second, we construct homogeneous graphs for drugs and microbes respectively, as well as drug-microbe heterogeneous graph, facilitating the integration of features from both homogeneous and heterogeneous neighbors of each target node. Third, we introduce a graph convolutional network with cross-graph feature propagation ability to transfer node features from homogeneous to heterogeneous graphs for enhanced neighbor feature representation learning. The propagation strategy aids in the deep fusion of features from both types of neighbors. Finally, we design spatial cross-attention to encode the attributes of drug-microbe pairs, revealing long-distance correlations among multiple pairwise attribute patches. The comprehensive comparison experiments showed our method outperformed state-of-the-art methods for drug-microbe association prediction. The ablation studies demonstrated the effectiveness of node category-sensitive hypergraph convolution network, graph convolutional network with cross-graph feature propagation, and spatial cross-attention. Case studies on three drugs further showed DHDMP's potential application in discovering the reliable candidate microbes for the interested drugs. AVAILABILITY AND IMPLEMENTATION: Source codes and supplementary materials are available at https://github.com/pingxuan-hlju/DHDMP.
Ping Xuan, Zelong Xu, Hui Cui 0002, Tiangang Zhang, Peiliang Wu
Bioinform.1
2024 Meta-Path Semantic and Global-Local Representation Learning Enhanced Graph Convolutional Model for Disease-Related miRNA Prediction
abstract
Dysregulation of miRNAs is closely related to the progression of various diseases, so identifying disease-related miRNAs is crucial. Most recently proposed methods are based on graph reasoning, while they did not completely exploit the topological structure composed of the higher-order neighbor nodes and the global and local features of miRNA and disease nodes. We proposed a prediction method, MDAP, to learn semantic features of miRNA and disease nodes based on various meta-paths, as well as node features from the entire heterogeneous network perspective, and node pair attributes. Firstly, for both the miRNA and disease nodes, node category-wise meta-paths were constructed to integrate the similarity and association connection relationships. Each target node has its specific neighbor nodes for each meta-path, and the neighbors of longer meta-paths constitute its higher-order neighbor topological structure. Secondly, we constructed a meta-path specific graph convolutional network module to integrate the features of higher-order neighbors and their topology, and then learned the semantic representations of nodes. Thirdly, for the entire miRNA-disease heterogeneous network, a global-aware graph convolutional autoencoder was built to learn the network-view feature representations of nodes. We also designed semantic-level and representation-level attentions to obtain informative semantic features and node representations. Finally, the strategy based on the parallel convolutional-deconvolutional neural networks were designed to enhance the local feature learning for a pair of miRNA and disease nodes. The experiment results showed that MDAP outperformed other state-of-the-art methods, and the ablation experiments demonstrated the effectiveness of MDAP's major innovations. MDAP's ability in discovering potential disease-related miRNAs was further analyzed by the case studies over three diseases.
Ping Xuan, Xiuju Wang, Hui Cui 0002, Xiangfeng Meng, Toshiya Nakaguchi, Tiangang Zhang
IEEE J. Biomed. Health Informatics1
2023 Shape-aware contrastive deep supervision for esophageal tumor segmentation from CT scans
abstract
Accurate tumor segmentation is crucial for esophageal cancer radiotherapy treatment planning. The low contrast among the esophagus, tumors, and surrounding tissues, and irregular tumor shapes limit the performance of automatic segmentation methods. In this paper, we aim to exploit the irregular shapes of tumors to facilitate accurate segmentation. We propose a simple and pluggable shape-aware contrastive deep supervision network (SCDSNet) with shape-aware regularization and voxel-to-voxel contrastive deep supervision. Specifically, the shape-aware regularization with an uncertainty minimization strategy encourages the precise predictions of an additional shape-aware head. The voxel-to-voxel contrastive deep supervision enhances the multi-scale shape-tumor contrast for better voxel-to-voxel prediction of shapes. The proposed method is simple and highly pluggable, which can easily be extended to other frameworks. Further, we establish a large in-house dataset on esophageal cancer to validate the effectiveness of our proposed method. The quantitative and qualitative experimental results demonstrate the effectiveness of SCDSNet on the esophageal cancer dataset.
Qiangguo Jin, Hui Cui 0002, Changming Sun, Jiapeng Huang, Ping Xuan, Yiyue Xu, Leilei Cao, Leyi Wei, Ran Su
BIBM5
2023 Multi-modality Contrastive Learning for Sarcopenia Screening from Hip X-rays and Clinical Information
Qiangguo Jin, Changjiang Zou, Hui Cui 0002, Changming Sun, Shu-Wei Huang, Yi-Jie Kuo, Ping Xuan, Leilei Cao, Ran Su, Leyi Wei, Henry Been-Lirn Duh, Yu-Pin Chen
MICCAI (6)7
2023 Semantic Meta-Path Enhanced Global and Local Topology Learning for lncRNA-Disease Association Prediction
abstract
Since abnormal expression of long non-coding RNAs (lncRNAs) is associated with various human diseases, identifying disease-related lncRNAs helps reveal the pathogenesis of diseases. Existing methods for lncRNA-disease association prediction mainly focus on multi-sourced data related to lncRNAs and diseases. The rich semantic information of meta-paths, composed of multiple kinds of connections between lncRNA and disease nodes, is neglected. We propose a new prediction method, MGLDA, to encode and integrate the semantics of multiple meta-paths, the global topology of heterogeneous graph, and pairwise attributes of lncRNA and disease nodes. First, a tri-layer heterogeneous graph is constructed to associate multi-sourced data across the lncRNA, disease, and miRNA nodes. Afterwards, we establish multiple meta-paths connecting the lncRNA and disease nodes to derive and denote various semantics. Each meta-path contains its specific semantics formulated by an embedding strategy, and each embedding covers local topology formed by the diverse semantic connections among the lncRNA, disease, and miRNA nodes. We construct multiple graph convolutional autoencoders (GCA) with topology-level attention to learn global and multiple local topologies from the tri-layer graph and each meta-path, respectively. The topology-level attention mechanism can learn the importance of various global and local topologies for adaptive pairwise topology fusion. Finally, a convolutional autoencoder learns the attribute representations of lncRNA-disease pairs, which integrates the learnt detailed and representative pairwise features. Experimental results show that MGLDA outperforms other state-of-the-art prediction methods in comparison and retrieves more real lncRNA-disease associations in the top-ranked candidates. The ablation study also demonstrates the important contributions of the local and global topology learning, and pairwise attribute learning. Case studies on three diseases further demonstrate MGLDA's ability to identify potential disease-related lncRNAs.
Ping Xuan, Hui Cui 0002, Linyun Zhan, Qiangguo Jin, Tiangang Zhang, Toshiya Nakaguchi
IEEE ACM Trans. Comput. Biol. Bioinform.1
2022 Prediction of drug-disease associations by integrating common topologies of heterogeneous networks and specific topologies of subnets
abstract
MOTIVATION: The development process of a new drug is time-consuming and costly. Thus, identifying new uses for approved drugs, named drug repositioning, is helpful for speeding up the drug development process and reducing development costs. Existing drug-related disease prediction methods mainly focus on single or multiple drug-disease heterogeneous networks. However, heterogeneous networks, and drug subnets and disease subnet contained in heterogeneous networks cover the common topology information between drug and disease nodes, the specific information between drug nodes and the specific information between disease nodes, respectively. RESULTS: We design a novel model, CTST, to extract and integrate common and specific topologies in multiple heterogeneous networks and subnets. Multiple heterogeneous networks composed of drug and disease nodes are established to integrate multiple kinds of similarities and associations among drug and disease nodes. These heterogeneous networks contain multiple drug subnets and a disease subnet. For multiple heterogeneous networks and subnets, we then define the common and specific representations of drug and disease nodes. The common representations of drug and disease nodes are encoded by a graph convolutional autoencoder with sharing parameters and they integrate the topological relationships of all nodes in heterogeneous networks. The specific representations of nodes are learned by specific graph convolutional autoencoders, respectively, and they fuse the topology and attributes of the nodes in each subnet. We then propose attention mechanisms at common representation level and specific representation level to learn more informative common and specific representations, respectively. Finally, an integration module with representation feature level attention is built to adaptively integrate these two representations for final association prediction. Extensive experimental results confirm the effectiveness of CTST. Comparison with six latest methods and case studies on five drugs further verify CTST has the ability to discover potential candidate diseases.
Hui Cui 0002, Tiangang Zhang, Nan Sheng, Ping Xuan
Briefings Bioinform.5
2022 ALDPI: adaptively learning importance of multi-scale topologies and multi-modality similarities for drug-protein interaction prediction
abstract
MOTIVATION: Effective computational methods to predict drug-protein interactions (DPIs) are vital for drug discovery in reducing the time and cost of drug development. Recent DPI prediction methods mainly exploit graph data composed of multiple kinds of connections among drugs and proteins. Each node in the graph usually has topological structures with multiple scales formed by its first-order neighbors and multi-order neighbors. However, most of the previous methods do not consider the topological structures of multi-order neighbors. In addition, deep integration of the multi-modality similarities of drugs and proteins is also a challenging task. RESULTS: We propose a model called ALDPI to adaptively learn the multi-scale topologies and multi-modality similarities with various significance levels. We first construct a drug-protein heterogeneous graph, which is composed of the interactions and the similarities with multiple modalities among drugs and proteins. An adaptive graph learning module is then designed to learn important kinds of connections in heterogeneous graph and generate new topology graphs. A module based on graph convolutional autoencoders is established to learn multiple representations, which imply the node attributes and multiple-scale topologies composed of one-order and multi-order neighbors, respectively. We also design an attention mechanism at neighbor topology level to distinguish the importance of these representations. Finally, since each similarity modality has its specific features, we construct a multi-layer convolutional neural network-based module to learn and fuse multi-modality features to obtain the attribute representation of each drug-protein node pair. Comprehensive experimental results show ALDPI's superior performance over six state-of-the-art methods. The results of recall rates of top-ranked candidates and case studies on five drugs further demonstrate the ability of ALDPI to discover potential drug-related protein candidates. CONTACT: [email protected].
Kaimiao Hu, Hui Cui 0002, Tiangang Zhang, Chang Sun 0002, Ping Xuan
Briefings Bioinform.5
2022 Multi-channel graph attention autoencoders for disease-related lncRNAs prediction
abstract
MOTIVATION: Predicting disease-related long non-coding RNAs (lncRNAs) can be used as the biomarkers for disease diagnosis and treatment. The development of effective computational prediction approaches to predict lncRNA-disease associations (LDAs) can provide insights into the pathogenesis of complex human diseases and reduce experimental costs. However, few of the existing methods use microRNA (miRNA) information and consider the complex relationship between inter-graph and intra-graph in complex-graph for assisting prediction. RESULTS: In this paper, the relationships between the same types of nodes and different types of nodes in complex-graph are introduced. We propose a multi-channel graph attention autoencoder model to predict LDAs, called MGATE. First, an lncRNA-miRNA-disease complex-graph is established based on the similarity and correlation among lncRNA, miRNA and diseases to integrate the complex association among them. Secondly, in order to fully extract the comprehensive information of the nodes, we use graph autoencoder networks to learn multiple representations from complex-graph, inter-graph and intra-graph. Thirdly, a graph-level attention mechanism integration module is adopted to adaptively merge the three representations, and a combined training strategy is performed to optimize the whole model to ensure the complementary and consistency among the multi-graph embedding representations. Finally, multiple classifiers are explored, and Random Forest is used to predict the association score between lncRNA and disease. Experimental results on the public dataset show that the area under receiver operating characteristic curve and area under precision-recall curve of MGATE are 0.964 and 0.413, respectively. MGATE performance significantly outperformed seven state-of-the-art methods. Furthermore, the case studies of three cancers further demonstrate the ability of MGATE to identify potential disease-correlated candidate lncRNAs. The source code and supplementary data are available at https://github.com/sheng-n/MGATE. CONTACT: [email protected], [email protected].
Nan Sheng, Lan Huang 0002, Yan Wang 0028, Ping Xuan, Yangkun Cao
Briefings Bioinform.5
2022 GVDTI: graph convolutional and variational autoencoders with attribute-level attention for drug-protein interaction prediction
abstract
MOTIVATION: Identifying proteins that interact with drugs plays an important role in the initial period of developing drugs, which helps to reduce the development cost and time. Recent methods for predicting drug-protein interactions mainly focus on exploiting various data about drugs and proteins. These methods failed to completely learn and integrate the attribute information of a pair of drug and protein nodes and their attribute distribution. RESULTS: We present a new prediction method, GVDTI, to encode multiple pairwise representations, including attention-enhanced topological representation, attribute representation and attribute distribution. First, a framework based on graph convolutional autoencoder is constructed to learn attention-enhanced topological embedding that integrates the topology structure of a drug-protein network for each drug and protein nodes. The topological embeddings of each drug and each protein are then combined and fused by multi-layer convolution neural networks to obtain the pairwise topological representation, which reveals the hidden topological relationships between drug and protein nodes. The proposed attribute-wise attention mechanism learns and adjusts the importance of individual attribute in each topological embedding of drug and protein nodes. Secondly, a tri-layer heterogeneous network composed of drug, protein and disease nodes is created to associate the similarities, interactions and associations across the heterogeneous nodes. The attribute distribution of the drug-protein node pair is encoded by a variational autoencoder. The pairwise attribute representation is learned via a multi-layer convolutional neural network to deeply integrate the attributes of drug and protein nodes. Finally, the three pairwise representations are fused by convolutional and fully connected neural networks for drug-protein interaction prediction. The experimental results show that GVDTI outperformed other seven state-of-the-art methods in comparison. The improved recall rates indicate that GVDTI retrieved more actual drug-protein interactions in the top ranked candidates than conventional methods. Case studies on five drugs further confirm GVDTI's ability in discovering the potential candidate drug-related proteins. CONTACT: [email protected] Supplementary information: Supplementary data are available at Briefings in Bioinformatics online.
Ping Xuan, Mengsi Fan, Hui Cui 0002, Tiangang Zhang, Toshiya Nakaguchi
Briefings Bioinform.1
2022 Fully connected autoencoder and convolutional neural network with attention-based method for inferring disease-related lncRNAs
abstract
Since abnormal expression of long noncoding RNAs (lncRNAs) is often closely related to various human diseases, identification of disease-associated lncRNAs is helpful for exploring the complex pathogenesis. Most of recent methods concentrate on exploiting multiple kinds of data related to lncRNAs and diseases for predicting candidate disease-related lncRNAs. These methods, however, failed to deeply integrate the topology information from the meta-paths that are composed of lncRNA, disease and microRNA (miRNA) nodes. We proposed a new method based on fully connected autoencoders and convolutional neural networks, called ACLDA, for inferring potential disease-related lncRNA candidates. A heterogeneous graph that consists of lncRNA, disease and miRNA nodes were firstly constructed to integrate similarities, associations and interactions among them. Fully connected autoencoder-based module was established to extract the low-dimensional features of lncRNA, disease and miRNA nodes in the heterogeneous graph. We designed the attention mechanisms at the node feature level and at the meta-path level to learn more informative features and meta-paths. A module based on convolutional neural networks was constructed to encode the local topologies of lncRNA and disease nodes from multiple meta-path perspectives. The comprehensive experimental results demonstrated ACLDA achieves superior performance than several state-of-the-art prediction methods. Case studies on breast, lung and colon cancers demonstrated that ACLDA is able to discover the potential disease-related lncRNAs.
Ping Xuan, Hui Cui 0002, Bochong Li, Tiangang Zhang
Briefings Bioinform.1
2022 Heterogeneous multi-scale neighbor topologies enhanced drug-disease association prediction
abstract
MOTIVATION: Identifying new uses of approved drugs is an effective way to reduce the time and cost of drug development. Recent computational approaches for predicting drug-disease associations have integrated multi-sourced data on drugs and diseases. However, neighboring topologies of various scales in multiple heterogeneous drug-disease networks have yet to be exploited and fully integrated. RESULTS: We propose a novel method for drug-disease association prediction, called MGPred, used to encode and learn multi-scale neighboring topologies of drug and disease nodes and pairwise attributes from heterogeneous networks. First, we constructed three heterogeneous networks based on multiple kinds of drug similarities. Each network comprises drug and disease nodes and edges created based on node-wise similarities and associations that reflect specific topological structures. We also propose an embedding mechanism to formulate topologies that cover different ranges of neighbors. To encode the embeddings and derive multi-scale neighboring topology representations of drug and disease nodes, we propose a module based on graph convolutional autoencoders with shared parameters for each heterogeneous network. We also propose scale-level attention to obtain an adaptive fusion of informative topological representations at different scales. Finally, a learning module based on a convolutional neural network with various receptive fields is proposed to learn multi-view attribute representations of a pair of drug and disease nodes. Comprehensive experiment results demonstrate that MGPred outperforms other state-of-the-art methods in comparison to drug-related disease prediction, and the recall rates for the top-ranked candidates and case studies on five drugs further demonstrate the ability of MGPred to retrieve potential drug-disease associations.
Ping Xuan, Xiangfeng Meng, Tiangang Zhang, Toshiya Nakaguchi
Briefings Bioinform.1
2022 Integration of pairwise neighbor topologies and miRNA family and cluster attributes for miRNA-disease association prediction
abstract
Identifying disease-related microRNAs (miRNAs) assists the understanding of disease pathogenesis. Existing research methods integrate multiple kinds of data related to miRNAs and diseases to infer candidate disease-related miRNAs. The attributes of miRNA nodes including their family and cluster belonging information, however, have not been deeply integrated. Besides, the learning of neighbor topology representation of a pair of miRNA and disease is a challenging issue. We present a disease-related miRNA prediction method by encoding and integrating multiple representations of miRNA and disease nodes learnt from the generative and adversarial perspective. We firstly construct a bilayer heterogeneous network of miRNA and disease nodes, and it contains multiple types of connections among these nodes, which reflect neighbor topology of miRNA-disease pairs, and the attributes of miRNA nodes, especially miRNA-related families and clusters. To learn enhanced pairwise neighbor topology, we propose a generative and adversarial model with a convolutional autoencoder-based generator to encode the low-dimensional topological representation of the miRNA-disease pair and multi-layer convolutional neural network-based discriminator to discriminate between the true and false neighbor topology embeddings. Besides, we design a novel feature category-level attention mechanism to learn the various importance of different features for final adaptive fusion and prediction. Comparison results with five miRNA-disease association methods demonstrated the superior performance of our model and technical contributions in terms of area under the receiver operating characteristic curve and area under the precision-recall curve. The results of recall rates confirmed that our model can find more actual miRNA-disease associations among top-ranked candidates. Case studies on three cancers further proved the ability to detect potential candidate miRNAs.
Ping Xuan, Hui Cui 0002, Tiangang Zhang, Toshiya Nakaguchi
Briefings Bioinform.1
2022 Learning global dependencies and multi-semantics within heterogeneous graph for predicting disease-related lncRNAs
abstract
MOTIVATION: Long noncoding RNAs (lncRNAs) play an important role in the occurrence and development of diseases. Predicting disease-related lncRNAs can help to understand the pathogenesis of diseases deeply. The existing methods mainly rely on multi-source data related to lncRNAs and diseases when predicting the associations between lncRNAs and diseases. There are interdependencies among node attributes in a heterogeneous graph composed of all lncRNAs, diseases and micro RNAs. The meta-paths composed of various connections between them also contain rich semantic information. However, the existing methods neglect to integrate attribute information of intermediate nodes in meta-paths. RESULTS: We propose a novel association prediction model, GSMV, to learn and deeply integrate the global dependencies, semantic information of meta-paths and node-pair multi-view features related to lncRNAs and diseases. We firstly formulate the global representations of the lncRNA and disease nodes by establishing a self-attention mechanism to capture and learn the global dependencies among node attributes. Second, starting from the lncRNA and disease nodes, respectively, multiple meta-pathways are established to reveal different semantic information. Considering that each meta-path contains specific semantics and has multiple meta-path instances which have different contributions to revealing meta-path semantics, we design a graph neural network based module which consists of a meta-path instance encoding strategy and two novel attention mechanisms. The proposed meta-path instance encoding strategy is used to learn the contextual connections between nodes within a meta-path instance. One of the two new attention mechanisms is at the meta-path instance level, which learns rich and informative meta-path instances. The other attention mechanism integrates various semantic information from multiple meta-paths to learn the semantic representation of lncRNA and disease nodes. Finally, a dilated convolution-based learning module with adjustable receptive fields is proposed to learn multi-view features of lncRNA-disease node pairs. The experimental results prove that our method outperforms seven state-of-the-art comparing methods for lncRNA-disease association prediction. Ablation experiments demonstrate the contributions of the proposed global representation learning, semantic information learning, pairwise multi-view feature learning and the meta-path instance encoding strategy. Case studies on three cancers further demonstrate our method's ability to discover potential disease-related lncRNA candidates. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Briefings in Bioinformatics online.
Ping Xuan, Shuai Wang 0038, Hui Cui 0002, Tiangang Zhang, Peiliang Wu
Briefings Bioinform.1
2022 Integrating specific and common topologies of heterogeneous graphs and pairwise attributes for drug-related side effect prediction
abstract
MOTIVATION: Computerized methods for drug-related side effect identification can help reduce costs and speed up drug development. Multisource data about drug and side effects are widely used to predict potential drug-related side effects. Heterogeneous graphs are commonly used to associate multisourced data of drugs and side effects which can reflect similarities of the drugs from different perspectives. Effective integration and formulation of diverse similarities, however, are challenging. In addition, the specific topology of each heterogeneous graph and the common topology of multiple graphs are neglected. RESULTS: We propose a drug-side effect association prediction model, GCRS, to encode and integrate specific topologies, common topologies and pairwise attributes of drugs and side effects. First, multiple drug-side effect heterogeneous graphs are constructed using various kinds of similarities and associations related to drugs and side effects. As each heterogeneous graph has its specific topology, we establish separate module based on graph convolutional autoencoder (GCA) to learn the particular topology representation of each drug node and each side effect node, respectively. Since multiple graphs reflect the complex relationships among the drug and side effect nodes and contain common topologies, we construct a module based on GCA with sharing parameters to learn the common topology representations of each node. Afterwards, we design an attention mechanism to obtain more informative topology representations at the representation level. Finally, multi-layer convolutional neural networks with attribute-level attention are constructed to deeply integrate the similarity and association attributes of a pair of drug-side effect nodes. Comprehensive experiments show that GCRS's prediction performance is superior to other comparing state-of-the-art methods for predicting drug-side effect associations. The recall rates in top-ranked candidates and case studies on five drugs further demonstrate GCRS's ability in discovering potential drug-related side effects. CONTACT: [email protected].
Ping Xuan, Tiangang Zhang, Toshiya Nakaguchi
Briefings Bioinform.1
2022 multi-type neighbors enhanced global topology and pairwise attribute learning for drug-protein interaction prediction
abstract
MOTIVATION: Accurate identification of proteins interacted with drugs helps reduce the time and cost of drug development. Most of previous methods focused on integrating multisource data about drugs and proteins for predicting drug-target interactions (DTIs). There are both similarity connection and interaction connection between two drugs, and these connections reflect their relationships from different perspectives. Similarly, two proteins have various connections from multiple perspectives. However, most of previous methods failed to deeply integrate these connections. In addition, multiple drug-protein heterogeneous networks can be constructed based on multiple kinds of connections. The diverse topological structures of these networks are still not exploited completely. RESULTS: We propose a novel model to extract and integrate multi-type neighbor topology information, diverse similarities and interactions related to drugs and proteins. Firstly, multiple drug-protein heterogeneous networks are constructed according to multiple kinds of connections among drugs and those among proteins. The multi-type neighbor node sequences of a drug node (or a protein node) are formed by random walks on each network and they reflect the hidden neighbor topological structure of the node. Secondly, a module based on graph neural network (GNN) is proposed to learn the multi-type neighbor topologies of each node. We propose attention mechanisms at neighbor node level and at neighbor type level to learn more informative neighbor nodes and neighbor types. A network-level attention is also designed to enhance the context dependency among multiple neighbor topologies of a pair of drug and protein nodes. Finally, the attribute embedding of the drug-protein pair is formulated by a proposed embedding strategy, and the embedding covers the similarities and interactions about the pair. A module based on three-dimensional convolutional neural networks (CNN) is constructed to deeply integrate pairwise attributes. Extensive experiments have been performed and the results indicate GCDTI outperforms several state-of-the-art prediction methods. The recall rate estimation over the top-ranked candidates and case studies on 5 drugs further demonstrate GCDTI's ability in discovering potential drug-protein interactions.
Ping Xuan, Kaimiao Hu, Toshiya Nakaguchi, Tiangang Zhang
Briefings Bioinform.1
2022 Learning multi-scale heterogenous network topologies and various pairwise attributes for drug-disease association prediction
abstract
MOTIVATION: Identifying new therapeutic effects for the approved drugs is beneficial for effectively reducing the drug development cost and time. Most of the recent computational methods concentrate on exploiting multiple kinds of information about drugs and disease to predict the candidate associations between drugs and diseases. However, the drug and disease nodes have neighboring topologies with multiple scales, and the previous methods did not fully exploit and deeply integrate these topologies. RESULTS: We present a prediction method, multi-scale topology learning for drug-disease (MTRD), to integrate and learn multi-scale neighboring topologies and the attributes of a pair of drug and disease nodes. First, for multiple kinds of drug similarities, multiple drug-disease heterogenous networks are constructed respectively to integrate the similarities and associations related to drugs and diseases. Moreover, each heterogenous network has its specific topology structure, which is helpful for learning the corresponding specific topology representation. We formulate the topology embeddings for each drug node and disease node by random walking on each heterogeneous network, and the embeddings cover the neighboring topologies with different scopes. Because the multi-scale topology embeddings have context relationships, we construct Bi-directional long short-term memory-based module to encode these embeddings and their relationships and learn the neighboring topology representation. We also design the attention mechanisms at feature level and at scale level to obtain the more informative pairwise features and topology embeddings. A module based on multi-layer convolutional networks is constructed to learn the representative attributes of the drug-disease node pair according to their related similarity and association information. Comprehensive experimental results indicate that MTRD achieves the superior performance than several state-of-the-art methods for predicting drug-disease associations. MTRD also retrieves more actual drug-disease associations in the top-ranked candidates of the prediction result. Case studies on five drugs further demonstrate MTRD's ability in discovering the potential candidate diseases for the interested drugs.
Hongda Zhang, Hui Cui 0002, Tiangang Zhang, Yangkun Cao, Ping Xuan
Briefings Bioinform.5
2022 Dynamic graph convolutional autoencoder with node-attribute-wise attention for kidney and tumor segmentation from CT volumes
Ping Xuan, Hui Cui 0002, Hongda Zhang, Tiangang Zhang, Toshiya Nakaguchi, Henry Been-Lirn Duh
Knowl. Based Syst.1
2022 Prediction of Drug-Related Diseases Through Integrating Pairwise Attributes and Neighbor Topological Structures
abstract
Identifying new disease indications for the approved drugs can help reduce the cost and time of drug development. Most of the recent methods focus on exploiting the various information related to drugs and diseases for predicting the candidate drug-disease associations. However, the previous methods failed to deeply integrate the neighborhood topological structure and the node attributes of an interested drug-disease node pair. We propose a new prediction method, ANPred, to learn and integrate pairwise attribute information and neighbor topology information from the similarities and associations related to drugs and diseases. First, a bi-layer heterogeneous network with intra-layer and inter-layer connections is established to combine the drug similarities, the disease similarities, and the drug-disease associations. Second, the embedding of a pair of drug and disease is constructed based on integrating multiple biological premises about drugs and diseases. The learning framework based on multi-layer convolutional neural networks is designed to learn the attribute representation of the pair of drug and disease nodes from its embedding. The sequences composed of neighbor nodes are formed based on random walk on the heterogeneous network. A framework based on fully-connected autoencoder and skip-gram module is constructed to learn the neighbor topological representations of nodes. The cross-validation results indicate the performance of ANPred is superior to several state-of-the-art methods. The case studies on 5 drugs further confirm the ability of ANPred in discovering the potential drug-disease association candidates.
Yingying Song, Hui Cui 0002, Tiangang Zhang, Tingxiao Yang, Xiaokun Li, Ping Xuan
IEEE ACM Trans. Comput. Biol. Bioinform.6
2022 Graph Convolutional Autoencoder and Generative Adversarial Network-Based Method for Predicting Drug-Target Interactions
abstract
The computational prediction of novel drug-target interactions (DTIs) may effectively speed up the process of drug repositioning and reduce its costs. Most previous methods integrated multiple kinds of connections about drugs and targets by constructing shallow prediction models. These methods failed to deeply learn the low-dimension feature vectors for drugs and targets and ignored the distribution of these feature vectors. We proposed a graph convolutional autoencoder and generative adversarial network (GAN)-based method, GANDTI, to predict DTIs. We constructed a drug-target heterogeneous network to integrate various connections related to drugs and targets, i.e., the similarities and interactions between drugs or between targets and the interactions between drugs and targets. A graph convolutional autoencoder was established to learn the network embeddings of the drug and target nodes in a low-dimensional feature space, and the autoencoder deeply integrated different kinds of connections within the network. A GAN was introduced to regularize the feature vectors of nodes into a Gaussian distribution. Severe class imbalance exists between known and unknown DTIs. Thus, we constructed a classifier based on an ensemble learning model, LightGBM, to estimate the interaction propensities of drugs and targets. This classifier completely exploited all unknown DTIs and counteracted the negative effect of class imbalance. The experimental results indicated that GANDTI outperforms several state-of-the-art methods for DTI prediction. Additionally, case studies of five drugs demonstrated the ability of GANDTI to discover the potential targets for drugs.
Chang Sun 0002, Ping Xuan, Tiangang Zhang
IEEE ACM Trans. Comput. Biol. Bioinform.2
2022 Inferring Drug-Target Interactions Based on Random Walk and Convolutional Neural Network
abstract
Computational strategies for identifying new drug-target interactions (DTIs) can guide the process of drug discovery, reduce the cost and time of drug development, and thus promote drug development. Most recently proposed methods predict DTIs via integration of heterogeneous data related to drugs and proteins. However, previous methods have failed to deeply integrate these heterogeneous data and learn deep feature representations of multiple original similarities and interactions related to drugs and proteins. We therefore constructed a heterogeneous network by integrating a variety of connection relationships about drugs and proteins, including drugs, proteins, and drug side effects, as well as their similarities, interactions, and associations. A DTI prediction method based on random walk and convolutional neural network was proposed and referred to as DTIPred. DTIPred not only takes advantage of various original features related to drugs and proteins, but also integrates the topological information of heterogeneous networks. The prediction model is composed of two sides and learns the deep feature representation of a drug-protein pair. On the left side, random walk with restart is applied to learn the topological vectors of drug and protein nodes. The topological representation is further learned by the constructed deep learning frame based on convolutional neural network. The right side of the model focuses on integrating multiple original similarities and interactions of drugs and proteins to learn the original representation of the drug-protein pair. The results of cross-validation experiments demonstrate that DTIPred achieves better prediction performance than several state-of-the-art methods. During the validation process, DTIPred can retrieve more actual drug-protein interactions within the top part of the predicted results, which may be more helpful to biologists. In addition, case studies on five drugs further demonstrate the ability of DTIPred to discover potential drug-protein interactions.
Xiaoqiang Xu, Ping Xuan, Tiangang Zhang, Bingxu Chen, Nan Sheng
IEEE ACM Trans. Comput. Biol. Bioinform.2
2022 Learning Multi-Scale Heterogeneous Representations and Global Topology for Drug-Target Interaction Prediction
abstract
Identification of interactions between drugs and target proteins plays a critical role not only in drug discovery but also in drug repositioning. Deep integration of inter-connections and intra-similarities between heterogeneous multi-source data about drugs and targets, however, is a challenging issue. We propose a drug-target interaction (DTI) prediction model by learning from drug and protein related multi-scale attributes and global topology formed by heterogeneous connections. A drug-protein-disease heterogeneous network (RPD-Net) is firstly constructed to associate diverse similarities, interactions and associations across nodes. Secondly, we propose a multi-scale pairwise deep representation learning module consisting of a new embedding strategy to integrate diverse inter-relations and intra-relations, and dilation convolutions for multi-scale deep representation extraction. A global topology learning module is proposed which is composed of strategy based on non-negative matrix factorization (NMF) to extract topology from RPD-Net, and a new relational-level attention mechanism for discriminative topology embedding. Experimental results using public dataset demonstrate improved performance over state-of-the-art methods and contributions of our major innovations. Evaluation results by top k recall rates and case studies on five drugs further show the effectiveness of our method in retrieving potential target candidates for drugs.
Ping Xuan, Kaimiao Hu, Hui Cui 0002, Tiangang Zhang, Toshiya Nakaguchi
IEEE J. Biomed. Health Informatics1
2022 Graph Triple-Attention Network for Disease-Related LncRNA Prediction
abstract
Abnormal expressions of long non-coding RNAs (lncRNAs) are associated with various human diseases. Identifying disease-related lncRNAs can help clarify complex disease pathogeneses. The latest methods for lncRNA-disease association prediction rely on diverse data about lncRNAs and diseases. These methods, however, cannot adequately integrate the neighbour topological information of lncRNA and disease nodes. Moreover, more intrinsic features of lncRNA-disease node pairs can be explored to better predict their latent associations. We developed a novel method, named GTAN, to predict the association propensities between lncRNAs and diseases. GTAN integrates various information about lncRNAs and diseases, and exploits neighbour topology and attribute representations of a pair of lncRNA-disease nodes. We adopted in GTAN a graph neural network architecture with three attention mechanisms and multi-layer convolutional neural networks. First, a neighbour-level self-attention mechanism is constructed to learn the importance of each neighbour for an interested lncRNA or disease node. Second, topology-level attention is proposed to enhance contextual dependencies among multiple local topology representations. An attention-enhanced graph neural network framework is then established to learn a topology representation of top-ranked neighbours. GTAN also has attribute-level attention to distinguish various contributions of attributes of the lncRNA-disease pair. Finally, attribute representation is learned by multi-layer CNN to integrate detailed features and representative features of the pair. Extensive experimental results demonstrated that GTAN outperformed state-of-the-art methods. The ablation studies confirmed the important contributions of three attention mechanisms. Case studies on three cancers further showed GTAN's ability in discovering potential lncRNA candidates related to diseases.
Ping Xuan, Liyun Zhan, Hui Cui 0002, Tiangang Zhang, Toshiya Nakaguchi, Weixiong Zhang
IEEE J. Biomed. Health Informatics1
2021 Co-graph Attention Reasoning Based Imaging and Clinical Features Integration for Lymph Node Metastasis Prediction
Hui Cui 0002, Ping Xuan, Qiangguo Jin, Mingjun Ding, Butuo Li, Bing Zou, Yiyue Xu, Bingjie Fan, Wanlong Li, Jinming Yu, Henry Been-Lirn Duh
MICCAI (5)2
2021 Predicting Esophageal Fistula Risks Using a Multimodal Self-attention Network
Yulu Guan, Hui Cui 0002, Yiyue Xu, Qiangguo Jin, Tian Feng 0001, Huawei Tu, Ping Xuan, Wanlong Li, Henry Been-Lirn Duh
MICCAI (5)7
2021 Attentional multi-level representation encoding based on convolutional and variance autoencoders for lncRNA-disease association prediction
abstract
As the abnormalities of long non-coding RNAs (lncRNAs) are closely related to various human diseases, identifying disease-related lncRNAs is important for understanding the pathogenesis of complex diseases. Most of current data-driven methods for disease-related lncRNA candidate prediction are based on diseases and lncRNAs. Those methods, however, fail to consider the deeply embedded node attributes of lncRNA-disease pairs, which contain multiple relations and representations across lncRNAs, diseases and miRNAs. Moreover, the low-dimensional feature distribution at the pairwise level has not been taken into account. We propose a prediction model, VADLP, to extract, encode and adaptively integrate multi-level representations. Firstly, a triple-layer heterogeneous graph is constructed with weighted inter-layer and intra-layer edges to integrate the similarities and correlations among lncRNAs, diseases and miRNAs. We then define three representations including node attributes, pairwise topology and feature distribution. Node attributes are derived from the graph by an embedding strategy to represent the lncRNA-disease associations, which are inferred via their common lncRNAs, diseases and miRNAs. Pairwise topology is formulated by random walk algorithm and encoded by a convolutional autoencoder to represent the hidden topological structural relations between a pair of lncRNA and disease. The new feature distribution is modeled by a variance autoencoder to reveal the underlying lncRNA-disease relationship. Finally, an attentional representation-level integration module is constructed to adaptively fuse the three representations for lncRNA-disease association prediction. The proposed model is tested over a public dataset with a comprehensive list of evaluations. Our model outperforms six state-of-the-art lncRNA-disease prediction models with statistical significance. The ablation study showed the important contributions of three representations. In particular, the improved recall rates under different top $k$ values demonstrate that our model is powerful in discovering true disease-related lncRNAs in the top-ranked candidates. Case studies of three cancers further proved the capacity of our model to discover potential disease-related lncRNAs.
Nan Sheng, Hui Cui 0002, Tiangang Zhang, Ping Xuan
Briefings Bioinform.4
2021 Integrating multi-scale neighbouring topologies and cross-modal similarities for drug-protein interaction prediction
abstract
MOTIVATION: Identifying the proteins that interact with drugs can reduce the cost and time of drug development. Existing computerized methods focus on integrating drug-related and protein-related data from multiple sources to predict candidate drug-target interactions (DTIs). However, multi-scale neighboring node sequences and various kinds of drug and protein similarities are neither fully explored nor considered in decision making. RESULTS: We propose a drug-target interaction prediction method, DTIP, to encode and integrate multi-scale neighbouring topologies, multiple kinds of similarities, associations, interactions related to drugs and proteins. We firstly construct a three-layer heterogeneous network to represent interactions and associations across drug, protein, and disease nodes. Then a learning framework based on fully-connected autoencoder is proposed to learn the nodes' low-dimensional feature representations within the heterogeneous network. Secondly, multi-scale neighbouring sequences of drug and protein nodes are formulated by random walks. A module based on bidirectional gated recurrent unit is designed to learn the neighbouring sequential information and integrate the low-dimensional features of nodes. Finally, we propose attention mechanisms at feature level, neighbouring topological level and similarity level to learn more informative features, topologies and similarities. The prediction results are obtained by integrating neighbouring topologies, similarities and feature attributes using a multiple layer CNN. Comprehensive experimental results over public dataset demonstrated the effectiveness of our innovative features and modules. Comparison with other state-of-the-art methods and case studies of five drugs further validated DTIP's ability in discovering the potential candidate drug-related proteins.
Ping Xuan, Hui Cui 0002, Tiangang Zhang, Maozu Guo 0001, Toshiya Nakaguchi
Briefings Bioinform.1
2021 Prediction of Drug-Target Interactions Based on Network Representation Learning and Ensemble Learning
abstract
Identifying interactions between drugs and target proteins is a critical step in the drug development process, as it helps identify new targets for drugs and accelerate drug development. The number of known drug-protein interactions (positive samples) is much lower than that of the unknown ones (negative samples), which forms a class imbalance. Most previous methods only utilised part of the negative samples to train the prediction model, so most of the information on negative samples was neglected. Therefore, a new method must be developed to predict candidate drug-related proteins and fully utilise negative samples to improve prediction performance. We present a method based on non-negative matrix factorisation and gradient boosting decision tree (GBDT), named NGDTP, to identify the candidate drug-protein interactions. NGDTP integrates multiple kinds of protein similarities, drugs-proteins interactions, and multiple kinds of drugs similarities at different levels, including target proteins of drugs, drug-related diseases, and side effects of drugs. We propose a network representation learning method based on matrix factorisation to learn low-dimensional vector representations of drug and protein nodes. On the basis of these low-dimensional node representations, a GBDT-based prediction model was constructed and it obtains the association scores through establishing multiple decision trees for a drug-protein pairs. NGDTP is an ensemble learning model that fully utilises all the negative samples to effectively alleviate the problem of class imbalance. NGDTP achieves superior prediction performance when it is compared with several state-of-the-art methods. The experimental results indicate that NGDTP also retrieves more actual drug-protein interactions in the top part of prediction result, which drew significant attention from the biologists. In addition, case studies on 10 drugs further confirmed the ability of the NGDTP to identify potential candidate proteins for drugs.
Ping Xuan, Bingxu Chen, Tiangang Zhang
IEEE ACM Trans. Comput. Biol. Bioinform.1
2021 Graph Convolutional Autoencoder and Fully-Connected Autoencoder with Attention Mechanism Based Method for Predicting Drug-Disease Associations
abstract
Predicting novel uses for approved drugs helps in reducing the costs of drug development and facilitates the development process. Most of previous methods focused on the multi-source data related to drugs and diseases to predict the candidate associations between drugs and diseases. There are multiple kinds of similarities between drugs, and these similarities reflect how similar two drugs are from the different views, whereas most of the previous methods failed to deeply integrate these similarities. In addition, the topology structures of the multiple drug-disease heterogeneous networks constructed by using the different kinds of drug similarities are not fully exploited. We therefore propose GFPred, a method based on a graph convolutional autoencoder and a fully-connected autoencoder with an attention mechanism, to predict drug-related diseases. GFPred integrates drug-disease associations, disease similarities, three kinds of drug similarities and attributes of the drug nodes. Three drug-disease heterogeneous networks are constructed based on the different kinds of drug similarities. We construct a graph convolutional autoencoder module, and integrate the attributes of the drug and disease nodes in each network to learn the topology representations of each drug node and disease node. As the different kinds of drug attributes contribute differently to the prediction of drug-disease associations, we construct an attribute-level attention mechanism. A fully-connected autoencoder module is established to learn the attribute representations of the drug and disease nodes. Finally, the original features of the drug-disease node pairs are also important auxiliary information for their association prediction. A combined strategy based on a convolutional neural network is proposed to fully integrate the topology representations, the attribute representations, and the original features of the drug-disease pairs. The ablation studies showed the contributions of data related to three types of drug attributes. Comparison with other methods confirmed that GFPred achieved better performance than several state-of-the-art prediction methods. In particular, case studies confirmed that GFPred is able to retrieve more actual drug-disease associations in the top k part of the prediction results. It is helpful for biologists to discover real associations by wet-lab experiments.
Ping Xuan, Nan Sheng, Tiangang Zhang, Toshiya Nakaguchi
IEEE J. Biomed. Health Informatics1
2020 Assistant diagnosis with Chinese electronic medical records based on CNN and BiLSTM with phrase-level and word-level attentions
abstract
BACKGROUND: Inferring diseases related to the patient's electronic medical records (EMRs) is of great significance for assisting doctor diagnosis. Several recent prediction methods have shown that deep learning-based methods can learn the deep and complex information contained in EMRs. However, they do not consider the discriminative contributions of different phrases and words. Moreover, local information and context information of EMRs should be deeply integrated. RESULTS: A new method based on the fusion of a convolutional neural network (CNN) and bidirectional long short-term memory (BiLSTM) with attention mechanisms is proposed for predicting a disease related to a given EMR, and it is referred to as FCNBLA. FCNBLA deeply integrates local information, context information of the word sequence and more informative phrases and words. A novel framework based on deep learning is developed to learn the local representation, the context representation and the combination representation. The left side of the framework is constructed based on CNN to learn the local representation of adjacent words. The right side of the framework based on BiLSTM focuses on learning the context representation of the word sequence. Not all phrases and words contribute equally to the representation of an EMR meaning. Therefore, we establish the attention mechanisms at the phrase level and word level, and the middle module of the framework learns the combination representation of the enhanced phrases and words. The macro average f-score and accuracy of FCNBLA achieved 91.29 and 92.78%, respectively. CONCLUSION: The experimental results indicate that FCNBLA yields superior performance compared with several state-of-the-art methods. The attention mechanisms and combination representations are also confirmed to be helpful for improving FCNBLA's prediction performance. Our method is helpful for assisting doctors in diagnosing diseases in patients.
Ping Xuan, Tiangang Zhang
BMC Bioinform.2
2020 Inferring Disease-Associated microRNAs in Heterogeneous Networks with Node Attributes
abstract
Identification of disease-associated microRNAs (disease miRNAs) is an essential step towards discovering causal miRNAs and understanding disease pathogenesis. Two sources of information can be exploited for predicting disease miRNAs: one includes the connections between miRNAs, between diseases, and between miRNAs and diseases, and the other has the attributes of miRNA nodes. The former contains information of miRNA similarities, disease similarities, and miRNA-disease associations. The latter includes the information of the families and clusters that miRNAs belong to. Similar diseases are usually associated with miRNAs that have similar functions and common attributes. However, most of the existing methods for disease miRNA prediction focus only on the connections of miRNAs and diseases. It remains challenging to adequately integrate the connections and miRNA node attributes to identify more reliable candidate disease miRNAs. We propose a non-negative matrix factorization based method, FamCluRank, for predicting disease miRNAs in heterogeneous networks with node attributes. One of the novelties of FamCluRank is to fully utilize these two oversighted characteristics of miRNAs and focuses particularly on a deep integration of miRNA families and cluster attributes. In particular, the integration was achieved by three different means. We first constructed a miRNA-disease heterogeneous network with node attributes where the miRNA nodes have their family and cluster attributes. Second, miRNAs sharing more common families and clusters are more likely to be associated with the diseases that are also related to these families and clusters. On the basis of the biological premise, we constructed a novel prediction model of FamCluRank to deeply integrate the family and cluster attributes of miRNAs. Third, two similar diseases tend to be associated with more common miRNA families and clusters, and vice versa. Hence, FamCluRank's prediction model is constructed by concerning not only the possible associations between miRNAs and diseases but also the possible disease-family and disease-cluster associations. Comparison with the state-of-the-art methods showed FamCluRank's superior performance not only on the well-characterized diseases but also on the new ones. Case studies on colorectal neoplasms, pancreatic neoplasms, lung neoplasms, and 32 new diseases demonstrated its ability for discovering potential disease miRNAs. FamCluRank is a potent prioritization tool for screening the reliable candidates for subsequent studies concerning their involvement in the pathogenesis of diseases. The web service of FamCluRank, the candidate disease miRNAs for 329 diseases, and the dataset used to develop FamCluRank are available at http://www.famclurank.top.
Ping Xuan, Tonghui Shen, Xiao Wang 0017, Tiangang Zhang, Weixiong Zhang
IEEE ACM Trans. Comput. Biol. Bioinform.1
2019 PRME-GTS: A New Successive POI Recommendation Model with Temporal and Social Influences
Rubai Mao, Zitu Liu, Yong Liu 0029, Xingfeng Lv, Ping Xuan
ADMA6
2019 Drug repositioning through integration of prior knowledge and projections of drugs and diseases
abstract
MOTIVATION: Identifying and developing novel therapeutic effects for existing drugs contributes to reduction of drug development costs. Most of the previous methods focus on integration of the heterogeneous data of drugs and diseases from multiple sources for predicting the candidate drug-disease associations. However, they fail to take the prior knowledge of drugs and diseases and their sparse characteristic into account. It is essential to develop a method that exploits the more useful information to predict the reliable candidate associations. RESULTS: We present a method based on non-negative matrix factorization, DisDrugPred, to predict the drug-related candidate disease indications. A new type of drug similarity is firstly calculated based on their associated diseases. DisDrugPred completely integrates two types of disease similarities, the associations between drugs and diseases, and the various similarities between drugs from different levels including the chemical structures of drugs, the target proteins of drugs, the diseases associated with drugs and the side effects of drugs. The prior knowledge of drugs and diseases and the sparse characteristic of drug-disease associations provide a deep biological perspective for capturing the relationships between drugs and diseases. Simultaneously, the possibility that a drug is associated with a disease is also dependant on their projections in the low-dimension feature space. Therefore, DisDrugPred deeply integrates the diverse prior knowledge, the sparse characteristic of associations and the projections of drugs and diseases. DisDrugPred achieves superior prediction performance than several state-of-the-art methods for drug-disease association prediction. During the validation process, DisDrugPred also can retrieve more actual drug-disease associations in the top part of prediction result which often attracts more attention from the biologists. Moreover, case studies on five drugs further confirm DisDrugPred's ability to discover potential candidate disease indications for drugs. AVAILABILITY AND IMPLEMENTATION: The fourth type of drug similarity and the predicted candidates for all the drugs are available at https://github.com/pingxuan-hlju/DisDrugPred. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ping Xuan, Yangkun Cao, Tiangang Zhang, Xiao Wang 0017, Shuxiang Pan, Tonghui Shen
Bioinform.1
2018 A non-negative matrix factorization based method for predicting disease-associated miRNAs in miRNA-disease bilayer network
abstract
MOTIVATION: Identification of disease-associated miRNAs (disease miRNAs) is critical for understanding disease etiology and pathogenesis. Since miRNAs exert their functions by regulating the expression of their target mRNAs, several methods based on the target genes were proposed to predict disease miRNA candidates. They achieved only limited success as they all suffered from the high false-positive rate of target prediction results. Alternatively, other prediction methods were based on the observation that miRNAs with similar functions tend to be associated with similar diseases and vice versa. The methods exploited the information about miRNAs and diseases, including the functional similarities between miRNAs, the similarities between diseases, and the associations between miRNAs and diseases. However, how to integrate the multiple kinds of information completely and consider the biological characteristic of disease miRNAs is a challenging problem. RESULTS: We constructed a bilayer network to represent the complex relationships among miRNAs, among diseases and between miRNAs and diseases. We proposed a non-negative matrix factorization based method to rank, so as to predict, the disease miRNA candidates. The method integrated the miRNA functional similarity, the disease similarity and the miRNA-disease associations seamlessly, which exploited the complex relationships within the bilayer network and the consensus relationship between multiple kinds of information. Considering the correlation between the candidates related to various diseases, it predicted their respective candidates for all the diseases simultaneously. In addition, the sparseness characteristic of disease miRNAs was introduced to generate more reliable prediction model that excludes those noisy candidates. The results on 15 common diseases showed a superior performance of the new method for not only well-characterized diseases but also new ones. A detailed case study on breast neoplasms, colorectal neoplasms, lung neoplasms and 32 other diseases demonstrated the ability of the method for discovering potential disease miRNAs. AVAILABILITY AND IMPLEMENTATION: The web service for the new method and the list of predicted candidates for all the diseases are available at http://www.bioinfolab.top. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yingli Zhong, Ping Xuan, Xiao Wang 0017, Tiangang Zhang, Jianzhong Li 0001, Yong Liu 0029, Weixiong Zhang
Bioinform.2
2015 Prediction of potential disease-associated microRNAs based on random walk
abstract
MOTIVATION: Identifying microRNAs associated with diseases (disease miRNAs) is helpful for exploring the pathogenesis of diseases. Because miRNAs fulfill function via the regulation of their target genes and because the current number of experimentally validated targets is insufficient, some existing methods have inferred potential disease miRNAs based on the predicted targets. It is difficult for these methods to achieve excellent performance due to the high false-positive and false-negative rates for the target prediction results. Alternatively, several methods have constructed a network composed of miRNAs based on their associated diseases and have exploited the information within the network to predict the disease miRNAs. However, these methods have failed to take into account the prior information regarding the network nodes and the respective local topological structures of the different categories of nodes. Therefore, it is essential to develop a method that exploits the more useful information to predict reliable disease miRNA candidates. RESULTS: miRNAs with similar functions are normally associated with similar diseases and vice versa. Therefore, the functional similarity between a pair of miRNAs is calculated based on their associated diseases to construct a miRNA network. We present a new prediction method based on random walk on the network. For the diseases with some known related miRNAs, the network nodes are divided into labeled nodes and unlabeled nodes, and the transition matrices are established for the two categories of nodes. Furthermore, different categories of nodes have different transition weights. In this way, the prior information of nodes can be completely exploited. Simultaneously, the various ranges of topologies around the different categories of nodes are integrated. In addition, how far the walker can go away from the labeled nodes is controlled by restarting the walking. This is helpful for relieving the negative effect of noisy data. For the diseases without any known related miRNAs, we extend the walking on a miRNA-disease bilayer network. During the prediction process, the similarity between diseases, the similarity between miRNAs, the known miRNA-disease associations and the topology information of the bilayer network are exploited. Moreover, the importance of information from different layers of network is considered. Our method achieves superior performance for 18 human diseases with AUC values ranging from 0.786 to 0.945. Moreover, case studies on breast neoplasms, lung neoplasms, prostatic neoplasms and 32 diseases further confirm the ability of our method to discover potential disease miRNAs. AVAILABILITY AND IMPLEMENTATION: A web service for the prediction and analysis of disease miRNAs is available at http://bioinfolab.stx.hk/midp/.
Ping Xuan, Yahong Guo, Jin Li 0024, Xia Li 0004, Yingli Zhong, Zhaogong Zhang
Bioinform.1
2013 Measuring gene functional similarity based on group-wise comparison of GO terms
abstract
MOTIVATION: Compared with sequence and structure similarity, functional similarity is more informative for understanding the biological roles and functions of genes. Many important applications in computational molecular biology require functional similarity, such as gene clustering, protein function prediction, protein interaction evaluation and disease gene prioritization. Gene Ontology (GO) is now widely used as the basis for measuring gene functional similarity. Some existing methods combined semantic similarity scores of single term pairs to estimate gene functional similarity, whereas others compared terms in groups to measure it. However, these methods may make error-prone judgments about gene functional similarity. It remains a challenge that measuring gene functional similarity reliably. RESULT: We propose a novel method called SORA to measure gene functional similarity in GO context. First of all, SORA computes the information content (IC) of a term making use of semantic specificity and coverage. Second, SORA measures the IC of a term set by means of combining inherited and extended IC of the terms based on the structure of GO. Finally, SORA estimates gene functional similarity using the IC overlap ratio of term sets. SORA is evaluated against five state-of-the-art methods in the file on the public platform for collaborative evaluation of GO-based semantic similarity measure. The carefully comparisons show SORA is superior to other methods in general. Further analysis suggests that it primarily benefits from the structure of GO, which implies expressive information about gene function. SORA offers an effective and reliable way to compare gene function. AVAILABILITY: The web service of SORA is freely available at http://nclab.hit.edu.cn/SORA/
Zhixia Teng, Maozu Guo 0001, Qiguo Dai, Chunyu Wang 0002, Ping Xuan
Bioinform.6
2011 PlantMiRNAPred: efficient classification of real and pseudo plant pre-miRNAs
abstract
MOTIVATION: MicroRNAs (miRNAs) are a set of short (21-24 nt) non-coding RNAs that play significant roles as post-transcriptional regulators in animals and plants. While some existing methods use comparative genomic approaches to identify plant precursor miRNAs (pre-miRNAs), others are based on the complementarity characteristics between miRNAs and their target mRNAs sequences. However, they can only identify the homologous miRNAs or the limited complementary miRNAs. Furthermore, since the plant pre-miRNAs are quite different from the animal pre-miRNAs, all the ab initio methods for animals cannot be applied to plants. Therefore, it is essential to develop a method based on machine learning to classify real plant pre-miRNAs and pseudo genome hairpins. RESULTS: A novel classification method based on support vector machine (SVM) is proposed specifically for predicting plant pre-miRNAs. To make efficient prediction, we extract the pseudo hairpin sequences from the protein coding sequences of Arabidopsis thaliana and Glycine max, respectively. These pseudo pre-miRNAs are extracted in this study for the first time. A set of informative features are selected to improve the classification accuracy. The training samples are selected according to their distributions in the high-dimensional sample space. Our classifier PlantMiRNAPred achieves >90% accuracy on the plant datasets from eight plant species, including A.thaliana, Oryza sativa, Populus trichocarpa, Physcomitrella patens, Medicago truncatula, Sorghum bicolor, Zea mays and G.max. The superior performance of the proposed classifier can be attributed to the extracted plant pseudo pre-miRNAs, the selected training dataset and the carefully selected features. The ability of PlantMiRNAPred to discern real and pseudo pre-miRNAs provides a viable method for discovering new non-homologous plant pre-miRNAs.
Ping Xuan, Maozu Guo 0001, Yangchao Huang, Yufei Huang 0001
Bioinform.1
2010 Two-stage clustering based effective sample selection for classification of pre-miRNAs
abstract
To solve the class imbalance problem in classification of pre-miRNAs with ab initio method, a novel sample selection method is proposed according to the characteristics of pre-miRNAs. Real/pseudo pre-miRNAs are clustered based on their stem similarity and their distribution in high dimensional sample space respectively. The training samples are selected according to the sample density of each cluster. Experimental results are validated by the cross validation and other testing datasets composed of human real/pseudo pre-miRNAs. When compared with the previous study, microPred, our classifier miRNAPred is nearly 12% greater in total accuracy. Our sample selection algorithm is useful to construct more efficient classifier for classification of real pre-miRNAs and pseudo hairpin sequences.
Ping Xuan, Maozu Guo 0001, Jun Wang 0035, Yingpeng Han
BIBM1
2004 Evolution of the GPGP/TÆMS Domain-Independent Coordination Framework
Victor R. Lesser, Keith S. Decker, Thomas Wagner 0001, Norman Carver, Alan Garvey, Bryan Horling, Daniel E. Neiman, Rodion M. Podorozhny, M. V. Nagendra Prasad, Anita Raja, Régis Vincent, Ping Xuan, Xiaoqin Zhang 0001
Auton. Agents Multi Agent Syst.12