Yangkun Cao

dblp:251/2693 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
7since 2021 · last 2024
0000-0001-7240-5486ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Estimating Individual Causal Treatment Effect by Variable Decomposition
abstract
Estimating individual-level causal effects is crucial for decision-making in various domains, such as personalized healthcare, social marketing, and public policy. Addressing confounding bias is a critical step in accurately estimating the causal effects of treatments on outcomes. However, many current causal inference approaches consider all observed variables as confounders without distinguishing them from colliders or indirect (two-order) colliders. This may lead to M-bias when improperly eliminating confounding bias. In this study, we propose a new framework to accurately estimate individual-level treatment effects by considering a causal structure that includes both confounding variables and indirect colliders. Specifically, we first perform a sample reweighting to approximately eliminate confounding bias. Then, we restore the covariate’ potential latent parents and extract the modules solely related to the outcome. Finally, we take both these modules with the treatment variables to infer counterfactuals for causal inference. To validate the effectiveness of our proposed approach, we conduct extensive experiments on synthetic and commonly used semi-synthetic benchmark datasets. The experimental results demonstrate that our method outperforms current state-of-the-art methods.
Hongyang Jiang 0002, Yonghe Zhao, Yangkun Cao, Huiyan Sun, Yi Chang 0001
IJCNN4
2023 MORGAT: A Model Based Knowledge-Informed Multi-omics Integration and Robust Graph Attention Network for Molecular Subtyping of Cancer
Haobo Shi, Yangkun Cao
ICIC (3)5
2023 Multi-task prediction-based graph contrastive learning for inferring the relationship among lncRNAs, miRNAs and diseases
abstract
MOTIVATION: Identifying the relationships among long non-coding RNAs (lncRNAs), microRNAs (miRNAs) and diseases is highly valuable for diagnosing, preventing, treating and prognosing diseases. The development of effective computational prediction methods can reduce experimental costs. While numerous methods have been proposed, they often to treat the prediction of lncRNA-disease associations (LDAs), miRNA-disease associations (MDAs) and lncRNA-miRNA interactions (LMIs) as separate task. Models capable of predicting all three relationships simultaneously remain relatively scarce. Our aim is to perform multi-task predictions, which not only construct a unified framework, but also facilitate mutual complementarity of information among lncRNAs, miRNAs and diseases. RESULTS: In this work, we propose a novel unsupervised embedding method called graph contrastive learning for multi-task prediction (GCLMTP). Our approach aims to predict LDAs, MDAs and LMIs by simultaneously extracting embedding representations of lncRNAs, miRNAs and diseases. To achieve this, we first construct a triple-layer lncRNA-miRNA-disease heterogeneous graph (LMDHG) that integrates the complex relationships between these entities based on their similarities and correlations. Next, we employ an unsupervised embedding model based on graph contrastive learning to extract potential topological feature of lncRNAs, miRNAs and diseases from the LMDHG. The graph contrastive learning leverages graph convolutional network architectures to maximize the mutual information between patch representations and corresponding high-level summaries of the LMDHG. Subsequently, for the three prediction tasks, multiple classifiers are explored to predict LDA, MDA and LMI scores. Comprehensive experiments are conducted on two datasets (from older and newer versions of the database, respectively). The results show that GCLMTP outperforms other state-of-the-art methods for the disease-related lncRNA and miRNA prediction tasks. Additionally, case studies on two datasets further demonstrate the ability of GCLMTP to accurately discover new associations. To ensure reproducibility of this work, we have made the datasets and source code publicly available at https://github.com/sheng-n/GCLMTP.
Nan Sheng, Yan Wang 0028, Lan Huang 0002, Yangkun Cao, Xuping Xie
Briefings Bioinform.5
2023 A Survey of Computational Methods and Databases for lncRNA-MiRNA Interaction Prediction
abstract
Long non-coding RNAs (lncRNAs) and microRNAs (miRNAs) are two prevalent non-coding RNAs in current research. They play critical regulatory roles in the life processes of animals and plants. Studies have shown that lncRNAs can interact with miRNAs to participate in post-transcriptional regulatory processes, mainly involved in regulating cancer development, metastatic progression, and drug resistance. Additionally, these interactions have significant effects on plant growth, development, and responses to biotic and abiotic stresses. Deciphering the potential relationships between lncRNAs and miRNAs may provide new insights into our understanding of the biological functions of lncRNAs and miRNAs, and the pathogenesis of complex diseases. In contrast, gathering information on lncRNA-miRNA interactions (LMIs) through biological experiments is expensive and time-consuming. With the accumulation of multi-omics data, computational models are extremely attractive in systematically exploring potential LMIs. To the best of our knowledge, this is the first comprehensive review of computational methods for identifying LMIs. Specifically, we first summarized the available public databases for predicting animal and plant LMIs. Second, we comprehensively reviewed the computational methods for predicting LMIs and classified them into two categories, including network-based methods and sequence-based methods. Third, we analyzed the standard evaluation methods and metrics used in LMI prediction. Finally, we pointed out some problems in the current study and discuss future research directions. Relevant databases and the latest advances in LMI prediction are summarized in a GitHub repository https://github.com/sheng-n/lncRNA-miRNA-interaction-methods, and we'll keep it updated.
Nan Sheng, Lan Huang 0002, Yangkun Cao, Xuping Xie, Yan Wang 0028
IEEE ACM Trans. Comput. Biol. Bioinform.4
2022 Multi-channel graph attention autoencoders for disease-related lncRNAs prediction
abstract
MOTIVATION: Predicting disease-related long non-coding RNAs (lncRNAs) can be used as the biomarkers for disease diagnosis and treatment. The development of effective computational prediction approaches to predict lncRNA-disease associations (LDAs) can provide insights into the pathogenesis of complex human diseases and reduce experimental costs. However, few of the existing methods use microRNA (miRNA) information and consider the complex relationship between inter-graph and intra-graph in complex-graph for assisting prediction. RESULTS: In this paper, the relationships between the same types of nodes and different types of nodes in complex-graph are introduced. We propose a multi-channel graph attention autoencoder model to predict LDAs, called MGATE. First, an lncRNA-miRNA-disease complex-graph is established based on the similarity and correlation among lncRNA, miRNA and diseases to integrate the complex association among them. Secondly, in order to fully extract the comprehensive information of the nodes, we use graph autoencoder networks to learn multiple representations from complex-graph, inter-graph and intra-graph. Thirdly, a graph-level attention mechanism integration module is adopted to adaptively merge the three representations, and a combined training strategy is performed to optimize the whole model to ensure the complementary and consistency among the multi-graph embedding representations. Finally, multiple classifiers are explored, and Random Forest is used to predict the association score between lncRNA and disease. Experimental results on the public dataset show that the area under receiver operating characteristic curve and area under precision-recall curve of MGATE are 0.964 and 0.413, respectively. MGATE performance significantly outperformed seven state-of-the-art methods. Furthermore, the case studies of three cancers further demonstrate the ability of MGATE to identify potential disease-correlated candidate lncRNAs. The source code and supplementary data are available at https://github.com/sheng-n/MGATE. CONTACT: [email protected], [email protected].
Nan Sheng, Lan Huang 0002, Yan Wang 0028, Ping Xuan, Yangkun Cao
Briefings Bioinform.7
2022 Learning multi-scale heterogenous network topologies and various pairwise attributes for drug-disease association prediction
abstract
MOTIVATION: Identifying new therapeutic effects for the approved drugs is beneficial for effectively reducing the drug development cost and time. Most of the recent computational methods concentrate on exploiting multiple kinds of information about drugs and disease to predict the candidate associations between drugs and diseases. However, the drug and disease nodes have neighboring topologies with multiple scales, and the previous methods did not fully exploit and deeply integrate these topologies. RESULTS: We present a prediction method, multi-scale topology learning for drug-disease (MTRD), to integrate and learn multi-scale neighboring topologies and the attributes of a pair of drug and disease nodes. First, for multiple kinds of drug similarities, multiple drug-disease heterogenous networks are constructed respectively to integrate the similarities and associations related to drugs and diseases. Moreover, each heterogenous network has its specific topology structure, which is helpful for learning the corresponding specific topology representation. We formulate the topology embeddings for each drug node and disease node by random walking on each heterogeneous network, and the embeddings cover the neighboring topologies with different scopes. Because the multi-scale topology embeddings have context relationships, we construct Bi-directional long short-term memory-based module to encode these embeddings and their relationships and learn the neighboring topology representation. We also design the attention mechanisms at feature level and at scale level to obtain the more informative pairwise features and topology embeddings. A module based on multi-layer convolutional networks is constructed to learn the representative attributes of the drug-disease node pair according to their related similarity and association information. Comprehensive experimental results indicate that MTRD achieves the superior performance than several state-of-the-art methods for predicting drug-disease associations. MTRD also retrieves more actual drug-disease associations in the top-ranked candidates of the prediction result. Case studies on five drugs further demonstrate MTRD's ability in discovering the potential candidate diseases for the interested drugs.
Hongda Zhang, Hui Cui 0002, Tiangang Zhang, Yangkun Cao, Ping Xuan
Briefings Bioinform.4
2021 Autoencoder-based drug-target interaction prediction by preserving the consistency of chemical properties and functions of drugs
abstract
MOTIVATION: Exploring the potential drug-target interactions (DTIs) is a key step in drug discovery and repurposing. In recent years, predicting the probable DTIs through computational methods has gradually become a research hot spot. However, most of the previous studies failed to judiciously take into account the consistency between the chemical properties of drug and its functions. The changes of these relationships may lead to a severely negative effect on the prediction of DTIs. RESULTS: We propose an autoencoder-based method, AEFS, under spatial consistency constraints to predict DTIs. A heterogeneous network is established to integrate the information of drugs, proteins and diseases. The original drug features are projected to an embedding (protein) space by a multi-layer encoder, and further projected into label (disease) space by a decoder. In this process, the clinical information of drugs is introduced to assist the DTI prediction. By maintaining the distribution of drug correlation in the original feature, embedding and label space, AEFS keeps the consistency between chemical properties and functions of drugs. Experimental comparisons indicate that AEFS is more robust for imbalanced data and of significantly superior performance in DTI prediction. Case studies further confirm its ability to mine the latent DTIs. AVAILABILITY AND IMPLEMENTATION: The code of AEFS is available at https://github.com/JackieSun818/AEFS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Chang Sun 0002, Yangkun Cao, Jinmao Wei 0001, Jian Liu 0040
Bioinform.2
2019 Drug repositioning through integration of prior knowledge and projections of drugs and diseases
abstract
MOTIVATION: Identifying and developing novel therapeutic effects for existing drugs contributes to reduction of drug development costs. Most of the previous methods focus on integration of the heterogeneous data of drugs and diseases from multiple sources for predicting the candidate drug-disease associations. However, they fail to take the prior knowledge of drugs and diseases and their sparse characteristic into account. It is essential to develop a method that exploits the more useful information to predict the reliable candidate associations. RESULTS: We present a method based on non-negative matrix factorization, DisDrugPred, to predict the drug-related candidate disease indications. A new type of drug similarity is firstly calculated based on their associated diseases. DisDrugPred completely integrates two types of disease similarities, the associations between drugs and diseases, and the various similarities between drugs from different levels including the chemical structures of drugs, the target proteins of drugs, the diseases associated with drugs and the side effects of drugs. The prior knowledge of drugs and diseases and the sparse characteristic of drug-disease associations provide a deep biological perspective for capturing the relationships between drugs and diseases. Simultaneously, the possibility that a drug is associated with a disease is also dependant on their projections in the low-dimension feature space. Therefore, DisDrugPred deeply integrates the diverse prior knowledge, the sparse characteristic of associations and the projections of drugs and diseases. DisDrugPred achieves superior prediction performance than several state-of-the-art methods for drug-disease association prediction. During the validation process, DisDrugPred also can retrieve more actual drug-disease associations in the top part of prediction result which often attracts more attention from the biologists. Moreover, case studies on five drugs further confirm DisDrugPred's ability to discover potential candidate disease indications for drugs. AVAILABILITY AND IMPLEMENTATION: The fourth type of drug similarity and the predicted candidates for all the drugs are available at https://github.com/pingxuan-hlju/DisDrugPred. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ping Xuan, Yangkun Cao, Tiangang Zhang, Xiao Wang 0017, Shuxiang Pan, Tonghui Shen
Bioinform.2