Chen Cao 0002

dblp:44/5148-2 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0001-5343-808XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 3 first-author · 13 since 2021
YearPublicationVenuePosition
2026 IHGCN-PLA: An interpretable heterogeneous graph convolutional network for protein-ligand binding affinity prediction with multimodal interaction fusion
Guishen Wang, Yuxiang Kong, Yuyouqiang Fu, Chen Cao 0002
J. Biomed. Informatics5
2025 TransFusionDR: A Framework for Drug Repositioning via Contrastive and High-Order Feature Fusion with Transformers
abstract
Drug repositioning can effectively reduce research and development costs and accelerate time to market by identifying new indications for existing drugs. In recent years, deep learning-based methods have achieved remarkable progress in the field of drug repositioning. However, current approaches still suffer from several limitations. First, many methods simply concatenate structural information and association information without fully capturing their intrinsic relationships, leading to suboptimal information fusion. Second, most models lack mechanisms for extracting high-order structural information and aligning heterogeneous and homogeneous features, which limits the model's expressiveness and predictive performance. To address these limitations, we propose TransFusionDR, a framework for drug repositioning via contrastive and high-order feature fusion with transformers. We employ a Graph Transformer to extract deep structural features from drug-drug and disease-disease graphs, and utilize a Heterogeneous Graph Transformer to capture semantic information from the drug-disease association graph. These features are then refined through contrastive learning to enhance semantic consistency and improve information fusion quality. We further introduce a Transformer Encoder to deeply integrate the homogeneous and heterogeneous features by dynamically modeling semantic dependencies, enabling the extraction of high-order interactions and achieving more effective feature alignment. Experimental results on two benchmark datasets demonstrate that the proposed framework significantly improves prediction accuracy and robustness in drug repositioning tasks, outperforming state-of-the-art methods.
Guishen Wang, Honghan Chen, Zhitong Guo, Chen Cao 0002, Xiaoxuan Gong
BIBM5
2025 MMDDI-GDSS: Multi-Modal Feature Fusion Drug-Drug Interaction Event Prediction Model
abstract
Combination drug therapy represents a cornerstone of modern medicine, offering enhanced therapeutic efficacy and reduced drug resistance. Predicting adverse drug-drug inter-action (DDI) events is crucial, and while multi-modal models show significant promise, the effective integration of diverse data sources remains an open challenge. In this work, we introduce MMDDI-GDSS, a novel multi-modal framework for DDI event prediction. Our framework leverages a multi-head attention mechanism to fuse diverse data modalities-including SMILES, protein targets, enzymes, and pharmacological information-into a unified, enhanced molecular attribute graph. To learn robust and expressive drug representations from this complex graph, MMDDI-GDSS employs a graph diffusion process over static subgraphs, generating a powerful and coherent representation for each drug. We benchmark MMDDI-GDSS against several state-of-the-art methods on two widely used DDI event prediction datasets. Experimental results demonstrate that our model consistently outperforms all baselines. Furthermore, comprehensive ablation studies validate the effectiveness of our key components, highlighting the distinct contributions of the attention-based feature integration and the graph diffusion mechanism.
Guishen Wang, Handan Wang, Honghan Chen, Chen Cao 0002
BIBM4
2025 MMDDI-SSE: A Novel Multi-Modal Feature Fusion Model With Static Subgraph Embedding for Drug-Drug Interaction Event Prediction
abstract
Artificial intelligence techniques play a pivotal role in the accurate identification of drug-drug interaction (DDI) events, thereby informing clinical decisions and treatment regimens. While existing DDI prediction models have made significant progress by leveraging sequence features such as chemical substructures, targets, and enzymes, they often face limitations in integrating and effectively utilizing multi-modal drug representations. To address these limitations, this study proposes a novel multi-modal feature fusion model for DDI event prediction: MMDDI-SSE. Our approach integrates drug sequence modality with DDI graph representations through a novel architecture that employs static subgraph generation to capture structural properties. The model utilizes a graph autoencoder architecture to learn both local and global topological features from these subgraphs, while simultaneously processing diverse sequence-based characteristics including semantically enhanced pharmacodynamic features, chemical substructures, target proteins, and enzyme information. Through comprehensive evaluation on two distinct datasets, MMDDI-SSE demonstrates superior predictive performance compared to state-of-the-art baselines. Ablation studies further validate the effectiveness of each architectural component in enhancing DDI prediction accuracy.
Guishen Wang, Honghan Chen, Handan Wang, Hairong Gao, Chen Cao 0002
IEEE J. Biomed. Health Informatics6
2025 Medical Graph Diffusion: Hybrid Graph Diffusion With Heterogeneous Graph Convolutional Networks for Medical Text Classification
abstract
Text classification is a critical task for understanding the knowledge behind text, especially in medical text. In this paper, we propose a medical graph diffusion model, named the MGD model, for the medical text classification task. To model more structural relationships within a document, our MGD model constructs a text heterogeneous graph to represent word-level, sentence-level, and word-sentence-level structural relationships. To overcome the limitation of only considering direct neighbors, a graph diffusion convolution is employed to reconstruct the text heterogeneous graph. Subsequently, a heterogeneous graph convolutional network and a multilayer perceptron are used to complete the medical text classification task. To evaluate the performance of our MGD model, various text classification benchmarks, including long text standard benchmarks, short text standard benchmarks, and medical text benchmarks, are used to comprehensively assess the effectiveness and robustness of our MGD model. Compared with other representative baselines, it achieved notable improvements in both Accuracy and F1 score evaluation metrics. Ablation experiment results further demonstrated that the construction of heterogeneous graphs and the use of diffusion graph convolutional networks significantly impact the performance of our MGD model.
Guishen Wang, Keshuang Liu, Chen Cao 0002
IEEE J. Biomed. Health Informatics5
2024 MMDDI-MGPFF: Multi-Modal Drug Representation Learning with Molecular Graph and Pharmacological Feature Fusion for Drug-Drug Interaction Event Prediction
abstract
Drug-drug interactions pose a significant challenge in healthcare, directly impacting patient safety and treatment efficacy. Although recent advances in computational methods have improved drug-drug interaction (DDI) event prediction, many existing models face difficulties in effectively fusing features across different DDI tasks. To address these limitations, we introduce MMDDI-MGPFF, a novel multi-modal drug representation learning framework that integrates molecular graphs and pharmacological feature fusion to improve DDI prediction. Our model leverages a graph isomorphism network (GIN) for efficient encoding of molecular structures, coupled with an autoencoder for learning sequence-based drug features, encompassing both biological and pharmacological characteristics. We propose a multi-modal fusion approach that employs multi-head attention and deep neural networks to seamlessly integrate these graph and sequence modalities. Comprehensive experiments on a benchmark dataset demonstrate that MMDDI-MGPFF significantly outperforms state-of-the-art methods in DDI prediction tasks. Ablation studies further corroborate the efficacy of our model’s components, particularly the GIN and the integration of multi-feature drug representations. This work advances the field of DDI prediction by providing a more holistic and accurate approach to drug representation and interaction modeling.
Guishen Wang, Zhitong Guo, Guilin You, Chen Cao 0002
BIBM5
2024 EHR-HGCN: An Enhanced Hybrid Approach for Text Classification Using Heterogeneous Graph Convolutional Networks in Electronic Health Records
abstract
Text classification is a central part of natural language processing, with important applications in understanding the knowledge behind biomedical texts including electronic health records (EHR). In this article, we propose a novel heterogeneous graph convolutional network method for classifying EHR texts. Our method, called EHR-HGCN, is able to combine context-sensitive word and sentence embeddings with structural sentence-level and word-level relation information to perform text classification. EHR-HGCN reframes EHR text classification as a graph classification task to better capture structural information about the document using a heterogeneous graph. To mine contextual information from a document, EHR-HGCN first applies a bidirectional recurrent neural network (BiRNN) on word embeddings obtained via Global Vectors for word representation (GloVe) to obtain context-sensitive word-level and sentence-level embeddings. To mine structural relationships from the document, EHR-HGCN then constructs a heterogeneous graph over the word and sentence embeddings, where sentence-word and word-word relationships are represented by graph edges. Finally, a heterogeneous graph convolutional neural network is used to classify documents by their graph representation. We evaluate EHR-HGCN on a variety of standard text classification benchmarks and find that EHR-HGCN has higher accuracy and F1-score than other representative machine learning and deep learning methods. We also apply EHR-HGCN to the MedLit benchmark and find it performs with high accuracy and F1-score on the task of section classification in EHR texts. Our ablation experiments show that the heterogeneous graph construction and heterogeneous graph convolutional network are critical to the performance of EHR-HGCN.
Guishen Wang, Xiaoxue Lou, Devin Kwok, Chen Cao 0002
IEEE J. Biomed. Health Informatics5
2023 Human-Spa: An Online Platform Based on Spatial Transcriptome Data for Diseases of Human Systems
abstract
Spatial transcriptomics has become a major method for high-throughput analysis of gene expression at the current level of cells, which can directly study gene expression changes in disease cells, reveal the occurrence and development mechanism of diseases, identify potential therapeutic targets related to diseases, and provide new clues for disease diagnosis and treatment. Although there have been major breakthroughs in the analysis and acquisition of transcriptome data, there are still many challenges in how to effectively manage, share and utilize these valuable data resources in the study of human systemic diseases. To solve this problem, we are committed to building a spatial transcriptome database website related to human systemic diseases, aiming to provide reliable datasets related to multiple diseases, and provide certain data information and analysis results to accelerate the progress of human systemic diseases research. This paper will introduce the construction process of the database website in detail, and discuss its application prospects in disease research, with the aim of promoting further development and innovation in the field of human health. Here, Human-spa mainly includes 12 Human systems, 38 disease types, 55 datasets, and Human-Spa provides a very friendly web interface for visualization and dataset parsing. In conclusion, the construction of Human-Spa will provide a powerful tool and resource platform for Human disease research. Human-Spa is available for free at http://www.human-spa.cn/
Yunyun Su, Feifei Cui, Shiyu Yan, Quan Zou 0001, Chen Cao 0002
BIBM5
2022 Botanical drugs: a new strategy for structure-based target prediction
abstract
Target identification of small molecules is an important and still changeling work in the area of drug discovery, especially for botanical drug development. Indistinct understanding of the relationships of ligand-protein interactions is one of the main obstacles for drug repurposing and identification of off-targets. In this study, we collected 9063 crystal structures of ligand-binding proteins released from January, 1995 to April, 2021 in PDB bank, and split the complexes into 5133 interaction pairs of ligand atoms and protein fragments (covalently linked three heavy atoms) with interatomic distance ≤5 Å. The interaction pairs were grouped into ligand atoms with the same SYBYL atom type surrounding each type of protein fragment, which were further clustered via Bayesian Gaussian Mixture Model (BGMM). Gaussian distributions with ligand atoms ≥20 were identified as significant interaction patterns. Reliability of the significant interaction patterns was validated by comparing the difference of number of significant interaction patterns between the docked poses with higher and lower similarity to the native crystal structures. Fifty-one candidate targets of brucine, strychnine and icajine involved in Semen Strychni (Mǎ Qián Zǐ) and eight candidate targets of astragaloside-IV, formononetin and calycosin-7-glucoside involved in Astragalus (Huáng Qí) were predicted by the significant interaction patterns, in combination with docking, which were consistent with the therapeutic effects of Semen Strychni and Astragalus for cancer and chronic pain. The new strategy in this study improves the accuracy of target identification for small molecules, which will facilitate discovery of botanical drugs.
Xuxu Wei, Zeyu Cheng, Qingming Wu, Chen Cao 0002, Hongcai Shang
Briefings Bioinform.5
2022 webSCST: an interactive web application for single-cell RNA-sequencing data and spatial transcriptomic data integration
abstract
SUMMARY: Integrative analysis of single-cell RNA-sequencing (scRNA-seq) data with spatial data for the same species and organ would provide each cell sample with a predictive spatial location, which would facilitate biological study. However, publicly available spatial sequencing datasets for specific species and organs are rare and are often displayed in different formats. In this study, we introduce a new web-based scRNA-seq analysis tool, webSCST, that integrates well-organized spatial transcriptome sequencing datasets categorized by species and organs, provides a user-friendly interface for raw single-cell processing with popular integration methods and allows users to submit their raw scRNA-seq data once to obtain predicted spatial locations for each cell type. AVAILABILITY AND IMPLEMENTATION: webSCST implemented in shiny with all major browsers supported is available at http://www.webscst.com. webSCST is also freely available as an R package at https://github.com/swsoyee/webSCST.
Feifei Cui, Lijun Dou, Chen Cao 0002, Quan Zou 0001
Bioinform.6
2022 A novel method for drug-target interaction prediction based on graph transformers model
abstract
BACKGROUND: Drug-target interactions (DTIs) prediction becomes more and more important for accelerating drug research and drug repositioning. Drug-target interaction network is a typical model for DTIs prediction. As many different types of relationships exist between drug and target, drug-target interaction network can be used for modeling drug-target interaction relationship. Recent works on drug-target interaction network are mostly concentrate on drug node or target node and neglecting the relationships between drug-target. RESULTS: We propose a novel prediction method for modeling the relationship between drug and target independently. Firstly, we use different level relationships of drugs and targets to construct feature of drug-target interaction. Then, we use line graph to model drug-target interaction. After that, we introduce graph transformer network to predict drug-target interaction. CONCLUSIONS: This method introduces a line graph to model the relationship between drug and target. After transforming drug-target interactions from links to nodes, a graph transformer network is used to accomplish the task of predicting drug-target interactions.
Mengyan Du, Guishen Wang, Chen Cao 0002
BMC Bioinform.5
2021 kTWAS: integrating kernel machine with transcriptome-wide association studies improves statistical power and reveals novel genes
abstract
The power of genotype-phenotype association mapping studies increases greatly when contributions from multiple variants in a focal region are meaningfully aggregated. Currently, there are two popular categories of variant aggregation methods. Transcriptome-wide association studies (TWAS) represent a set of emerging methods that select variants based on their effect on gene expressions, providing pretrained linear combinations of variants for downstream association mapping. In contrast to this, kernel methods such as sequence kernel association test (SKAT) model genotypic and phenotypic variance use various kernel functions that capture genetic similarity between subjects, allowing nonlinear effects to be included. From the perspective of machine learning, these two methods cover two complementary aspects of feature engineering: feature selection/pruning and feature aggregation. Thus far, no thorough comparison has been made between these categories, and no methods exist which incorporate the advantages of TWAS- and kernel-based methods. In this work, we developed a novel method called kernel-based TWAS (kTWAS) that applies TWAS-like feature selection to a SKAT-like kernel association test, combining the strengths of both approaches. Through extensive simulations, we demonstrate that kTWAS has higher power than TWAS and multiple SKAT-based protocols, and we identify novel disease-associated genes in Wellcome Trust Case Control Consortium genotyping array data and MSSNG (Autism) sequence data. The source code for kTWAS and our simulations are available in our GitHub repository (https://github.com/theLongLab/kTWAS).
Chen Cao 0002, Devin Kwok, Shannon Edie, Qing Li 0076, Bowei Ding, Pathum Kossinna, Simone Campbell, Jingjing Wu 0002, Matthew Greenberg
Briefings Bioinform.1
2021 WgLink: reconstructing whole-genome viral haplotypes using L0+L1-regularization
abstract
SUMMARY: Many tools can reconstruct viral sequences based on next-generation sequencing reads. Although existing tools effectively recover local regions, their accuracy suffers when reconstructing the whole viral genomes (strains). Moreover, they consume significant memory when the sequencing coverage is high or when the genome size is large. We present WgLink to meet this challenge. WgLink takes local reconstructions produced by other tools as input and patches the resulting segments together into coherent whole-genome strains. We accomplish this using an L0+L1-regularized regression, synthesizing variant allele frequency data with physical linkage between multiple variants spanning multiple regions simultaneously. WgLink achieves higher accuracy than existing tools both on simulated and on real datasets while using significantly less memory (RAM) and fewer CPU hours. AVAILABILITY AND IMPLEMENTATION: Source code and binaries are freely available at https://github.com/theLongLab/wglink. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Chen Cao 0002, Matthew Greenberg
Bioinform.1
2019 PRESM: personalized reference editor for somatic mutation discovery in cancer genomics
abstract
MOTIVATION: Accurate detection of somatic mutations is a crucial step toward understanding cancer. Various tools have been developed to detect somatic mutations from cancer genome sequencing data by mapping reads to a universal reference genome and inferring likelihoods from complex statistical models. However, read mapping is frequently obstructed by mismatches between germline and somatic mutations on a read and the reference genome. Previous attempts to develop personalized genome tools are not compatible with downstream statistical models for somatic mutation detection. RESULTS: We present PRESM, a tool that builds personalized reference genomes by integrating germline mutations into the reference genome. The aforementioned obstacle is circumvented by using a two-step germline substitution procedure, maintaining positional fidelity using an innovative workaround. Reads derived from tumor tissue can be positioned more accurately along a personalized reference than a universal reference due to the reduced genetic distance between the subject (tumor genome) and the target (the personalized genome). Application of PRESM's personalized genome reduced false-positive (FP) somatic mutation calls by as much as 55.5%, and facilitated the discovery of a novel somatic point mutation on a germline insertion in PDE1A, a phosphodiesterase associated with melanoma. Moreover, all improvements in calling accuracy were achieved without parameter optimization, as PRESM itself is parameter-free. Hence, similar increases in read mapping and decreases in the FP rate will persist when PRESM-built genomes are applied to any user-provided dataset. AVAILABILITY AND IMPLEMENTATION: The software is available at https://github.com/precisionomics/PRESM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Chen Cao 0002, Lauren Mak, Guangxu Jin, Paul Gordon, Kai Ye 0001
Bioinform.1
2019 Using discriminative vector machine model with 2DPCA to predict interactions among proteins
abstract
BACKGROUND: The interactions among proteins act as crucial roles in most cellular processes. Despite enormous effort put for identifying protein-protein interactions (PPIs) from a large number of organisms, existing firsthand biological experimental methods are high cost, low efficiency, and high false-positive rate. The application of in silico methods opens new doors for predicting interactions among proteins, and has been attracted a great deal of attention in the last decades. RESULTS: Here we present a novelty computational model with the adoption of our proposed Discriminative Vector Machine (DVM) model and a 2-Dimensional Principal Component Analysis (2DPCA) descriptor to identify candidate PPIs only based on protein sequences. To be more specific, a 2DPCA descriptor is employed to capture discriminative feature information from Position-Specific Scoring Matrix (PSSM) of amino acid sequences by the tool of PSI-BLAST. Then, a robust and powerful DVM classifier is employed to infer PPIs. When applied on both gold benchmark datasets of Yeast and H. pylori, our model obtained mean prediction accuracies as high as of 97.06 and 92.89%, respectively, which demonstrates a noticeable improvement than some state-of-the-art methods. Moreover, we constructed Support Vector Machines (SVM) based predictive model and made comparison it with our model on Human benchmark dataset. In addition, to further demonstrate the predictive reliability of our proposed method, we also carried out extensive experiments for identifying cross-species PPIs on five other species datasets. CONCLUSIONS: All the experimental results indicate that our method is very effective for identifying potential PPIs and could serve as a practical approach to aid bioexperiment in proteomics research.
Zhengwei Li 0001, Ru Nie, Zhu-Hong You, Chen Cao 0002, Jiashu Li
BMC Bioinform.4