Xiangzheng Fu

dblp:231/2047 · DBLP profile ↗
← Back
39ranked-venue papers
4as first author
38since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 33 · 4 first-author · 32 since 2021Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 CAML: A Conflict-Aware Molecular Language Model Merging Framework for Multi-Constraint Molecular Generation
abstract
Xuanbai Ren, Luoda Tan, Pei Liu, Tengfei Ma, Xiangzheng Fu, Longyue Wang, Yiping Liu, Xiangxiang Zeng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xuanbai Ren, Luoda Tan, Pei Liu 0008, Tengfei Ma 0002, Xiangzheng Fu, Longyue Wang, Xiangxiang Zeng
ACL (1)5
2026 A deep adversarial network model for multi-task analysis of single-cell omics data
abstract
Single-cell multi-omics data reveal complex cellular states and deepen our understanding of tissue cell phenotypes and functions. However, data analysis remains challenging due to the discrete nature and high noise level of the data, as well as the lack of modality. Here, we propose scMultiNet, a multi-task deep adversarial neural network that can integrate different tasks to analyze single-cell multi-modal data. In particular, we achieve joint training of multi-modal integration and cross-modal prediction tasks by introducing a cross-modal bi-prediction module and a multi-head self-attention module. Data denoising is further enhanced by integrating an indicator matrix that constrains and precisely reconstructs the original expression values. Extensive simulations and real data experiments demonstrate that scMultiNet outperforms existing state-of-the-art methods in dimensionality reduction, visualization, clustering, batch elimination, data denoising, multi-modal integration, single-cell cross-modality translation, and in revealing cell type-specific biological insights. In addition, we demonstrate that scMultiNet can effectively transfer the complex relationships between modalities from one batch to another. In summary, scMultiNet stands as a comprehensive end-to-end framework, ideally suited for analyzing single-cell multi-omics data.
Junlin Xu, Yajie Meng, Shuting Jin, Changcheng Lu, Feifei Cui, Xiangzheng Fu, Quan Zou 0001, Xiangxiang Zeng
Briefings Bioinform.8
2026 DynaTCR: dynamic hard-negative ensemble graph learning improves TCR-epitope binding prediction
abstract
MOTIVATION: T-cell receptors (TCRs) recognize antigenic peptides presented by major histocompatibility complex (MHC) molecules and are central to adaptive immunity. Computational prediction of TCR-epitope binding (TEB) can accelerate immunotherapy development, yet remains hampered by limited labeled data, false-negative noise in unobserved pairs, and over-smoothing in graph-based models. RESULTS: We present DynaTCR, a dynamic graph ensemble learning framework for TEB prediction. DynaTCR encodes TCR and epitope sequences with protein language model embeddings and organizes them into a bipartite interaction graph. A graph regularization-variance-preserving aggregation (GR-VPA) encoder stabilizes message propagation and alleviates over-smoothing, while a global attention layer captures long-range dependencies. Multiple base learners are trained with iteratively updated hard-negative samples to reduce false-negative predictions. Under the StrictTCR evaluation protocol on four public datasets, DynaTCR achieves AUC improvements of 4.0-8.2 percentage points over the strongest existing method and up to 15.8 percentage points in AUPR. On the most stringently curated dataset, DynaTCR attains an AUC of 95.1%. Furthermore, on an independent structure-derived test set, DynaTCR achieves the highest AUC (72.6%) among all compared methods, demonstrating its robustness and effectiveness for TEB prediction and candidate prioritization. AVAILABILITY: Source code and data can be downloaded from: https://github.com/2014402680/TEB/.
Xiangzheng Fu, Xinyu Zhang 0012, Linlin Zhuo, Dong-Sheng Cao 0001, Quan Zou 0001
Bioinform.1
2026 DrugKANs: A Paradigm to Enhance Drug-Target Interaction Prediction With KANs
abstract
Identifyingpotential drug-target interactions (DTIs) is crucial for understanding drug mechanisms, and recent computational methods have yielded promising results in this area. However, these methods face several challenges, including limited model generalization due to heavy reliance on multiple similarity datasets and complex feature extraction, as well as a lack of interpretability by ignoring intrinsic information about drugs and targets. To address these challenges, we propose DrugKANs, a novel DTI prediction model that enhances both the quality and interpretability of DTI representations by integrating a dual-tower architecture with Kolmogorov-Arnold Network (KAN) technology. Our model involves utilizing a pre-trained model to derive initial representations of drugs and targets, and employing a lightweight attention mechanism to capture key features, thereby improving representation quality. We leverage the dual-tower architecture and a lightweight feature interaction mechanism to extract high-level representations separately for drugs and targets, aiming to reduce complex feature interactions and mitigate overfitting. Additionally, we incorporate a contrastive learning strategy within the drug-target bipartite graph to address sparse neighborhood effects and enhance topological information. The inclusion of KAN technology further improves the interpretability of the DTI prediction model. Experimental results on public datasets demonstrate that our model predicts DTIs effectively, underscoring its potential as a valuable tool in drug discovery. This comprehensive methodology presents a balanced approach to overcoming the identified challenges in DTI prediction.
Xiangzheng Fu, Zhenya Du, Haiting Chen, Linlin Zhuo, Aiping Lu, Dong-Sheng Cao 0001
IEEE J. Biomed. Health Informatics1
2026 BloodPatrol: Revolutionizing Blood Cancer Diagnosis - Advanced Real-Time Detection Leveraging Deep Learning & Cloud Technologies
abstract
Cloud computing and Internet of Things (IoT) technologies are gradually becoming the technological changemakers in cancer diagnosis. Blood cancer is an aggressive disease affecting the blood, bone marrow, and lymphatic system, and its early detection is crucial for subsequent treatment. Flow cytometry has been widely studied as a commonly used method for detecting blood cancer. However, the high computation and resource consumption severely limit its practical application, especifically in regions with limited medical and computational resources. In this study, with the help of cloud computing and IoT technologies, we develop a novel blood cancer dynamic monitoring diagnostic model named BloodPatrol based on an intelligent feature weight fusion mechanism. The proposed model is capable of capturing the dual-view importance relationship between cell samples and features, greatly improving prediction accuracy and significantly surpassing previous models. Besides, benefiting from the powerful processing ability of cloud computing, BloodPatrol can run on a distributed network to efficiently process large-scale cell data, which provides immediate and scalable blood cancer diagnostic services.
Jinhang Wei, Longyue Wang, Zhecheng Zhou, Linlin Zhuo, Xiangxiang Zeng, Xiangzheng Fu, Quan Zou 0001, Keqin Li 0001, Zhongjun Zhou
IEEE J. Biomed. Health Informatics6
2025 GSToxi: Gated Cross-Modal Modeling With Graph-Sequence Encoders for Peptide Toxicity Prediction
abstract
Peptide-based therapeutics hold great potential, yet their cytotoxicity remains a key challenge in drug development. Most existing toxicity prediction models rely solely on sequence information, often overlooking the fusion of multimodal submolecular patterns. We propose GSToxi, a multimodal deep learning framework that leverage sequence and molecular graph featurizer to enhance peptide toxicity prediction. A shared gating mechanism is employed to facilitate semantic alignment and cross-modal integration, while a contrastive regularization loss further optimizes latent-space consistency throughout the training process. Furthermore, GSToxi incorporates embeddings from pre-trained protein language models alongside low-level compositional priors, enabling the capture of both global contextual semantics and local structural features. Experimental results show that GSToxi outperforms state-of-the-art baselines across multiple evaluation metrics on an independent test set. Ablation studies underscore the critical contributions of each component, with the molecular graph encoder and pre-trained embeddings proving particularly impactful. This work offers a generalizable and robust framework for peptide toxicity prediction and provides valuable insights for future multimodal modeling of biological molecules.
Li Wang 0145, Xiangzheng Fu, Xiucai Ye, Tetsuya Sakurai
BIBM2
2025 ET-PROTACs: modeling ternary complex interactions using cross-modal learning and ternary attention for accurate PROTAC-induced degradation prediction
abstract
MOTIVATION: Accurately predicting the degradation capabilities of proteolysis-targeting chimeras (PROTACs) for given target proteins and E3 ligases is important for PROTAC design. The distinctive ternary structure of PROTACs presents a challenge to traditional drug-target interaction prediction methods, necessitating more innovative approaches. While current state-of-the-art (SOTA) methods using graph neural networks (GNNs) can discern the molecular structure of PROTACs and proteins, thus enabling the efficient prediction of PROTACs' degradation capabilities, they rely heavily on limited crystal structure data of the POI-PROTAC-E3 ternary complex. This reliance underutilizes rich PROTAC experimental data and neglects intricate interaction relationships within ternary complexes. RESULTS: In this study, we propose a model based on cross-modal strategy and ternary attention technology, ET-PROTACs, to predict the targeted degradation capabilities of PROTACs. Our model capitalizes on the strengths of cross-modal methods by using equivariant GNN graph neural networks to process the graph structure and spatial coordinates of PROTAC molecules concurrently while utilizing sequence-based methods to learn the protein sequence information. This integration of cross-modal information is cohesively harnessed and channeled into a ternary attention mechanism, specially tailored for the unique structure of PROTACs, enabling the congruent modeling of both PROTAC and protein modalities. Experimental results demonstrate that the ET-PROTACs model outperforms existing SOTA methods. Moreover, visualizing attention scores illuminates crucial residues and atoms pivotal in specific POI-PROTAC-E3 interactions, thus offering invaluable insights and guidance for future pharmaceutical research. AVAILABILITY AND IMPLEMENTATION: The codes of our model are available at https://github.com/GuanyuYue/ET-PROTACs.
Guanyu Yue, Li Wang 0145, Quan Zou 0001, Xiangzheng Fu, Dong-Sheng Cao 0001
Briefings Bioinform.7
2025 GRAPE: graph-regularized protein language modeling unlocks TCR-epitope binding specificity
abstract
T-cell receptor (TCR)-epitope binding prediction is critical for immunotherapies but remains challenged by sparse interaction networks and severe class imbalance in training data. Current graph neural network (GNN) approaches for predicting TCR-epitope binding (TEB) fail to address two key limitations: over-smoothing during message propagation in sparse TCR-epitope graphs and biased predictions toward dominant epitope-TCR pairs. Here, we present GRAPE (Graph-Regularized Attentive Protein Embeddings), a framework unifying spectral graph regularization and imbalance-aware learning. GRAPE first leverages protein language models (ESM-2) to generate evolutionary-informed TCR/epitope embeddings, constructing a topology-aware interaction graph. To mitigate over-smoothing, we introduce spectral graph regularization, explicitly constraining node feature smoothness to preserve discriminative patterns in sparse neighborhoods. Simultaneously, a dynamic edge reweighting module prioritizes unobserved TCR-epitope edges during graph propagation, coupled with a differentiable area under the ROC curve-maximization objective that directly optimizes for imbalance resilience. Extensive benchmarking on public datasets demonstrates that GRAPE significantly outperforms state-of-the-art methods in TEB prediction. This work establishes GRAPE as a robust framework for elucidating TCR-epitope interactions, with broad applications in immunology research and therapeutic design.
Xiangzheng Fu, Mingqiang Rong, Dong-Sheng Cao 0001, Sisi Yuan, Aiping Lu
Briefings Bioinform.1
2025 ST-GCP: a graph convolutional network model with contrastive consistency and permutation for spatial transcriptomics
abstract
Spatial transcriptomics (STs) technology is a powerful technique that simultaneously preserves gene expression profiles and spatial information, enabling deeper exploration of tissue organization and function. However, many existing computational approaches often rely on labeled ST data and overlook the rich spatial information, resulting in limited representations and suboptimal clustering. In this paper, we propose ST-GCP, a self-supervised graph representation learning framework for ST data, which incorporates a structure-feature perturbation mechanism. First, ST-GCP applies feature-level random permutation of the gene expression matrix and random edge dropout in the spatial neighbor network, creating two complementary augmented graph views of ST data. ST-GCP then employs a two-layer graph convolutional network (GCN) encoder-decoder to extract spatial representations and reconstruct gene expression. Finally, a cosine-similarity-based contrastive objective aligns the view-specific representations, and the overall loss jointly optimizes reconstruction fidelity and contrastive consistency, thereby coupling graph topology with transcriptomic profiles in a shared low-dimensional space. Experimental results on multiple ST datasets demonstrate that ST-GCP can uncover biologically meaningful patterns, such as tumor heterogeneity, brain developmental architecture, and cellular developmental trajectories.
Yajie Meng, Xianfang Tang, Feifei Cui, Xiangzheng Fu, Quan Zou 0001, Junlin Xu
Briefings Bioinform.7
2025 MOFormer: navigating the antimicrobial peptide design space with Pareto-based multi-objective transformer
abstract
Antimicrobial peptide (AMP) design through deep learning holds the potential to revolutionize antibiotic development. Despite recent progress in AMP generation, designing peptide antibiotics with multiple optimal properties remains a significant challenge. We present MOFormer, an advanced multi-objective AMP design pipeline capable of optimizing multiple AMP properties simultaneously. By leveraging a conditional Transformer, the model refines the AMP sequence-property landscape for efficient multi-objective generation. It also incorporates regularization techniques to maintain a highly structured space, enabling the sampling of precise and desirable candidates. Comparative analyses reveal that MOFormer achieves the optimal hypervolume in the multi-objective space, surpassing advanced methods in simultaneously maximizing antimicrobial activity (minimum inhibitory concentration) and minimizing hemolysis and toxicity, thereby yielding the most promising and desirable set of candidate peptides. When extended to a tri-objective scenario, MOFormer continues to exhibit remarkable optimization performance. Finally, we execute a hierarchical and rapid ranking of generated candidates based on Pareto fronts. We conducted a comprehensive validation of the physicochemical properties and target attributes of the candidates, while AlphaFold structure predictions revealed notably reliable predicted local distance difference test scores ranging from 70% to 87%. Our findings suggest that MOFormer holds potential to accelerate the discovery of efficacious peptide antibiotics by optimizing multi-objective trade-offs.
Li Wang 0145, Xiangzheng Fu, Xiucai Ye, Tetsuya Sakurai, Xiangxiang Zeng
Briefings Bioinform.2
2025 An image-based protein-ligand binding representation learning framework via multi-level flexible dynamics trajectory pre-training
abstract
MOTIVATION: Accurate prediction of protein-ligand binding (PLB) relationships plays a crucial role in drug discovery, which helps identify drugs that modulate the activity of specific targets. Traditional biological assays for measuring PLB relationships are time consuming and costly. In addition, models for predicting PLB relationships have been developed and widely used in drug discovery tasks. However, learning more accurate PLB representations is essential to meet the stringent standards required for drug discovery. RESULTS: We propose an image-based PLB representation learning framework, called ImagePLB, which equips ligand representation learner (LRL) and protein representation learner (PRL) to accept 3D multi-view ligand images and protein graphs as input, respectively, and learns rich interaction information between ligand and protein through a binding representation learner (BRL). Considering the scarcity of protein-ligand pairs, we further propose a multi-level next trajectory prediction (MLNTP) task to pre-train ImagePLB on the 4D flexible dynamics trajectory of 16 972 complexes, including ligand level, protein level, and complex level, to learn information related to trajectories. Besides, by introducing trajectory regularization (TR), we effectively alleviate the problem of high (even almost identical) feature similarity caused by adjacent trajectories. Compared with the current state-of-the-art methods, ImagePLB has achieved competitive improvements on PLB-related prediction tasks, including protein-ligand affinity and efficacy prediction tasks. This study opens the door to the image-based PLB learning paradigm. AVAILABILITY AND IMPLEMENTATION: All data and implementation details of code can be obtained from https://github.com/HongxinXiang/ImagePLB.
Hongxin Xiang, Mingquan Liu, Linlin Hou, Shuting Jin, Jianmin Wang 0016, Jun Xia 0001, Wenjie Du 0003, Sisi Yuan, Xiangzheng Fu, Lei Xu 0047
Bioinform.9
2025 SGPS-IMR: Efficiently inferring microbial resistance using self-supervised graph perturbation strategy
Linlin Zhuo, Zhecheng Zhou, Xiangzheng Fu, Quan Zou 0001
Expert Syst. Appl.4
2025 AEGNN-M:A 3D Graph-Spatial Co-Representation Model for Molecular Property Prediction
abstract
Improving the drug development process can expedite the introduction of more novel drugs that cater to the demands of precision medicine. Accurately predicting molecular properties remains a fundamental challenge in drug discovery and development. Currently, a plethora of computer-aided drug discovery (CADD) methods have been widely employed in the field of molecular prediction. However, most of these methods primarily analyze molecules using low-dimensional representations such as SMILES notations, molecular fingerprints, and molecular graph-based descriptors. Only a few approaches have focused on incorporating and utilizing high-dimensional spatial structural representations of molecules. In light of the advancements in artificial intelligence, we introduce a 3D graph-spatial co-representation model called AEGNN-M, which combines two graph neural networks, GAT and EGNN. AEGNN-M enables learning of information from both molecular graphs representations and 3D spatial structural representations to predict molecular properties accurately. We conducted experiments on seven public datasets, three regression datasets and 14 breast cancer cell line phenotype screening datasets, comparing the performance of AEGNN-M with state-of-the-art deep learning methods. Extensive experimental results demonstrate the satisfactory performance of the AEGNN-M model. Furthermore, we analyzed the performance impact of different modules within AEGNN-M and the influence of spatial structural representations on the model's performance. The interpretability analysis also revealed the significance of specific atoms in determining particular molecular properties.
Xiangzheng Fu, Linlin Zhuo, Quan Zou 0001
IEEE J. Biomed. Health Informatics3
2025 PKAN: Leveraging Kolmogorov-Arnold Networks and Multi-Modal Learning for Peptide Prediction With Advanced Language Models
abstract
Peptides can offer highly specific biological activities, serving as essential mediators of intercellular signaling, which are critical for advancing precision medicine and drug development. Their primary structure can be depicted either as an amino acid sequence or as a chemical molecules consisting of atoms and chemical bonds. Large language models (LLMs) hold the potential to thoroughly elucidate the intricate intrinsic properties of peptides. Here we present the Peptide Kolmogorov-Arnold Network (PKAN), a framework leveraging multi-modal representations inspired by advanced language models for peptide activity and functionality prediction. Comparative experiments across tasks show that PKAN outperforms state-of-the-art models while maintaining a streamlined design with superior predictive capabilities. The multi-modal feature importance scoring, anchored in global structures and the significant marginal impacts of derived features on the model, coupled with intricate symbolic regression of specific activation functions, further demonstrates the robustness and precision of the PKAN framework in identifying and elucidating key determinants of peptide functionality. This work provides scientific evidence for investigating the complex mechanisms of peptide materials and supports the progression of peptide language paradigms in biology.
Li Wang 0145, Xiangzheng Fu, Xiucai Ye, Tetsuya Sakurai, Xiangxiang Zeng
IEEE J. Biomed. Health Informatics2
2025 CardiOT: Towards Interpretable Drug Cardiotoxicity Prediction Using Optimal Transport and Kolmogorov-Arnold Networks
abstract
Investigating the inhibitory effects of compounds on cardiac ion channels is essential for assessing cardiac drug safety. Consequently, researchers have developed computational models to evaluate combined cardiotoxicity (CCT) on cardiac ion channels. However, limitations in experimental data often cause issues like uneven data distribution and scarcity. Additionally, existing models primarily emphasize atomic information flow within graph neural networks (GNNs) while overlooking chemical bonds, leading to inadequate recognition of key structures. Therefore, this study integrates optimal transport (OT), structure remapping (SR), and Kolmogorov-Arnold networks (KANs) into a GNN-based CCT prediction model, CardiOT. First, the proposed CardiOT model employs OT pooling to optimize sample-feature joint distribution using expectation maximization, identifying "important" sample-feature pairs. Additionally, SR technology is used to emphasize the role of chemical bond information in message propagation. KAN technology is integrated to greatly enhance model interpretability. In summary, the model mitigates challenges related to uneven data distribution and scarcity. Multiple experiments on public datasets confirm the model's robust performance. We anticipate that this model will provide deeper insights into compound inhibition mechanisms on cardiac ion channels and reduce toxicity risks.
Xinyu Zhang 0012, Zhenya Du, Linlin Zhuo, Xiangzheng Fu, Dong-Sheng Cao 0001, Boqia Xie, Keqin Li 0001
IEEE J. Biomed. Health Informatics5
2024 Dual-Stream Heterogeneous Graph Neural Network Based on Zero-Shot Embeddings for Predicting miRNA-Drug Sensitivity
abstract
MicroRNAs (miRNAs) are a class of non-coding RNA molecules that have been shown to be closely associated with the sensitivity of chemotherapeutic drugs in cancer treatment. Given the high cost and extended duration of traditional biological experiments, there is an urgent need to develop computational models to predict the sensitivity scores between miRNAs and drugs. In this study, we proposed a dual-stream graph neural network method based on Zero-Shot Embeddings, named DSHGZS, to explore the potential sensitivity scores between miRNAs and drugs. DSHGZS first constructs two heterogeneous graphs with different isomorphic subgraphs based on zero-shot embeddings obtained from large language models (LLMs) and known miRNA-drug association data. It then utilized the enhanced LLM-derived node feature representations, embedding them into the layer feature learning process of the two heterogeneous graphs to generate high-quality vector representations of miRNAs and drugs. The learned high-quality feature embeddings are subsequently used in a segmented inner product decoder to evaluate the sensitivity association scores between miRNAs and drugs. To address the model’s excessive reliance on high-quality feature representations, we employed PCA to extract the core representations of the LLM-derived node features for data augmentation. Case studies demonstrated that DSHGZS is an effective tool for predicting potential sensitivity scores between miRNAs and drugs.
Wang Wang, Wenhui Xiao, Xiangzheng Fu
BIBM5
2024 A Novel Approach for Subtype Identification via Multi-omics Data Using Adversarial Autoencoder
Hao Nie, Quanwei Chen, Zixing He, Xiuxiu Chao, Weihao Ou, Xiangzheng Fu
ISBRA (1)8
2024 MS-BACL: enhancing metabolic stability prediction through bond graph augmentation and contrastive learning
abstract
MOTIVATION: Accurately predicting molecular metabolic stability is of great significance to drug research and development, ensuring drug safety and effectiveness. Existing deep learning methods, especially graph neural networks, can reveal the molecular structure of drugs and thus efficiently predict the metabolic stability of molecules. However, most of these methods focus on the message passing between adjacent atoms in the molecular graph, ignoring the relationship between bonds. This makes it difficult for these methods to estimate accurate molecular representations, thereby being limited in molecular metabolic stability prediction tasks. RESULTS: We propose the MS-BACL model based on bond graph augmentation technology and contrastive learning strategy, which can efficiently and reliably predict the metabolic stability of molecules. To our knowledge, this is the first time that bond-to-bond relationships in molecular graph structures have been considered in the task of metabolic stability prediction. We build a bond graph based on 'atom-bond-atom', and the model can simultaneously capture the information of atoms and bonds during the message propagation process. This enhances the model's ability to reveal the internal structure of the molecule, thereby improving the structural representation of the molecule. Furthermore, we perform contrastive learning training based on the molecular graph and its bond graph to learn the final molecular representation. Multiple sets of experimental results on public datasets show that the proposed MS-BACL model outperforms the state-of-the-art model. AVAILABILITY AND IMPLEMENTATION: The code and data are publicly available at https://github.com/taowang11/MS.
Zhen Li 0015, Linlin Zhuo, Xiangzheng Fu, Quan Zou 0001
Briefings Bioinform.5
2024 Diff-AMP: tailored designed antimicrobial peptide framework with all-in-one generation, identification, prediction and optimization
abstract
Antimicrobial peptides (AMPs), short peptides with diverse functions, effectively target and combat various organisms. The widespread misuse of chemical antibiotics has led to increasing microbial resistance. Due to their low drug resistance and toxicity, AMPs are considered promising substitutes for traditional antibiotics. While existing deep learning technology enhances AMP generation, it also presents certain challenges. Firstly, AMP generation overlooks the complex interdependencies among amino acids. Secondly, current models fail to integrate crucial tasks like screening, attribute prediction and iterative optimization. Consequently, we develop a integrated deep learning framework, Diff-AMP, that automates AMP generation, identification, attribute prediction and iterative optimization. We innovatively integrate kinetic diffusion and attention mechanisms into the reinforcement learning framework for efficient AMP generation. Additionally, our prediction module incorporates pre-training and transfer learning strategies for precise AMP identification and screening. We employ a convolutional neural network for multi-attribute prediction and a reinforcement learning-based iterative optimization strategy to produce diverse AMPs. This framework automates molecule generation, screening, attribute prediction and optimization, thereby advancing AMP research. We have also deployed Diff-AMP on a web server, with code, data and server details available in the Data Availability section.
Rui Wang 0168, Linlin Zhuo, Jinhang Wei, Xiangzheng Fu, Quan Zou 0001
Briefings Bioinform.5
2024 Joint deep autoencoder and subgraph augmentation for inferring microbial responses to drugs
abstract
Exploring microbial stress responses to drugs is crucial for the advancement of new therapeutic methods. While current artificial intelligence methodologies have expedited our understanding of potential microbial responses to drugs, the models are constrained by the imprecise representation of microbes and drugs. To this end, we combine deep autoencoder and subgraph augmentation technology for the first time to propose a model called JDASA-MRD, which can identify the potential indistinguishable responses of microbes to drugs. In the JDASA-MRD model, we begin by feeding the established similarity matrices of microbe and drug into the deep autoencoder, enabling to extract robust initial features of both microbes and drugs. Subsequently, we employ the MinHash and HyperLogLog algorithms to account intersections and cardinality data between microbe and drug subgraphs, thus deeply extracting the multi-hop neighborhood information of nodes. Finally, by integrating the initial node features with subgraph topological information, we leverage graph neural network technology to predict the microbes' responses to drugs, offering a more effective solution to the 'over-smoothing' challenge. Comparative analyses on multiple public datasets confirm that the JDASA-MRD model's performance surpasses that of current state-of-the-art models. This research aims to offer a more profound insight into the adaptability of microbes to drugs and to furnish pivotal guidance for drug treatment strategies. Our data and code are publicly available at: https://github.com/ZZCrazy00/JDASA-MRD.
Zhecheng Zhou, Linlin Zhuo, Xiangzheng Fu, Quan Zou 0001
Briefings Bioinform.3
2024 GraphADT: empowering interpretable predictions of acute dermal toxicity with multi-view graph pooling and structure remapping
abstract
MOTIVATION: Accurate prediction of acute dermal toxicity (ADT) is essential for the safe and effective development of contact drugs. Currently, graph neural networks, a form of deep learning technology, accurately model the structure of compound molecules, enhancing predictions of their ADT. However, many existing methods emphasize atom-level information transfer and overlook crucial data conveyed by molecular bonds and their interrelationships. Additionally, these methods often generate "equal" node representations across the entire graph, failing to accentuate "important" substructures like functional groups, pharmacophores, and toxicophores, thereby reducing interpretability. RESULTS: We introduce a novel model, GraphADT, utilizing structure remapping and multi-view graph pooling (MVPool) technologies to accurately predict compound ADT. Initially, our model applies structure remapping to better delineate bonds, transforming "bonds" into new nodes and "bond-atom-bond" interactions into new edges, thereby reconstructing the compound molecular graph. Subsequently, we use MVPool to amalgamate data from various perspectives, minimizing biases inherent to single-view analyses. Following this, the model generates a robust node ranking collaboratively, emphasizing critical nodes or substructures to enhance model interpretability. Lastly, we apply a graph comparison learning strategy to train both the original and structure remapped molecular graphs, deriving the final molecular representation. Experimental results on public datasets indicate that the GraphADT model outperforms existing state-of-the-art models. The GraphADT model has been demonstrated to effectively predict compound ADT, offering potential guidance for the development of contact drugs and related treatments. AVAILABILITY AND IMPLEMENTATION: Our code and data are accessible at: https://github.com/mxqmxqmxq/GraphADT.git.
Xinqian Ma, Xiangzheng Fu, Linlin Zhuo, Quan Zou 0001
Bioinform.2
2024 Revisiting drug-protein interaction prediction: a novel global-local perspective
abstract
MOTIVATION: Accurate inference of potential drug-protein interactions (DPIs) aids in understanding drug mechanisms and developing novel treatments. Existing deep learning models, however, struggle with accurate node representation in DPI prediction, limiting their performance. RESULTS: We propose a new computational framework that integrates global and local features of nodes in the drug-protein bipartite graph for efficient DPI inference. Initially, we employ pre-trained models to acquire fundamental knowledge of drugs and proteins and to determine their initial features. Subsequently, the MinHash and HyperLogLog algorithms are utilized to estimate the similarity and set cardinality between drug and protein subgraphs, serving as their local features. Then, an energy-constrained diffusion mechanism is integrated into the transformer architecture, capturing interdependencies between nodes in the drug-protein bipartite graph and extracting their global features. Finally, we fuse the local and global features of nodes and employ multilayer perceptrons to predict the likelihood of potential DPIs. A comprehensive and precise node representation guarantees efficient prediction of unknown DPIs by the model. Various experiments validate the accuracy and reliability of our model, with molecular docking results revealing its capability to identify potential DPIs not present in existing databases. This approach is expected to offer valuable insights for furthering drug repurposing and personalized medicine research. AVAILABILITY AND IMPLEMENTATION: Our code and data are accessible at: https://github.com/ZZCrazy00/DPI.
Zhecheng Zhou, Qingquan Liao, Jinhang Wei, Linlin Zhuo, Xiaonan Wu, Xiangzheng Fu, Quan Zou 0001
Bioinform.6
2024 Multi-source data integration for explainable miRNA-driven drug discovery
Zhen Li 0015, Qingquan Liao, Peng Xu 0004, Linlin Zhuo, Xiangzheng Fu, Quan Zou 0001
Future Gener. Comput. Syst.6
2024 Developing explainable models for lncRNA-Targeted drug discovery using graph autoencoders
Xiangzheng Fu, Haiting Chen, Jun Shang, Haoyu Zhou, Wang Zhe
Future Gener. Comput. Syst.2
2024 ECD-CDGI: An efficient energy-constrained diffusion model for cancer driver gene identification
abstract
The identification of cancer driver genes (CDGs) poses challenges due to the intricate interdependencies among genes and the influence of measurement errors and noise. We propose a novel energy-constrained diffusion (ECD)-based model for identifying CDGs, termed ECD-CDGI. This model is the first to design an ECD-Attention encoder by combining the ECD technique with an attention mechanism. ECD-Attention encoder excels at generating robust gene representations that reveal the complex interdependencies among genes while reducing the impact of data noise. We concatenate topological embedding extracted from gene-gene networks through graph transformers to these gene representations. We conduct extensive experiments across three testing scenarios. Extensive experiments show that the ECD-CDGI model possesses the ability to not only be proficient in identifying known CDGs but also efficiently uncover unknown potential CDGs. Furthermore, compared to the GNN-based approach, the ECD-CDGI model exhibits fewer constraints by existing gene-gene networks, thereby enhancing its capability to identify CDGs. Additionally, ECD-CDGI is open-source and freely available. We have also launched the model as a complimentary online tool specifically crafted to expedite research efforts focused on CDGs identification.
Linlin Zhuo, Xiangzheng Fu, Xiangxiang Zeng, Quan Zou 0001
PLoS Comput. Biol.4
2024 Dual-View Learning Based on Images and Sequences for Molecular Property Prediction
abstract
The prediction of molecular properties remains a challenging task in the field of drug design and development. Recently, there has been a growing interest in the analysis of biological images. Molecular images, as a novel representation, have proven to be competitive, yet they lack explicit information and detailed semantic richness. Conversely, semantic information in SMILES sequences is explicit but lacks spatial structural details. Therefore, in this study, we focus on and explore the relationship between these two types of representations, proposing a novel multimodal architecture named ISMol. ISMol relies on a cross-attention mechanism to extract information representations of molecules from both images and SMILES strings, thereby predicting molecular properties. Evaluation results on 14 small molecule ADMET datasets indicate that ISMol outperforms machine learning (ML) and deep learning (DL) models based on single-modal representations. In addition, we analyze our method through a large number of experiments to test the superiority, interpretability and generalizability of the method. In summary, ISMol offers a powerful deep learning toolbox for drug discovery in a variety of molecular properties.
Xiang Zhang 0008, Hongxin Xiang, Xixi Yang, Jingxin Dong 0002, Xiangzheng Fu, Xiangxiang Zeng, Keqin Li 0001
IEEE J. Biomed. Health Informatics5
2023 MPCLCDA: predicting circRNA-disease associations by using automatically selected meta-path and contrastive learning
abstract
Circular RNA (circRNA) is closely associated with human diseases. Accordingly, identifying the associations between human diseases and circRNA can help in disease prevention, diagnosis and treatment. Traditional methods are time consuming and laborious. Meanwhile, computational models can effectively predict potential circRNA-disease associations (CDAs), but are restricted by limited data, resulting in data with high dimension and imbalance. In this study, we propose a model based on automatically selected meta-path and contrastive learning, called the MPCLCDA model. First, the model constructs a new heterogeneous network based on circRNA similarity, disease similarity and known association, via automatically selected meta-path and obtains the low-dimensional fusion features of nodes via graph convolutional networks. Then, contrastive learning is used to optimize the fusion features further, and obtain the node features that make the distinction between positive and negative samples more evident. Finally, circRNA-disease scores are predicted through a multilayer perceptron. The proposed method is compared with advanced methods on four datasets. The average area under the receiver operating characteristic curve, area under the precision-recall curve and F1 score under 5-fold cross-validation reached 0.9752, 0.9831 and 0.9745, respectively. Simultaneously, case studies on human diseases further prove the predictive ability and application value of this method.
Wei Liu 0150, Ting Tang, Xu Lu 0002, Xiangzheng Fu
Briefings Bioinform.4
2023 NSRGRN: a network structure refinement method for gene regulatory network inference
abstract
The elucidation of gene regulatory networks (GRNs) is one of the central challenges of systems biology, which is crucial for understanding pathogenesis and curing diseases. Various computational methods have been developed for GRN inference, but identifying redundant regulation remains a fundamental problem. Although considering topological properties and edge importance measures simultaneously can identify and reduce redundant regulations, how to address their respective weaknesses whilst leveraging their strengths is a critical problem faced by researchers. Here, we propose a network structure refinement method for GRN (NSRGRN) that effectively combines the topological properties and edge importance measures during GRN inference. NSRGRN has two major parts. The first part constructs a preliminary ranking list of gene regulations to avoid starting the GRN inference from a directed complete graph. The second part develops a novel network structure refinement (NSR) algorithm to refine the network structure from local and global topology perspectives. Specifically, the Conditional Mutual Information with Directionality and network motifs are applied to optimise the local topology, and the lower and upper networks are used to balance the bilateral relationship between the local topology's optimisation and the global topology's maintenance. NSRGRN is compared with six state-of-the-art methods on three datasets (26 networks in total), and it shows the best all-round performance. Furthermore, when acting as a post-processing step, the NSR algorithm can improve the results of other methods in most datasets.
Wei Liu 0150, Xu Lu 0002, Xiangzheng Fu, Ruiqing Sun, Li Yang 0026
Briefings Bioinform.4
2023 GCFMCL: predicting miRNA-drug sensitivity using graph collaborative filtering and multi-view contrastive learning
abstract
Studies have shown that the mechanism of action of many drugs is related to miRNA. In-depth research on the relationship between miRNA and drugs can provide theoretical foundations and practical approaches for various areas, such as drug target discovery, drug repositioning and biomarker research. Traditional biological experiments to test miRNA-drug susceptibility are costly and time-consuming. Thus, sequence- or topology-based deep learning methods are recognized in this field for their efficiency and accuracy. However, these methods have limitations in dealing with sparse topologies and higher-order information of miRNA (drug) feature. In this work, we propose GCFMCL, a model for multi-view contrastive learning based on graph collaborative filtering. To the best of our knowledge, this is the first attempt that incorporates contrastive learning strategy into the graph collaborative filtering framework to predict the sensitivity relationships between miRNA and drug. The proposed multi-view contrastive learning method is divided into topological contrastive objective and feature contrastive objective: (1) For the homogeneous neighbors of the topological graph, we propose a novel topological contrastive learning method via constructing the contrastive target through the topological neighborhood information of nodes. (2) The proposed model obtains feature contrastive targets from high-order feature information according to the correlation of node features, and mines potential neighborhood relationships in the feature space. The proposed multi-view comparative learning effectively alleviates the impact of heterogeneous node noise and graph data sparsity in graph collaborative filtering, and significantly enhances the performance of the model. Our study employs a dataset derived from the NoncoRNA and ncDR databases, encompassing 2049 experimentally validated miRNA-drug sensitivity associations. Five-fold cross-validation shows that the Area Under the Curve (AUC), Area Under the Precision-Recall Curve (AUPR) and F1-score (F1) of GCFMCL reach 95.28%, 95.66% and 89.77%, which outperforms the state-of-the-art (SOTA) method by the margin of 2.73%, 3.42% and 4.96%, respectively. Our code and data can be accessed at https://github.com/kkkayle/GCFMCL.
Jinhang Wei, Linlin Zhuo, Zhecheng Zhou, Xinzhe Lian, Xiangzheng Fu
Briefings Bioinform.5
2022 NSCGRN: a network structure control method for gene regulatory network inference
abstract
Accurate inference of gene regulatory networks (GRNs) is an essential premise for understanding pathogenesis and curing diseases. Various computational methods have been developed for GRN inference, but the identification of redundant regulation remains a challenge faced by researchers. Although combining global and local topology can identify and reduce redundant regulations, the topologies' specific forms and cooperation modes are unclear and real regulations may be sacrificed. Here, we propose a network structure control method [network-structure-controlling-based GRN inference method (NSCGRN)] that stipulates the global and local topology's specific forms and cooperation mode. The method is carried out in a cooperative mode of 'global topology dominates and local topology refines'. Global topology requires layering and sparseness of the network, and local topology requires consistency of the subgraph association pattern with the network motifs (fan-in, fan-out, cascade and feedforward loop). Specifically, an ordered gene list is obtained by network topology centrality sorting. A Bernaola-Galvan mutation detection algorithm applied to the list gives the hierarchy of GRNs to control the upstream and downstream regulations within the global scope. Finally, four network motifs are integrated into the hierarchy to optimize local complex regulations and form a cooperative mode where global and local topologies play the dominant and refined roles, respectively. NSCGRN is compared with state-of-the-art methods on three different datasets (six networks in total), and it achieves the highest F1 and Matthews correlation coefficient. Experimental results show its unique advantages in GRN inference.
Wei Liu 0150, Xingen Sun, Li Yang 0026, Xiangzheng Fu
Briefings Bioinform.6
2022 DAESTB: inferring associations of small molecule-miRNA via a scalable tree boosting model based on deep autoencoder
abstract
MicroRNAs (miRNAs) are closely related to a variety of human diseases, not only regulating gene expression, but also having an important role in human life activities and being viable targets of small molecule drugs for disease treatment. Current computational techniques to predict the potential associations between small molecule and miRNA are not that accurate. Here, we proposed a new computational method based on a deep autoencoder and a scalable tree boosting model (DAESTB), to predict associations between small molecule and miRNA. First, we constructed a high-dimensional feature matrix by integrating small molecule-small molecule similarity, miRNA-miRNA similarity and known small molecule-miRNA associations. Second, we reduced feature dimensionality on the integrated matrix using a deep autoencoder to obtain the potential feature representation of each small molecule-miRNA pair. Finally, a scalable tree boosting model is used to predict small molecule and miRNA potential associations. The experiments on two datasets demonstrated the superiority of DAESTB over various state-of-the-art methods. DAESTB achieved the best AUC value. Furthermore, in three case studies, a large number of predicted associations by DAESTB are confirmed with the public accessed literature. We envision that DAESTB could serve as a useful biological model for predicting potential small molecule-miRNA associations.
Yuan Tu, Xiangzheng Fu, Xiang Chen 0029
Briefings Bioinform.5
2022 RNMFLP: Predicting circRNA-disease associations based on robust nonnegative matrix factorization and label propagation
abstract
Circular RNAs (circRNAs) are a class of structurally stable endogenous noncoding RNA molecules. Increasing studies indicate that circRNAs play vital roles in human diseases. However, validating disease-related circRNAs in vivo is costly and time-consuming. A reliable and effective computational method to identify circRNA-disease associations deserves further studies. In this study, we propose a computational method called RNMFLP that combines robust nonnegative matrix factorization (RNMF) and label propagation algorithm (LP) to predict circRNA-disease associations. First, to reduce the impact of false negative data, the original circRNA-disease adjacency matrix is updated by matrix multiplication using the integrated circRNA similarity and the disease similarity information. Subsequently, the RNMF algorithm is used to obtain the restricted latent space to capture potential circRNA-disease pairs from the association matrix. Finally, the LP algorithm is utilized to predict more accurate circRNA-disease associations from the integrated circRNA similarity network and integrated disease similarity network, respectively. Fivefold cross-validation of four datasets shows that RNMFLP is superior to the state-of-the-art methods. In addition, case studies on lung cancer, hepatocellular carcinoma and colorectal cancer further demonstrate the reliability of our method to discover disease-related circRNAs.
Xiang Chen 0029, Xiangzheng Fu, Wei Liu 0150
Briefings Bioinform.5
2022 Predicting ncRNA-protein interactions based on dual graph convolutional network and pairwise learning
abstract
Noncoding RNAs (ncRNAs) have recently attracted considerable attention due to their key roles in biology. The ncRNA-proteins interaction (NPI) is often explored to reveal some biological activities that ncRNA may affect, such as biological traits, diseases, etc. Traditional experimental methods can accomplish this work but are often labor-intensive and expensive. Machine learning and deep learning methods have achieved great success by exploiting sufficient sequence or structure information. Graph Neural Network (GNN)-based methods consider the topology in ncRNA-protein graphs and perform well on tasks like NPI prediction. Based on GNN, some pairwise constraint methods have been developed to apply on homogeneous networks, but not used for NPI prediction on heterogeneous networks. In this paper, we construct a pairwise constrained NPI predictor based on dual Graph Convolutional Network (GCN) called NPI-DGCN. To our knowledge, our method is the first to train a heterogeneous graph-based model using a pairwise learning strategy. Instead of binary classification, we use a rank layer to calculate the score of an ncRNA-protein pair. Moreover, our model is the first to predict NPIs on the ncRNA-protein bipartite graph rather than the homogeneous graph. We transform the original ncRNA-protein bipartite graph into two homogenous graphs on which to explore second-order implicit relationships. At the same time, we model direct interactions between two homogenous graphs to explore explicit relationships. Experimental results on the four standard datasets indicate that our method achieves competitive performance with other state-of-the-art methods. And the model is available at https://github.com/zhuoninnin1992/NPIPredict.
Linlin Zhuo, Bosheng Song, Yuansheng Liu, Xiangzheng Fu
Briefings Bioinform.5
2022 A dynamic population reduction differential evolution algorithm combining linear and nonlinear strategy piecewise functions
abstract
Abstract The population size has a great impact on the performance of the Differential Evolution (DE) algorithm, but in many classic DE algorithms, the population size is usually determined by the user based on the experience value, and it remains unchanged during the evolution process, which greatly affect the performance of DE. To this end, a dynamic population reduction differential evolution algorithm (DPSHADE) that combines linear and nonlinear strategy piecewise functions is proposed. The algorithm uses a dynamic population size reduction method to dynamically adjust the population size during operation, and construct a combination of linear and nonlinear piecewise functions for dynamic scale adaptive adjustment. In this paper, we proposed the DPSHADE algorithm and is compared with the four traditional algorithms in the CEC2017 benchmark set. The experimental results show that DPSHADE performs better in overall performance, which is significant better than the performance of SHADE.
Kangshun Li, Xiangzheng Fu, Hassan Jalil
Concurr. Comput. Pract. Exp.2
2021 Nonnegative Matrix Factorization Framework for Disease-Related CircRNA Prediction
Wei Liu 0150, Xiangzheng Fu
ICA3PP (3)4
2021 Drug repositioning based on the heterogeneous information fusion graph convolutional network
abstract
In silico reuse of old drugs (also known as drug repositioning) to treat common and rare diseases is increasingly becoming an attractive proposition because it involves the use of de-risked drugs, with potentially lower overall development costs and shorter development timelines. Therefore, there is a pressing need for computational drug repurposing methodologies to facilitate drug discovery. In this study, we propose a new method, called DRHGCN (Drug Repositioning based on the Heterogeneous information fusion Graph Convolutional Network), to discover potential drugs for a certain disease. To make full use of different topology information in different domains (i.e. drug-drug similarity, disease-disease similarity and drug-disease association networks), we first design inter- and intra-domain feature extraction modules by applying graph convolution operations to the networks to learn the embedding of drugs and diseases, instead of simply integrating the three networks into a heterogeneous network. Afterwards, we parallelly fuse the inter- and intra-domain embeddings to obtain the more representative embeddings of drug and disease. Lastly, we introduce a layer attention mechanism to combine embeddings from multiple graph convolution layers for further improving the prediction performance. We find that DRHGCN achieves high performance (the average AUROC is 0.934 and the average AUPR is 0.539) in four benchmark datasets, outperforming the current approaches. Importantly, we conducted molecular docking experiments on DRHGCN-predicted candidate drugs, providing several novel approved drugs for Alzheimer's disease (e.g. benzatropine) and Parkinson's disease (e.g. trihexyphenidyl and haloperidol).
Changcheng Lu, Junlin Xu, Yajie Meng, Peng Wang 0035, Xiangzheng Fu, Xiangxiang Zeng, Yansen Su
Briefings Bioinform.6
2021 ITP-Pred: an interpretable method for predicting, therapeutic peptides with fused features low-dimension representation
abstract
The peptide therapeutics market is providing new opportunities for the biotechnology and pharmaceutical industries. Therefore, identifying therapeutic peptides and exploring their properties are important. Although several studies have proposed different machine learning methods to predict peptides as being therapeutic peptides, most do not explain the decision factors of model in detail. In this work, an Interpretable Therapeutic Peptide Prediction (ITP-Pred) model based on efficient feature fusion was developed. First, we proposed three kinds of feature descriptors based on sequence and physicochemical property encoded, namely amino acid composition (AAC), group AAC and coding autocorrelation, and concatenated them to obtain the feature representation of therapeutic peptide. Then, we input it into the CNN-Bi-directional Long Short-Term Memory (BiLSTM) model to automatically learn recognition of therapeutic peptides. The cross-validation and independent verification experiments results indicated that ITP-Pred has a higher prediction performance on the benchmark dataset than other comparison methods. Finally, we analyzed the output of the model from two aspects: sequence order and physical and chemical properties, mining important features as guidance for the design of better models that can complement existing methods.
Li Wang 0145, Xiangzheng Fu, Chenxing Xia, Xiangxiang Zeng, Quan Zou 0001
Briefings Bioinform.3
2021 iEnhancer-XG: interpretable sequence-based enhancers and their strength predictor
abstract
MOTIVATION: Enhancers are non-coding DNA fragments with high position variability and free scattering. They play an important role in controlling gene expression. As machine learning has become more widely used in identifying enhancers, a number of bioinformatic tools have been developed. Although several models for identifying enhancers and their strengths have been proposed, their accuracy and efficiency have yet to be improved. RESULTS: We propose a two-layer predictor called 'iEnhancer-XG.' It comprises a one-layer predictor (for identifying enhancers) and a second classifier (for their strength) and uses 'XGBoost' as a base classifier and five feature extraction methods, namely, k-Spectrum Profile, Mismatch k-tuple, Subsequence Profile, Position-specific scoring matrix (PSSM) and Pseudo dinucleotide composition (PseDNC). Each method has an independent output. We place the feature vector matrix into the ensemble learning for fusion. This experiment involves the method of 'SHapley Additive explanations' to provide interpretability for the previous black box machine learning methods and improve their credibility. The accuracies of the ensemble learning method are 0.811 (first layer) and 0.657 (second layer). The rigorous 10-fold cross-validation confirms that the proposed method is significantly better than existing technologies. AVAILABILITY AND IMPLEMENTATION: The source code and dataset for the enhancer predictions have been uploaded to https://github.com/jimmyrate/ienhancer-xg. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xuanbai Ren, Xiangzheng Fu, Mingyu Gao 0004, Xiangxiang Zeng
Bioinform.3
2020 StackCPPred: a stacking and pairwise energy content-based prediction of cell-penetrating peptides and their uptake efficiency
abstract
MOTIVATION: Cell-penetrating peptides (CPPs) are a vehicle for transporting into living cells pharmacologically active molecules, such as short interfering RNAs, nanoparticles, plasmid DNAs and small peptides, thus offering great potential as future therapeutics. Existing experimental techniques for identifying CPPs are time-consuming and expensive. Thus, the prediction of CPPs from peptide sequences by using computational methods can be useful to annotate and guide the experimental process quickly. Many machine learning-based methods have recently emerged for identifying CPPs. Although considerable progress has been made, existing methods still have low feature representation capabilities, thereby limiting further performance improvements. RESULTS: We propose a method called StackCPPred, which proposes three feature methods on the basis of the pairwise energy content of the residue as follows: RECM-composition, PseRECM and RECM-DWT. These features are used to train stacking-based machine learning methods to effectively predict CPPs. On the basis of the CPP924 and CPPsite3 datasets with jackknife validation, StackDPPred achieved 94.5% and 78.3% accuracy, which was 2.9% and 5.8% higher than the state-of-the-art CPP predictors, respectively. StackCPPred can be a powerful tool for predicting CPPs and their uptake efficiency, facilitating hypothesis-driven experimental design and accelerating their applications in clinical therapy. AVAILABILITY AND IMPLEMENTATION: Source code and data can be downloaded from https://github.com/Excelsior511/StackCPPred. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiangzheng Fu, Xiangxiang Zeng, Quan Zou 0001
Bioinform.1