Xiangxiang Zeng

dblp:20/3839 · DBLP profile ↗
← Back
174ranked-venue papers
18as first author
117since 2021 · last 2026
0000-0003-1081-7658ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 102 · 12 first-author · 70 since 2021Artificial intelligence and machine learning · 51 · 4 first-author · 34 since 2021Theory of computation · 12 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 10 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2026 TRACE: Transformation-Aware Graph Refinement for Reaction Condition Prediction
abstract
Identifying suitable reaction conditions is critical for chemical synthesis, as they directly affect yield, selectivity, and transformation feasibility. While recent methods have shown promising results, most approaches either encode reactants and products independently or rely on rule-based reaction graphs, both of which constrain the ability of the model to capture condition-relevant structural transformations. In this work, we propose TRACE, a transformation-aware graph refinement framework for reaction condition prediction. TRACE constructs atom-level joint graphs that integrate both reactant and product structures to represent condition-relevant transformations. A structure-aware encoder enriches atom features with local chemical context, followed by a dynamic interaction refinement module that adaptively infers task-specific edges. To further guide the model toward condition-relevant patterns, a mechanism regularized graph encoder incorporates reaction center information, enabling more accurate modeling of transformation mechanisms. Experiments on benchmark datasets show that TRACE achieves state-of-the-art performance across multiple condition types. The integration of transformation-aware refinement leads to improvements in prediction accuracy and generalization, while maintaining robust performance in challenging and realistic synthesis planning scenarios.
Yujie Chen 0002, Tengfei Ma 0002, Yuansheng Liu, Leyi Wei, Dong-Sheng Cao 0001, Xiangxiang Zeng
AAAI8
2026 Expert-Inspired Multi-Agent Coordination for Multi-Objective Molecular Optimization
abstract
Multi-objective molecular optimization is a fundamental yet inherently challenging task in drug discovery, as it requires simultaneously optimizing multiple, often conflicting, molecular properties. Although recent deep learning methods have shown promise, they often lack objective-specific specialization and dynamic coordination, making them ineffective in handling competing objectives and difficult to scale in complex, high-dimensional molecular design tasks. Inspired by the division of labor among domain experts in medicinal chemistry, we propose MAMO, a multi-agent framework for molecular design that simulates expert collaboration. Each agent specializes in optimizing a single objective, and their interactions are orchestrated by a central scheduling module that dynamically reallocates tasks based on evaluation feedback. This coordination mechanism enables interpretable and goal-conditioned optimization while adaptively balancing conflicting objectives. Extensive experiments on benchmark datasets demonstrate that MAMO consistently achieves superior performance in both objective quality and Pareto diversity, particularly in scenarios with strong inter-objective conflict. Our results highlight the potential of multi-agent coordination strategies for scalable and conflict-aware molecular design.
Daojian Zeng, Tianle Li, Jiacai Yi, Lincheng Jiang, Tengfei Ma 0002, Xiangxiang Zeng
AAAI8
2026 CAML: A Conflict-Aware Molecular Language Model Merging Framework for Multi-Constraint Molecular Generation
abstract
Xuanbai Ren, Luoda Tan, Pei Liu, Tengfei Ma, Xiangzheng Fu, Longyue Wang, Yiping Liu, Xiangxiang Zeng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xuanbai Ren, Luoda Tan, Pei Liu 0008, Tengfei Ma 0002, Xiangzheng Fu, Longyue Wang, Xiangxiang Zeng
ACL (1)8
2026 M²PO: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation
abstract
Hao Wang, Linlong Xu, Heng Liu, Yangyang Liu, Xiaohu Zhao, Bo Zeng, Liangying Shao, Yichen Dong, Xinwei Wu, Jiang Zhou, Tianyu Dong, Xiangxiang Zeng, Longyue Wang, Weihua Luo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Linlong Xu, Liangying Shao, Yichen Dong, Xinwei Wu 0001, Tianyu Dong, Xiangxiang Zeng, Longyue Wang, Weihua Luo
ACL (1)12
2026 A deep adversarial network model for multi-task analysis of single-cell omics data
abstract
Single-cell multi-omics data reveal complex cellular states and deepen our understanding of tissue cell phenotypes and functions. However, data analysis remains challenging due to the discrete nature and high noise level of the data, as well as the lack of modality. Here, we propose scMultiNet, a multi-task deep adversarial neural network that can integrate different tasks to analyze single-cell multi-modal data. In particular, we achieve joint training of multi-modal integration and cross-modal prediction tasks by introducing a cross-modal bi-prediction module and a multi-head self-attention module. Data denoising is further enhanced by integrating an indicator matrix that constrains and precisely reconstructs the original expression values. Extensive simulations and real data experiments demonstrate that scMultiNet outperforms existing state-of-the-art methods in dimensionality reduction, visualization, clustering, batch elimination, data denoising, multi-modal integration, single-cell cross-modality translation, and in revealing cell type-specific biological insights. In addition, we demonstrate that scMultiNet can effectively transfer the complex relationships between modalities from one batch to another. In summary, scMultiNet stands as a comprehensive end-to-end framework, ideally suited for analyzing single-cell multi-omics data.
Junlin Xu, Yajie Meng, Shuting Jin, Changcheng Lu, Feifei Cui, Xiangzheng Fu, Quan Zou 0001, Xiangxiang Zeng
Briefings Bioinform.11
2026 Learning drug synergy through environment-conditioned feature modulation
abstract
MOTIVATION: Drug combinations are crucial for overcoming resistance in cancer therapy. Although deep learning has achieved strong performance in synergy prediction, existing models often treat cell-specific features and paired drugs as a static background and fail to capture how the specific cell-drug environment dynamically modulates drug representations, thereby hindering the modeling of environment-specific synergistic effects. RESULTS: We propose Env-Syn, a framework for modeling drug-drug-cell interactions through Environment-Conditioned Feature Modulation, which incorporates a Residual Feature-wise Linear Modulation (R-FiLM) module to perform precise affine transformations on drug representations conditioned on paired drugs and cellular environments. Benchmark evaluations show that Env-Syn consistently outperforms state-of-the-art methods. Notably, the model exhibits exceptional generalization performance in rigorous inductive scenarios. It maintains high predictive accuracy for unseen drugs with AUROC and AUPRC exceeding 0.81 in the Leave-drug-out setting and further demonstrates strong cross-dataset reliability by surpassing a recall of 0.7 on independent test set. Furthermore, among 15 novel predicted drug combinations, 8 are directly supported by literature evidence. These results demonstrate that Env-Syn is an effective computational tool for drug synergy discovery. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/AnQi-87/Env-Syn.
Shuting Jin, Yajie Meng, Zhonghang Zhu, Yinghui Jiang, Junlin Xu, Xiangxiang Zeng
Bioinform.7
2026 FKSUDDAPre: A drug-disease association prediction framework based on F-TEST feature selection and AMDKSU resampling with interpretability analysis
abstract
In drug discovery and therapeutic research, the prediction of drug-disease associations (DDAs) holds significant scientific and clinical value. Drug molecules exert their effects by precisely identifying disease-related biological targets, systematically modulating the entire pharmacological process from absorption, distribution, and metabolism to final efficacy. Accurate prediction of drug-disease associations not only facilitates an in-depth understanding of molecular mechanisms of drug action but also provides critical theoretical foundations for drug repositioning and personalized medicine. While traditional prediction methods based on in vitro experiments and clinical statistics yield reliable results, they suffer from inherent drawbacks such as long development cycles, substantial resource consumption, and low throughput. In contrast, emerging machine learning techniques offer a promising solution to these bottlenecks, enabling the intelligent and efficient discovery of potential drug-disease association networks and significantly improving drug development efficiency. However, it is noteworthy that existing machine learning methods still face significant challenges in practical applications: the complexity of feature construction raises the threshold for data processing; data sparsity constrains the depth of information mining; and the pervasive issue of sample imbalance poses a severe challenge to the model's predictive accuracy and generalization performance. In this study, we developed an efficient and accurate framework for drug-disease association prediction named FKSUDDAPre. The model employs a multi-modal feature fusion strategy: on one hand, it leverages an ensemble of Mol2vec and K- BERT to deeply capture the semantic features of drug molecular fingerprints; on the other hand, it integrates Medical Subject Headings (MeSH) with DeepWalk to effectively reduce the dimensionality of disease features while preserving their relational structure. To address the class imbalance problem, FKSUDDAPre designed an optimization algorithm called AMDKSU, which combined clustering with an improved distance metric strategy, significantly enhancing the discriminative power of the sample set. For data processing, F-test was employed for feature importance ranking, effectively reducing data dimensionality and improving model generalization. For the predictive architecture, FKSUDDAPre proposed a novel ensemble framework composed of XGBoost, Decision Tree, Random Forest, and HyperFast. By employing a dynamic weight allocation strategy, this ensemble effectively harnesses the complementary strengths of these models to achieve significantly enhanced predictive performance. Rigorous validation demonstrated the system's outstanding performance across multiple evaluation metrics, with an average AUC of 0.9725, improving the AUC by approximately 3.88% compared to the best-performing baseline model. In the prediction of Alzheimer's disease and Parkinson's disease, 80% and 60% of the top 10 candidate drugs recommended by FKSUDDAPre, respectively, had been confirmed by literature, demonstrating the model's good practical application potential. Furthermore, we conducted a LIME-based feature importance analysis on the model's predictions, visualizing the correlations between features and the target variable to demonstrate the model's interpretability. A cross-platform, user-friendly visualization tool had also been developed using the PyQt5 framework.
Yun Zuo 0001, Ge Hua, Xiangrong Liu, Xiangxiang Zeng, Zhaohong Deng
PLoS Comput. Biol.6
2026 BloodPatrol: Revolutionizing Blood Cancer Diagnosis - Advanced Real-Time Detection Leveraging Deep Learning & Cloud Technologies
abstract
Cloud computing and Internet of Things (IoT) technologies are gradually becoming the technological changemakers in cancer diagnosis. Blood cancer is an aggressive disease affecting the blood, bone marrow, and lymphatic system, and its early detection is crucial for subsequent treatment. Flow cytometry has been widely studied as a commonly used method for detecting blood cancer. However, the high computation and resource consumption severely limit its practical application, especifically in regions with limited medical and computational resources. In this study, with the help of cloud computing and IoT technologies, we develop a novel blood cancer dynamic monitoring diagnostic model named BloodPatrol based on an intelligent feature weight fusion mechanism. The proposed model is capable of capturing the dual-view importance relationship between cell samples and features, greatly improving prediction accuracy and significantly surpassing previous models. Besides, benefiting from the powerful processing ability of cloud computing, BloodPatrol can run on a distributed network to efficiently process large-scale cell data, which provides immediate and scalable blood cancer diagnostic services.
Jinhang Wei, Longyue Wang, Zhecheng Zhou, Linlin Zhuo, Xiangxiang Zeng, Xiangzheng Fu, Quan Zou 0001, Keqin Li 0001, Zhongjun Zhou
IEEE J. Biomed. Health Informatics5
2026 Enhanced Protein Network Representation With Explicit Structural Binding for Protein-Protein Interaction Prediction
abstract
Protein-protein interactions (PPIs) are fundamental molecular events in the human body, playing a pivotal role in disease treatment and intervention. However, existing approaches for protein representation often rely on simplistic PPI network models, which face two key challenges: (i) neglecting explicit residue-based binding relationships critical to protein interactions, and (ii) failing to integrate residue-level binding data with protein interaction networks, limiting their ability to uncover the binding mechanisms of PPIs. To address these issues, we propose an Enhanced protein network representation framework with Explicit structural binding information for improved PPI prediction, named E$^{2}$PPI. Specifically, E$^{2}$PPI extracts residue-level interactions between paired proteins using both single-protein structural analysis and inter-protein binding representation modules. To seamlessly integrate residue-level binding data with the semantics of protein interactions, we introduce an enhanced protein network representation module. This design enables the model to capture the interaction and binding mechanisms of PPIs, thereby improving its classification performance. Benchmark experiments demonstrate that E$^{2}$PPI outperforms state-of-the-art models, especially for few-shot and novel proteins, showcasing its superior generalization capabilities.
Zhuowen Zhen, Tengfei Ma 0002, Xiangxiang Zeng
IEEE J. Biomed. Health Informatics4
2025 S²DN: Learning to Denoise Unconvincing Knowledge for Inductive Knowledge Graph Completion
abstract
Inductive Knowledge Graph Completion (KGC) aims to infer missing facts between newly emerged entities within knowledge graphs (KGs), posing a significant challenge. While recent studies have shown promising results in inferring such entities through knowledge subgraph reasoning, they suffer from (i) the semantic inconsistencies of similar relations, and (ii) noisy interactions inherent in KGs due to the presence of unconvincing knowledge for emerging entities. To address these challenges, we propose a Semantic Structure-aware Denoising Network (S2DN) for inductive KGC. Our goal is to learn adaptable general semantics and reliable structures to distill consistent semantic knowledge while preserving reliable interactions within KGs. Specifically, we introduce a semantic smoothing module over the enclosing subgraphs to retain the universal semantic knowledge of relations. We incorporate a structure refining module to filter out unreliable interactions and offer additional knowledge, retaining robust structure surrounding target links. Extensive experiments conducted on three benchmark KGs demonstrate that S2DN surpasses the performance of state-of-the-art models. These results demonstrate the effectiveness of S2DN in preserving semantic consistency and enhancing the robustness of filtering out unreliable interactions in contaminated KGs.
Tengfei Ma 0002, Yujie Chen 0002, Xuan Lin, Bosheng Song, Xiangxiang Zeng
AAAI6
2025 Multi-Objective Molecular Design Through Learning Latent Pareto Set
abstract
Molecular design inherently involves the optimization of multiple conflicting objectives, such as enhancing bio-activity and ensuring synthesizability. Evaluating these objectives often requires resource-intensive computations or physical experiments. Current molecular design methodologies typically approximate the Pareto set using a limited number of molecules. In this paper, we present an innovative approach, called Multi-Objective Molecular Design through Learning Latent Pareto Set (MLPS). MLPS initially utilizes an encoder-decoder model to seamlessly transform the discrete chemical space into a continuous latent space. We then employ local Bayesian optimization models to efficiently search for local optimal solutions (i.e., molecules) within predefined trust regions. Using surrogate objective values derived from these local models, we train a global Pareto set learning model to understand the mapping between direction vectors (called “preferences”) in the objective space and the entire Pareto set in the continuous latent space. Both the global Pareto set learning model and local Bayesian optimization models collaborate to discover high-quality solutions and adapt the trust regions dynamically. Our work is an effective endeavor towards learning the Pareto set for multi-objective molecular design, providing decision-makers with the capability to fine-tune their preferences and thoroughly explore the Pareto set. Experimental results demonstrate that MLPS achieves state-of-the-art performance across various multi-objective scenarios, encompassing diverse objective types and varying numbers of objectives. The effectiveness of MLPS was further validated through real-world challenges in discovering antifungal peptides with low toxicity and high activity.
Xuanbai Ren, Yuansheng Liu, Bosheng Song, Xiangxiang Zeng, Hisao Ishibuchi
AAAI7
2025 Large Language and Protein Assistant for Protein-Protein Interactions Prediction
abstract
Peng Zhou, Pengsen Ma, Jianmin Wang, Xibao Cai, Haitao Huang, Wei Liu, Longyue Wang, Lai Hou Tim, Xiangxiang Zeng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Peng Zhou 0011, Pengsen Ma, Jianmin Wang 0016, Xibao Cai, Wei Liu 0005, Longyue Wang, Lai Hou Tim, Xiangxiang Zeng
ACL (1)9
2025 VQH-AM: Vector Quantization-Augmented Heterogeneous Graph Learning for Antibody Affinity Maturation Prediction
abstract
Modeling the impact of amino acid mutations on the binding affinity between antibodies and antigens plays a pivotal role in antibody therapeutics development. While numerous deep learning-based methods have shown promising results, accurately modeling the impact of mutations on binding interfaces remains a significant challenge, as even small changes can substantially affect binding affinity. To address this issue, we propose VQH-AM, a Vector Quantization-based Heterogeneous graph neural network for antibody Affinity Maturation prediction. Specifically, we construct atom-level heterogeneous graphs for both wild-type and mutant antibody-antigen complexes to model their binding interfaces. Atomic representations extracted by the heterogeneous graph neural networks are then quantized via a context-aware codebook designed to capture semantic atomic patterns. Finally, both the original and quantized atomic representations are integrated and used to predict the affinity maturation outcomes. Benchmark experiments demonstrate that VQH-AM consistently outperforms state-of-the-art methods across multiple evaluation metrics. The source code is available at: https://github.com/zwzhen-hnu/VQH-AM.
Zhuowen Zhen, Tengfei Ma 0002, Jiaxuan Li 0003, Dashun Zheng, Patrick Pang 0001, Xiangxiang Zeng
BIBM7
2025 Towards Synergistic Path-based Explanations for Knowledge Graph Completion: Exploration and Evaluation
abstract
Knowledge graph completion (KGC) aims to alleviate the inherent incompleteness of knowledge graphs (KGs), a crucial task for numerous applications such as recommendation systems and drug repurposing. The success of knowledge graph embedding (KGE) models provokes the question about the explainability: ``\textit{Which the patterns of the input KG are most determinant to the prediction}?'' Particularly, path-based explainers prevail in existing methods because of their strong capability for human understanding. In this paper, based on the observation that a fact is usually determined by the synergy of multiple reasoning chains, we propose a novel explainable framework, dubbed KGExplainer, to explore synergistic pathways. KGExplainer is a model-agnostic approach that employs a perturbation-based greedy search algorithm to identify the most crucial synergistic paths as explanations within the local structure of target predictions. To evaluate the quality of these explanations, KGExplainer distills an evaluator from the target KGE model, allowing for the examination of their fidelity. We experimentally demonstrate that the distilled evaluator has comparable predictive performance to the target KGE. Experimental results on benchmark datasets demonstrate the effectiveness of KGExplainer, achieving a human evaluation accuracy of 83.3\% and showing promising improvements in explainability. Code is available at \url{https://github.com/xiaomingaaa/KGExplainer}
Tengfei Ma 0002, Xiang Song 0003, Wen Tao, Mufei Li, Jiani Zhang 0003, Xiaoqin Pan, Yijun Wang 0002, Bosheng Song, Xiangxiang Zeng
ICLR9
2025 An All-Atom Generative Model for Designing Protein Complexes
abstract
Proteins typically exist in complexes, interacting with other proteins or biomolecules to perform their specific biological roles. Research on single-chain protein modeling has been extensively and deeply explored, with advancements seen in models like the series of ESM and AlphaFold2. Despite these developments, the study and modeling of multi-chain proteins remain largely uncharted, though they are vital for understanding biological functions. Recognizing the importance of these interactions, we introduce APM (all-Atom Protein generative Model), a model specifically designed for modeling multi-chain proteins. By integrating atom-level information and leveraging data on multi-chain proteins, APM is capable of precisely modeling inter-chain interactions and designing protein complexes with binding capabilities from scratch. It also performs folding and inverse-folding tasks for multi-chain proteins. Moreover, APM demonstrates versatility in downstream applications: it achieves enhanced performance through supervised fine-tuning (SFT) while also supporting zero-shot sampling in certain tasks, achieving state-of-the-art results. We released our code at https://github.com/bytedance/apm.
Ruizhe Chen, Dongyu Xue, Xiangxin Zhou, Zaixiang Zheng, Xiangxiang Zeng, Quanquan Gu
ICML5
2025 Enhancing Chemical Reaction and Retrosynthesis Prediction with Large Language Model and Dual-task Learning
abstract
Chemical reaction and retrosynthesis prediction are fundamental tasks in drug discovery. Recently, large language models (LLMs) have shown potential in many domains. However, directly applying LLMs to these tasks faces two major challenges: (i) lacking a large-scale chemical synthesis-related instruction dataset; (ii) ignoring the close correlation between reaction and retrosynthesis prediction for the existing fine-tuning strategies. To address these challenges, we propose ChemDual, a novel LLM framework for accurate chemical synthesis. Specifically, considering the high cost of data acquisition for reaction and retrosynthesis, ChemDual regards the reaction-and-retrosynthesis of molecules as a related recombination-and-fragmentation process and constructs a large-scale of 4.4 million instruction dataset. Furthermore, ChemDual introduces an enhanced LLaMA, equipped with a multi-scale tokenizer and dual-task learning strategy, to jointly optimize the process of recombination and fragmentation as well as the tasks between reaction and retrosynthesis prediction. Extensive experiments on Mol-Instruction and USPTO-50K datasets demonstrate that ChemDual achieves state-of-the-art performance in both predictions of reaction and retrosynthesis, outperforming the existing conventional single-task approaches and the general open-source LLMs. Through molecular docking analysis, ChemDual generates compounds with diverse and strong protein binding affinity, further highlighting its strong potential in drug design.
Xuan Lin, Qingrui Liu, Hongxin Xiang, Daojian Zeng, Xiangxiang Zeng
IJCAI5
2025 Electron Density-enhanced Molecular Geometry Learning
abstract
Electron density (ED), which describes the probability distribution of electrons in space, is crucial for accurately understanding the energy and force distribution in molecular force fields (MFF). Existing machine learning force fields (MLFF) focus on mining appropriate physical quantities from the atom-level conformation to enhance the molecular geometry representation while ignoring the unique information from microscopic electrons. In this work, we propose an efficient Electronic Density representation framework to enhance molecular Geometric learning (called EDG), which leverages images rendered from ED to boost molecular geometric representations in MLFF. Specifically, we construct a novel image-based ED representation, which consists of 2 million 6-view images with RGB-D channels, and design an ED representation learning model, called ImageED, to learn ED-related knowledge from these images. We further propose an efficient ED-aware teacher and introduce a cross-modal distillation strategy to transfer knowledge from the image-based teacher to the geometry-based students. Extensive experiments on QM9 and rMD17 demonstrate that EDG can be directly integrated into existing geometry-based models and significantly improves the capabilities of these models (e.g., SchNet, EGNN, SphereNet, ViSNet) for geometry representation learning in MLFF with a maximum average performance increase of 33.7%. Code and appendix are available at https://github.com/HongxinXiang/EDG
Hongxin Xiang, Jun Xia 0001, Xin Jin 0014, Wenjie Du 0003, Xiangxiang Zeng
IJCAI6
2025 Surface-based Molecular Design with Multi-modal Flow Matching
abstract
Therapeutic peptides show promise in targeting previously undruggable binding sites, with recent advancements in deep generative models enabling full-atom peptide co-design for specific protein receptors.However, the critical role of molecular surfaces in proteinprotein interactions (PPIs) has been underexplored.To bridge this gap, we propose an omni-design peptides generation paradigm, called SurfFlow, a novel surface-based generative algorithm that enables comprehensive co-design of sequence, structure, and surface for peptides.SurfFlow employs a multi-modality conditional flow matching (CFM) architecture to learn distributions of surface geometries and biochemical properties, enhancing peptide binding accuracy.Evaluated on the comprehensive PepMerge benchmark, SurfFlow consistently outperforms full-atom baselines across all metrics.These results highlight the advantages of considering molecular surfaces in de novo peptide discovery and demonstrate the potential of integrating multiple protein modalities for more effective therapeutic peptide discovery.
Fang Wu 0002, Zhengyuan Zhou, Shuting Jin, Xiangxiang Zeng, Jure Leskovec, Jinbo Xu
KDD (2)4
2025 Self-supervised Blending Structural Context of Visual Molecules for Robust Drug Interaction Prediction
abstract
Identifying drug-drug interactions (DDIs) is critical for ensuring drug safety and advancing drug development, a topic that has garnered significant research interest. While existing methods have made considerable progress, approaches relying solely on known DDIs face a key challenge when applied to drugs with limited data: insufficient exploration of the space of unlabeled pairwise drugs. To address these issues, we innovatively introduce S$^2$VM, a Self-supervised Visual pretraining framework for pair-wise Molecules, to fully fuse structural representations and explore the space of drug pairs for DDI prediction. S$^2$VM incorporates the explicit structure and correlations of visual molecules, such as the positional relationships and connectivity between functional substructures. Specifically, we blend the visual fragments of drug pairs into a unified input for joint encoding and then recover molecule-specific visual information for each drug individually. This approach integrates fine-grained structural representations from unlabeled drug pair data. By using visual fragments as anchors, S$^2$VM effectively captures the spatial information of local molecular components within visual molecules, resulting in more comprehensive embeddings of drug pairs. Experimental results show that S$^2$VM achieves state-of-the-art performance on widely used benchmarks, with Macro-F1 score improvements of 4.21% and 3.31%, respectively. Further extensive results and theoretical analysis demonstrate the effectiveness of S$^2$VM for both few-shot and novel drugs.
Tengfei Ma 0002, Yongsheng Zang, Yujie Chen 0002, Xuanbai Ren, Bosheng Song, Hongxin Xiang, Xiangxiang Zeng
NeurIPS9
2025 EDBench: Large-Scale Electron Density Data for Molecular Modeling
abstract
Existing molecular machine learning force fields (MLFFs) generally focus on the learning of atoms, molecules, and simple quantum chemical properties (such as energy and force), but ignore the importance of electron density (ED) $\rho(r)$ in accurately understanding molecular force fields (MFFs). ED describes the probability of finding electrons at specific locations around atoms or molecules, which uniquely determines all ground state properties (such as energy, molecular structure, etc.) of interactive multi-particle systems according to the Hohenberg-Kohn theorem. However, the calculation of ED relies on the time-consuming first-principles density functional theory (DFT), which leads to the lack of large-scale ED data and limits its application in MLFFs. In this paper, we introduce EDBench, a large-scale, high-quality dataset of ED designed to advance learning-based research at the electronic scale. Built upon the PCQM4Mv2, EDBench provides accurate ED data, covering 3.3 million molecules. To comprehensively evaluate the ability of models to understand and utilize electronic information, we design a suite of ED-centric benchmark tasks spanning prediction, retrieval, and generation. Our evaluation of several state-of-the-art methods demonstrates that learning from EDBench is not only feasible but also achieves high accuracy. Moreover, we show that learning-based methods can efficiently calculate ED with comparable precision while significantly reducing the computational cost relative to traditional DFT calculations. All data and benchmarks from EDBench will be freely available, laying a robust foundation for ED-driven drug discovery and materials science.
Hongxin Xiang, Mingquan Liu, Zhixiang Cheng, Wenjie Du 0003, Jun Xia 0001, Xin Jin 0014, Xiangxiang Zeng
NeurIPS10
2025 MOFormer: navigating the antimicrobial peptide design space with Pareto-based multi-objective transformer
abstract
Antimicrobial peptide (AMP) design through deep learning holds the potential to revolutionize antibiotic development. Despite recent progress in AMP generation, designing peptide antibiotics with multiple optimal properties remains a significant challenge. We present MOFormer, an advanced multi-objective AMP design pipeline capable of optimizing multiple AMP properties simultaneously. By leveraging a conditional Transformer, the model refines the AMP sequence-property landscape for efficient multi-objective generation. It also incorporates regularization techniques to maintain a highly structured space, enabling the sampling of precise and desirable candidates. Comparative analyses reveal that MOFormer achieves the optimal hypervolume in the multi-objective space, surpassing advanced methods in simultaneously maximizing antimicrobial activity (minimum inhibitory concentration) and minimizing hemolysis and toxicity, thereby yielding the most promising and desirable set of candidate peptides. When extended to a tri-objective scenario, MOFormer continues to exhibit remarkable optimization performance. Finally, we execute a hierarchical and rapid ranking of generated candidates based on Pareto fronts. We conducted a comprehensive validation of the physicochemical properties and target attributes of the candidates, while AlphaFold structure predictions revealed notably reliable predicted local distance difference test scores ranging from 70% to 87%. Our findings suggest that MOFormer holds potential to accelerate the discovery of efficacious peptide antibiotics by optimizing multi-objective trade-offs.
Li Wang 0145, Xiangzheng Fu, Xiucai Ye, Tetsuya Sakurai, Xiangxiang Zeng
Briefings Bioinform.7
2025 CMOMO: a deep multi-objective optimization framework for constrained molecular multi-property optimization
abstract
Molecular optimization, aiming to identify molecules with improved properties from a huge chemical search space, is a critical step in drug development. This task is challenging due to the need to optimize multiple properties while adhering to stringent drug-like criteria. Recently, numerous effective artificial intelligence methods have been proposed for molecular optimization. However, most of them neglect the constraints in molecular optimization, thereby limiting the development of high-quality molecules that simultaneously satisfy property objectives and constraint compliance. To address this issue, we proposed a deep multi-objective optimization framework, termed CMOMO, for constrained molecular multi-property optimization. The proposed CMOMO divides the optimization process into two stages, which enables it to use a dynamic constraint handling strategy to balance multi-property optimization and constraint satisfaction. Besides, a latent vector fragmentation based evolutionary reproduction strategy is designed to generate promising molecules effectively. Experimental results on two benchmark tasks show that the proposed CMOMO outperforms five state-of-the-art methods to obtain more successfully optimized molecules with multiple desired properties and satisfying drug-like constraints. Moreover, the superiority of CMOMO is verified on two practical tasks, including a potential protein-ligand optimization task of 4LDE protein, which is the structure of $\beta $2-adrenoceptor GPCR receptor, and a potential inhibitor optimization task of glycogen synthase kinase-3$\beta $ target (GSK3$\beta $). Notably, CMOMO demonstrates a two-fold improvement in success rate for the GSK3$\beta $ optimization task, successfully identifying molecules with favorable bioactivity, drug-likeness, synthetic accessibility, and adherence to structural constraints.
Xiangxiang Zeng, Xingyi Zhang 0001, Chun-Hou Zheng 0001, Yansen Su
Briefings Bioinform.3
2025 DrugAssist: a large language model for molecule optimization
abstract
Recently, the impressive performance of large language models (LLMs) on a wide range of tasks has attracted an increasing number of attempts to apply LLMs in drug discovery. However, molecule optimization, a critical task in the drug discovery pipeline, is currently an area that has seen little involvement from LLMs. Most of existing approaches focus solely on capturing the underlying patterns in chemical structures provided by the data, without taking advantage of expert feedback. These non-interactive approaches overlook the fact that the drug discovery process is actually one that requires the integration of expert experience and iterative refinement. To address this gap, we propose DrugAssist, an interactive molecule optimization model which performs optimization through human-machine dialogue by leveraging LLM's strong interactivity and generalizability. DrugAssist has achieved leading results in both single and multiple property optimization, simultaneously showcasing immense potential in transferability and iterative optimization. In addition, we publicly release a large instruction-based dataset called 'MolOpt-Instructions' for fine-tuning language models on molecule optimization tasks. We have made our code and data publicly available at https://github.com/blazerye/DrugAssist, which we hope to pave the way for future research in LLMs' application for drug discovery.
Geyan Ye, Xibao Cai, Houtim Lai, Xing Wang 0007, Junhong Huang, Longyue Wang, Wei Liu 0005, Xiangxiang Zeng
Briefings Bioinform.8
2025 MlyPredCSED: based on extreme point deviation compensated clustering combined with cross-scale convolutional neural networks to predict multiple lysine sites in human
abstract
In post-translational modification, covalent bonds on lysine and attached chemical groups significantly change proteins' physical and chemical properties. They shape protein structures, enhance function and stability, and are vital for physiological processes, affecting health and disease through mechanisms like gene expression, signal transduction, protein degradation, and cell metabolism. Although lysine (K) modification sites are considered among the most common types of post-translational modifications in proteins, research on K-PTMs has largely overlooked the synergistic effects between different modifications and lacked the techniques to address the problem of sample imbalance. Based on this, the Extreme Point Deviation Compensated Clustering (EPDCC) Undersampling algorithm was proposed in this study and combined with Cross-Scale Convolutional Neural Networks (CSCNNs) to develop a novel computational tool, MlyPredCSED, for simultaneously predicting multiple lysine modification sites. MlyPredCSED employs Multi-Label Position-Specific Triad Amino Acid Propensity and the physicochemical properties of amino acids to enhance the richness of sequence information. To address the challenge of sample imbalance, the innovative EPDCC Undersampling technique was introduced to adjust the majority class samples. The model's training and testing phase relies on the advanced CSCNN framework. MlyPredCSED, through cross-validation and testing, outperformed existing models, especially in complex categories with multiple modification sites. This research not only provides an efficient method for the identification of lysine modification sites but also demonstrates its value in biological research and drug development. To facilitate efficient use of MlyPredCSED by researchers, we have specifically developed an accessible free web tool: http://www.mlypredcsed.com.
Yun Zuo 0001, Xingze Fang, Jiankang Chen, Jiayi Ji, Xiangrong Liu, Xiangxiang Zeng, Zhaohong Deng, Hongwei Yin, Anjing Zhao
Briefings Bioinform.8
2025 ZeroGEN: leveraging language models for zero-shot ligand design from protein sequences
abstract
MOTIVATION: Deep generative methods based on language models have the capability to generate new data that resemble a given distribution and have begun to gain traction in ligand design. However, existing models face significant challenges when it comes to generating ligands for unseen targets, a scenario known as zero-shot learning. The ability to effectively generate ligands for novel targets is crucial for accelerating drug discovery and expanding the applicability of ligand design. Therefore, there is a pressing need to develop robust deep generative frameworks that can operate efficiently in zero-shot scenarios. RESULTS: In this study, we introduce ZeroGEN, a novel zero-shot deep generative framework based on protein sequences. ZeroGEN analyzes extensive data on protein-ligand inter-relationships and incorporates contrastive learning to align known protein-ligand features, thereby enhancing the model's understanding of potential interactions between proteins and ligands. Additionally, ZeroGEN employs self-distillation to filter the initially generated data, retaining only the ligands deemed reliable by the model. It also implements data augmentation techniques to aid the model in identifying ligands that match unseen targets. Experimental results demonstrate that ZeroGEN successfully generates ligands for unseen targets with strong affinity and desirable drug-like properties. Furthermore, visualizations of molecular docking and attention matrices reveal that ZeroGEN can autonomously focus on key residues of proteins, underscoring its capability to understand and generate effective ligands for novel targets. AVAILABILITY AND IMPLEMENTATION: The source code and data of this work is freely available in the https://github.com/viko-3/ZeroGEN.
Yangyang Chen 0006, Pengyong Li, Xiangxiang Zeng, Lei Xu 0002
Bioinform.5
2025 AJGM: joint learning of heterogeneous gene networks with adaptive graphical model
abstract
MOTIVATION: Inferring gene networks provides insights into biological pathways and functional relationships among genes. When gene expression samples exhibit heterogeneity, they may originate from unknown subtypes, prompting the utilization of mixture Gaussian graphical model (GGM) for simultaneous subclassification and gene network inference. However, this method overlooks the heterogeneity of network relationships across subtypes and does not sufficiently emphasize shared relationships. Additionally, GGM assumes data follows a multivariate Gaussian distribution, which is often not the case with zero-inflated scRNA-seq data. RESULTS: We propose an Adaptive Joint Graphical Model (AJGM) for estimating multiple gene networks from single-cell or bulk data with unknown heterogeneity. In AJGM, an overall network is introduced to capture relationships shared by all samples. The model establishes connections between the subtype networks and the overall network through adaptive weights, enabling it to focus more effectively on gene relationships shared across all networks, thereby enhancing the accuracy of network estimation. On synthetic data, the proposed approach outperforms existing methods in terms of sample classification and network inference, particularly excelling in the identification of shared relationships. Applying this method to gene expression data from triple-negative breast cancer confirms known gene pathways and hub genes, while also revealing novel biological insights. AVAILABILITY AND IMPLEMENTATION: The Python code and demonstrations of the proposed approaches are available at https://github.com/yyytim/AJGM, and the software is archived in Zenodo with DOI: 10.5281/zenodo.14740972.
Shunqi Yang, Lingyi Hu, Pengzhou Chen, Xiangxiang Zeng, Shanjun Mao
Bioinform.4
2025 Self-supervised learning in drug discovery
Yangyang Chen 0006, Jianmin Wang 0016, Yanyi Chu, Qingpeng Zhang, Zhong Alan Li, Xiangxiang Zeng
Sci. China Inf. Sci.7
2025 The computational properties of P systems with mutative membrane structures
Bosheng Song, Chuanlong Hu, David Orellana-Martín, Antonio Ramírez-de-Arellano, Mario J. Pérez-Jiménez, Xiangxiang Zeng
Inf. Comput.6
2025 HyperACP: A cutting-edge hybrid framework for anticancer peptide classification via scalable feature extraction and adaptive neighbor-based synthesis
abstract
Cancer remains a major contributor to global mortality, constituting a significant and escalating threat to human health. Anticancer peptides (ACPs) have emerged as promising therapeutic agents due to their specific mechanisms of action, pronounced tumor-targeting capability, and low toxicity. Nevertheless, traditional approaches for ACP identification are constrained by their reliance on shallow, hand-crafted sequence features, which fail to capture deeper semantic and structural characteristics. Moreover, such models exhibit limited robustness and interpretability when confronted with practical challenges such as severe class imbalance. To address these limitations, this study proposes HyperACP, an innovative framework for ACP recognition that integrates deep representation learning, adaptive sampling, and mechanistic interpretability. The framework leverages the ESMC protein language model to extract comprehensive sequence features and employs a novel adaptive algorithm, ANBS, to mitigate class imbalance at the decision boundary. For enhanced model transparency, SHAP-Res is incorporated to elucidate the contributions of individual residues to the final predictions. Comprehensive evaluations demonstrate that HyperACP consistently outperforms state-of-the-art methods across multiple datasets and validation protocols-including 10-fold cross-validation and independent test sets-according to metrics such as Accuracy (ACC), Sensitivity (SN), Specificity (SP), Matthews Correlation Coefficient (MCC), and Area Under the Curve (AUC). Furthermore, the model yields biologically interpretable results, pinpointing key residues (K, L, F, G) known to play pivotal roles in anticancer activity. These findings provide not only a robust predictive tool (available at www.hyperacp.com) but also novel insights into the structure-function relationships underlying ACPs.
Bangyi Zhang, Yun Zuo 0001, Jiayue Liu, Xiangrong Liu, Xiangxiang Zeng, Zhaohong Deng
PLoS Comput. Biol.6
2025 Molecular Dynamics-Powered Hierarchical Geometric Deep Learning Framework for Protein-Ligand Interaction
abstract
Accurate prediction of the drug binding between proteins and ligands can significantly advance the development of structure-based drug design. Recent advances have shown great potential in applying equivariant graph neural network (EGNN) -based methods to learn representations of protein-ligand (PL) complexes. However, most of them typically focus on atom-level graph representations and omit the residue-level information in PL complexes, which are considered essential for understanding the binding mechanism. In this article, we develop a SO(3)-equivariant hierarchical graph neural network (EHGNN) that effectively captures the intrinsic hierarchy of biomolecular structures to enhance the predictive performance of PL interactions. Based on the SO(3)-EHGNN, we further propose a molecular dynamics-powered and energy-guided deep learning framework, called Dynamics-PLI, to capture the spatial structures and energetic information inside molecular dynamic (MD) trajectories. Extensive experimental results show significant improvements over current state-of-the-art methods, with a decrease of 4.03% in RMSE for the binding affinity problem and an average increase of 3.95% in AUROC and AUPRC for the ligand efficacy problem, demonstrating the superiority of Dynamics-PLI for PL interaction prediction. Our findings indicate that the SO(3)-EHGNN exhibits enhanced performance without the necessity of pre-training, emphasizing the inherent analytical strength of SO(3)-EHGNN.
Mingquan Liu, Shuting Jin, Houtim Lai, Longyue Wang, Jianmin Wang 0016, Zhixiang Cheng, Xiangxiang Zeng
IEEE Trans. Comput. Biol. Bioinform.7
2025 Quality Scores Compression of Genomic Sequencing Data: A Comprehensive Review and Performance Evaluation
abstract
Advanced sequencing technologies have profoundly revolutionized biology and produced vast amounts of raw sequencing data during the past decades. The enormous amount of sequencing data proposed significant challenges of data storage and transmission. Compressing a big file into a small file is an encouraging method to tackle these challenges. Howver, it has been found that traditional text data compression algorithms are not well-suited for handing the vast sequencing datasets. Therefore, several algorithms are designed specifically for the efficient compression of sequencing data. Recently, considerable research has been devoted to compressing quality scores stored in the FASTQ format file, resulting in substantial advances in compression performance. Despite these advances, there has been no systematic review and evaluation of these algorithms or software. In this review, we aim to conduct a broad review of the existing quality score compression algorithms. We mainly discuss those algorithms from two categories, i.e., lossless and lossy compression. Additionally, we benchmark the compression performance of 12 tools using 14 real datasets. We anticipate that our review will provide practical guidance for others seeking to design an appropriate algorithm for compressing quality scores.
Yuansheng Liu, Zexuan Zhu 0001, Xiangxiang Zeng, Quan Zou 0001, Keqin Li 0001
IEEE Trans. Comput. Biol. Bioinform.4
2025 Multi-Modal Deep Representation Learning Accurately Identifies and Interprets Drug-Target Interactions
abstract
Deep learning offers efficient solutions for drug-target interaction prediction, but current methods often fail to capture the full complexity of multi-modal data (i.e., sequence, graphs, and three-dimensional structures), limiting both performance and generalization. Here, we present UnitedDTA, a novel explainable deep learning framework capable of integrating multi-modal biomolecule data to improve the binding affinity prediction, especially for novel (unseen) drugs and targets. UnitedDTA enables automatic learning unified discriminative representations from multi-modality data via contrastive learning and cross-attention mechanisms for cross-modality alignment and integration. Comparative results on multiple benchmark datasets show that UnitedDTA significantly outperforms the state-of-the-art drug-target affinity prediction methods and exhibits better generalization ability in predicting unseen drug-target pairs. More importantly, unlike most "black-box" deep learning methods, our well-established model offers better interpretability which enables us to directly infer the important substructures of the drug-target complexes that influence the binding activity, thus providing the insights in unveiling the binding preferences. Moreover, by extending UnitedDTA to other downstream tasks (e.g., molecular property prediction), we showcase the proposed multi-modal representation learning is capable of capturing the latent molecular representations that are closely associated with the molecular property, demonstrating the broad application potential for advancing the drug discovery process.
Jiayue Hu, Xiangxiang Zeng, Quan Zou 0001, Ran Su, Leyi Wei
IEEE J. Biomed. Health Informatics3
2025 PKAN: Leveraging Kolmogorov-Arnold Networks and Multi-Modal Learning for Peptide Prediction With Advanced Language Models
abstract
Peptides can offer highly specific biological activities, serving as essential mediators of intercellular signaling, which are critical for advancing precision medicine and drug development. Their primary structure can be depicted either as an amino acid sequence or as a chemical molecules consisting of atoms and chemical bonds. Large language models (LLMs) hold the potential to thoroughly elucidate the intricate intrinsic properties of peptides. Here we present the Peptide Kolmogorov-Arnold Network (PKAN), a framework leveraging multi-modal representations inspired by advanced language models for peptide activity and functionality prediction. Comparative experiments across tasks show that PKAN outperforms state-of-the-art models while maintaining a streamlined design with superior predictive capabilities. The multi-modal feature importance scoring, anchored in global structures and the significant marginal impacts of derived features on the model, coupled with intricate symbolic regression of specific activation functions, further demonstrates the robustness and precision of the PKAN framework in identifying and elucidating key determinants of peptide functionality. This work provides scientific evidence for investigating the complex mechanisms of peptide materials and supports the progression of peptide language paradigms in biology.
Li Wang 0145, Xiangzheng Fu, Xiucai Ye, Tetsuya Sakurai, Xiangxiang Zeng
IEEE J. Biomed. Health Informatics5
2025 Multiview Deep Learning-Based Molecule Design and Structural Optimization Accelerates Inhibitor Discover
abstract
In this work, we propose MEDICO, a multiview deep generative model for molecule generation, structural optimization, and the SARS-CoV-2 inhibitor discovery. To the best of our knowledge, MEDICO is the first-of-this-kind graph generative model that can generate molecular graphs similar to the structure of targeted molecules, with a multiview representation learning framework to sufficiently and adaptively learn comprehensive structural semantics from targeted molecular topology and geometry. We show that our MEDICO significantly outperforms the state-of-the-art methods in generating valid, novel, and unique molecules under benchmarking comparisons, particularly achieving $\tilde {8}5 \%$ improvement compared with the state-of-the-art methods in terms of validity. Importantly, we showcase that the multiview deep learning model enables us to generate not only the molecules structurally similar to the targeted molecules but also the molecules with desired chemical properties. Moreover, case study results on targeted molecule generation for the SARS-CoV-2 main protease (Mpro) show that we successfully generate new small molecules with desired drug-like properties for the Mpro by integrating molecular docking into our model as a chemical priori, potentially accelerating the de novo design of COVID-19 drugs. Furthermore, we apply MEDICO to the structural optimization of three well-known Mpro inhibitors (N3, 11a, and GC376) and achieve $\tilde {8}8 \%$ improvement compared with the origin inhibitors in their binding affinity to Mpro, demonstrating the application value of our model for the development of therapeutics for SARS-CoV-2 infection.
Ruheng Wang, Quan Zou 0001, Xiangxiang Zeng, Ran Su, Leyi Wei
IEEE Trans. Neural Networks Learn. Syst.7
2024 CSCL-DTI: predicting drug-target interaction through cross-view and self-supervised contrastive learning
abstract
Accurately predicting drug-target interactions (DTI) is a critical step in drug discovery. Existing methods of DTI prediction primarily employ Simplified Molecular-Input Line-Entry System (SMILES) sequences or molecular graphs to learn drug representations. However, the features learned by such single-view approach is prone to incomplete. While some multiview methods that consider the views of both SMILES sequences and molecular graphs have been developed, these methods often fall in short in capturing potential interactions between views. In this work, we propose a novel dual contrastive learning framework CSCL-DTI for DTI prediction. First, we design a contrastive-enhanced cross-view representation learning (CVRL) to learn representations for drugs. In this module, Transformer-based and graph convolutional network (GCN)-based encoders are separately adopted to learn view-specific representations, followed by contrastive learning to enrich the representations by accounting for the potential interplay between local chemical context and topological structure. Second, we combine Transformer with self-supervised contrastive learning (SSCL) to learn representations for targets by modelling protein amino acids sequences. The scheme allows to effectively preserve the intrinsic characteristics of the sequences. Finally, we introduce a bilinear attention network to obtain an integrated representation by adaptively incorporating drug and target representations. Benchmarking experiments on two datasets demonstrated that CSCL-DTI1outperforms six state-of-the-art methods.
Xuan Lin, Xi Zhang 0008, Yahui Long, Xiangxiang Zeng, Philip S. Yu
BIBM5
2024 An Image-enhanced Molecular Graph Representation Learning Framework
Hongxin Xiang, Shuting Jin, Jun Xia 0001, Jianmin Wang 0016, Xiangxiang Zeng
IJCAI7
2024 Multi-Objective Molecular Design in Constrained Latent Space
abstract
In recent times, molecular design has undergone significant advancements, particularly with the integration of artificial intelligence (AI) techniques for discovering molecules with various desired attributes. The use of generative models, especially variational autoencoders (VAEs), has proven to be a potent and efficient method. These models facilitate the rapid identification of new molecules that align with specific research goals. However, sequence-based generative models, like SMILES-based VAEs, often generate invalid molecules, presenting substantial challenges due to the inherent constraints in their latent spaces. To overcome this issue, we introduce a novel multi-objective molecular design approach that incorporates a corrector-based constraint handling technique. This technique employs a transformer as a corrector to convert invalid molecules into valid ones during the search process. Following correction, the latent space is segmented into distinct zones, organized via a spatial partition tree. Utilizing Monte Carlo tree search, we pinpoint the most promising zones for evolutionary-based sampling. Our method applies multi-objective molecular design within these constrained latent spaces. Our experimental results demonstrate that this approach markedly improves the quality of molecules generated from the latent space.
Li Wang 0145, Xiangxiang Zeng
IJCNN6
2024 SSR-DTA: Substructure-aware multi-layer graph neural networks for drug-target binding affinity prediction
Yuansheng Liu, Xinyan Xia, Yongshun Gong, Bosheng Song, Xiangxiang Zeng
Artif. Intell. Medicine5
2024 Attribute-guided prototype network for few-shot molecular property prediction
abstract
The molecular property prediction (MPP) plays a crucial role in the drug discovery process, providing valuable insights for molecule evaluation and screening. Although deep learning has achieved numerous advances in this area, its success often depends on the availability of substantial labeled data. The few-shot MPP is a more challenging scenario, which aims to identify unseen property with only few available molecules. In this paper, we propose an attribute-guided prototype network (APN) to address the challenge. APN first introduces an molecular attribute extractor, which can not only extract three different types of fingerprint attributes (single fingerprint attributes, dual fingerprint attributes, triplet fingerprint attributes) by considering seven circular-based, five path-based, and two substructure-based fingerprints, but also automatically extract deep attributes from self-supervised learning methods. Furthermore, APN designs the Attribute-Guided Dual-channel Attention module to learn the relationship between the molecular graphs and attributes and refine the local and global representation of the molecules. Compared with existing works, APN leverages high-level human-defined attributes and helps the model to explicitly generalize knowledge in molecular graphs. Experiments on benchmark datasets show that APN can achieve state-of-the-art performance in most cases and demonstrate that the attributes are effective for improving few-shot MPP performance. In addition, the strong generalization ability of APN is verified by conducting experiments on data from different domains.
Linlin Hou, Hongxin Xiang, Xiangxiang Zeng, Dong-Sheng Cao 0001, Bosheng Song
Briefings Bioinform.3
2024 AMGDTI: drug-target interaction prediction based on adaptive meta-graph learning in heterogeneous network
abstract
Prediction of drug-target interactions (DTIs) is essential in medicine field, since it benefits the identification of molecular structures potentially interacting with drugs and facilitates the discovery and reposition of drugs. Recently, much attention has been attracted to network representation learning to learn rich information from heterogeneous data. Although network representation learning algorithms have achieved success in predicting DTI, several manually designed meta-graphs limit the capability of extracting complex semantic information. To address the problem, we introduce an adaptive meta-graph-based method, termed AMGDTI, for DTI prediction. In the proposed AMGDTI, the semantic information is automatically aggregated from a heterogeneous network by training an adaptive meta-graph, thereby achieving efficient information integration without requiring domain knowledge. The effectiveness of the proposed AMGDTI is verified on two benchmark datasets. Experimental results demonstrate that the AMGDTI method overall outperforms eight state-of-the-art methods in predicting DTI and achieves the accurate identification of novel DTIs. It is also verified that the adaptive meta-graph exhibits flexibility and effectively captures complex fine-grained semantic information, enabling the learning of intricate heterogeneous network topology and the inference of potential drug-target relationship.
Yansen Su, Zhiyang Hu, Fei Wang 0095, Yannan Bin, Chun-Hou Zheng 0001, Haitao Li 0004, Xiangxiang Zeng
Briefings Bioinform.8
2024 MSlocPRED: deep transfer learning-based identification of multi-label mRNA subcellular localization
abstract
Subcellular localization of messenger ribonucleic acid (mRNA) is a universal mechanism for precise and efficient control of the translation process. Although many computational methods have been constructed by researchers for predicting mRNA subcellular localization, very few of these computational methods have been designed to predict subcellular localization with multiple localization annotations, and their generalization performance could be improved. In this study, the prediction model MSlocPRED was constructed to identify multi-label mRNA subcellular localization. First, the preprocessed Dataset 1 and Dataset 2 are transformed into the form of images. The proposed MDNDO-SMDU resampling technique is then used to balance the number of samples in each category in the training dataset. Finally, deep transfer learning was used to construct the predictive model MSlocPRED to identify subcellular localization for 16 classes (Dataset 1) and 18 classes (Dataset 2). The results of comparative tests of different resampling techniques show that the resampling technique proposed in this study is more effective in preprocessing for subcellular localization. The prediction results of the datasets constructed by intercepting different NC end (Both the 5' and 3' untranslated regions that flank the protein-coding sequence and influence mRNA function without encoding proteins themselves.) lengths show that for Dataset 1 and Dataset 2, the prediction performance is best when the NC end is intercepted by 35 nucleotides, respectively. The results of both independent testing and five-fold cross-validation comparisons with established prediction tools show that MSlocPRED is significantly better than established tools for identifying multi-label mRNA subcellular localization. Additionally, to understand how the MSlocPRED model works during the prediction process, SHapley Additive exPlanations was used to explain it. The predictive model and associated datasets are available on the following github: https://github.com/ZBYnb1/MSlocPRED/tree/main.
Yun Zuo 0001, Bangyi Zhang, Wenying He, Yue Bi, Xiangrong Liu, Xiangxiang Zeng, Zhaohong Deng
Briefings Bioinform.6
2024 Anti-symmetric framework for balanced learning of protein-protein interactions
abstract
MOTIVATION: Protein-protein interactions (PPIs) are essential for the regulation and facilitation of virtually all biological processes. Computational tools, particularly those based on deep learning, are preferred for the efficient prediction of PPIs. Despite recent progress, two challenges remain unresolved: (i) the imbalanced nature of PPI characteristics is often ignored and (ii) there exists a high computational cost associated with capturing long-range dependencies within protein data, typically exhibiting quadratic complexity relative to the length of the protein sequence. RESULT: Here, we propose an anti-symmetric graph learning model, BaPPI, for the balanced prediction of PPIs and extrapolation of the involved patterns in PPI network. In BaPPI, the contextualized information of protein data is efficiently handled by an attention-free mechanism formed by recurrent convolution operator. The anti-symmetric graph convolutional network is employed to model the uneven distribution within PPI networks, aiming to learn a more robust and balanced representation of the relationships between proteins. Ultimately, the model is updated using asymmetric loss. The experimental results on classical baseline datasets demonstrate that BaPPI outperforms four state-of-the-art PPI prediction methods. In terms of Micro-F1, BaPPI exceeds the second-best method by 6.5% on SHS27K and 5.3% on SHS148K. Further analysis of the generalization ability and patterns of predicted PPIs also demonstrates our model's generalizability and robustness to the imbalanced nature of PPI datasets. AVAILABILITY AND IMPLEMENTATION: The source code of this work is publicly available at https://github.com/ttan6729/BaPPI.
Weizhuo Li, Yuansheng Liu, Xiangxiang Zeng
Bioinform.6
2024 DualSyn: A dual-level feature interaction method to predict synergistic drug combinations
Xiangzhen Shen, Yuansheng Liu, Xuan Lin, Daojian Zeng, Xiangxiang Zeng
Expert Syst. Appl.7
2024 Geometric deep learning for drug discovery
Mingquan Liu, Chunyan Li 0002, Ruizhe Chen, Dong-Sheng Cao 0001, Xiangxiang Zeng
Expert Syst. Appl.5
2024 Effective drug-target affinity prediction via generative active learning
Yuansheng Liu, Zhenran Zhou, Dong-Sheng Cao 0001, Xiangxiang Zeng
Inf. Sci.5
2024 ECD-CDGI: An efficient energy-constrained diffusion model for cancer driver gene identification
abstract
The identification of cancer driver genes (CDGs) poses challenges due to the intricate interdependencies among genes and the influence of measurement errors and noise. We propose a novel energy-constrained diffusion (ECD)-based model for identifying CDGs, termed ECD-CDGI. This model is the first to design an ECD-Attention encoder by combining the ECD technique with an attention mechanism. ECD-Attention encoder excels at generating robust gene representations that reveal the complex interdependencies among genes while reducing the impact of data noise. We concatenate topological embedding extracted from gene-gene networks through graph transformers to these gene representations. We conduct extensive experiments across three testing scenarios. Extensive experiments show that the ECD-CDGI model possesses the ability to not only be proficient in identifying known CDGs but also efficiently uncover unknown potential CDGs. Furthermore, compared to the GNN-based approach, the ECD-CDGI model exhibits fewer constraints by existing gene-gene networks, thereby enhancing its capability to identify CDGs. Additionally, ECD-CDGI is open-source and freely available. We have also launched the model as a complimentary online tool specifically crafted to expedite research efforts focused on CDGs identification.
Linlin Zhuo, Xiangzheng Fu, Xiangxiang Zeng, Quan Zou 0001
PLoS Comput. Biol.5
2024 PreMLS: The undersampling technique based on ClusterCentroids to predict multiple lysine sites
abstract
The translated protein undergoes a specific modification process, which involves the formation of covalent bonds on lysine residues and the attachment of small chemical moieties. The protein's fundamental physicochemical properties undergo a significant alteration. The change significantly alters the proteins' 3D structure and activity, enabling them to modulate key physiological processes. The modulation encompasses inhibiting cancer cell growth, delaying ovarian aging, regulating metabolic diseases, and ameliorating depression. Consequently, the identification and comprehension of post-translational lysine modifications hold substantial value in the realms of biological research and drug development. Post-translational modifications (PTMs) at lysine (K) sites are among the most common protein modifications. However, research on K-PTMs has been largely centered on identifying individual modification types, with a relative scarcity of balanced data analysis techniques. In this study, a classification system is developed for the prediction of concurrent multiple modifications at a single lysine residue. Initially, a well-established multi-label position-specific triad amino acid propensity algorithm is utilized for feature encoding. Subsequently, PreMLS: a novel ClusterCentroids undersampling algorithm based on MiniBatchKmeans was introduced to eliminate redundant or similar major class samples, thereby mitigating the issue of class imbalance. A convolutional neural network architecture was specifically constructed for the analysis of biological sequences to predict multiple lysine modification sites. The model, evaluated through five-fold cross-validation and independent testing, was found to significantly outperform existing models such as iMul-kSite and predML-Site. The results presented here aid in prioritizing potential lysine modification sites, facilitating subsequent biological assays and advancing pharmaceutical research. To enhance accessibility, an open-access predictive script has been crafted for the multi-label predictive model developed in this study.
Yun Zuo 0001, Xingze Fang, Jiayong Wan, Wenying He, Xiangrong Liu, Xiangxiang Zeng, Zhaohong Deng
PLoS Comput. Biol.6
2024 On the migrativity properties between uni-nullnorms and overlap (grouping) functions
Xiangxiang Zeng, Kuanyun Zhu
Soft Comput.1
2024 scCAN: Clustering With Adaptive Neighbor-Based Imputation Method for Single-Cell RNA-Seq Data
abstract
Single-cell RNA sequencing (scRNA-seq) is widely used to study cellular heterogeneity in different samples. However, due to technical deficiencies, dropout events often result in zero gene expression values in the gene expression matrix. In this paper, we propose a new imputation method called scCAN, based on adaptive neighborhood clustering, to estimate the zero value of dropouts. Our method continuously updates cell-cell similarity information by simultaneously learning similarity relationships, clustering structures, and imposing new rank constraints on the Laplacian matrix of the similarity matrix, improving the imputation of dropout zero values. To evaluate the performance of this method, we used four simulated and eight real scRNA-seq data for downstream analyses, including cell clustering, recovered gene expression, and reconstructed cell trajectories. Our method improves the performance of the downstream analysis and is better than other imputation methods.
Shujie Dong, Yuansheng Liu, Yongshun Gong, Xiangjun Dong 0001, Xiangxiang Zeng
IEEE ACM Trans. Comput. Biol. Bioinform.5
2024 GenoM7GNet: An Efficient N7-Methylguanosine Site Prediction Approach Based on a Nucleotide Language Model
abstract
N-methylguanosine (m7G), one of the mainstream post-transcriptional RNA modifications, occupies an exceedingly significant place in medical treatments. However, classic approaches for identifying m7G sites are costly both in time and equipment. Meanwhile, the existing machine learning methods extract limited hidden information from RNA sequences, thus making it difficult to improve the accuracy. Therefore, we put forward to a deep learning network, called "GenoM7GNet," for m7G site identification. This model utilizes a Bidirectional Encoder Representation from Transformers (BERT) and is pretrained on nucleotide sequences data to capture hidden patterns from RNA sequences for m7G site prediction. Moreover, through detailed comparative experiments with various deep learning models, we discovered that the one-dimensional convolutional neural network (CNN) exhibits outstanding performance in sequence feature learning and classification. The proposed GenoM7GNet model achieved 0.953in accuracy, 0.932in sensitivity, 0.976in specificity, 0.907in Matthews Correlation Coefficient and 0.984in Area Under the receiver operating characteristic Curve on performance evaluation. Extensive experimental results further prove that our GenoM7GNet model markedly surpasses other state-of-the-art models in predicting m7G sites, exhibiting high computing performance.
Chuang Li 0004, Heshi Wang, Yanhua Wen, Rui Yin 0002, Xiangxiang Zeng, Keqin Li 0001
IEEE ACM Trans. Comput. Biol. Bioinform.5
2024 Dynamic threshold spiking neural P systems with weights and multiple channels
Bosheng Song, Yuansheng Liu, Xiangxiang Zeng, Shengye Huang
Theor. Comput. Sci.4
2024 Evolutionary Multimodal Multiobjective Optimization for Traveling Salesman Problems
abstract
Multimodal multiobjective optimization problems (MMOPs) are commonly seen in real-world applications. Many evolutionary algorithms have been proposed to solve continuous MMOPs. However, little effort has been made to solve combinatorial (or discrete) MMOPs. Searching for equivalent Pareto-optimal solutions in the discrete decision space is challenging. Moreover, the true Pareto-optimal solutions of a combinatorial MMOP are usually difficult to know, which has limited the development of its optimizer. In this article, we first propose a test problem generator for multimodal multiobjective traveling salesman problems (MMTSPs). It can readily generate MMTSPs with known Pareto-optimal solutions. Then, we propose a novel evolutionary algorithm to solve MMTSPs. In our proposed algorithm, we develop two new edge assembly crossover operators, which are specialized in searching for superior solutions to MMTSPs. Moreover, the proposed algorithm uses a new environmental selection operator to maintain a good balance between the objective space diversity and decision space diversity. We compare our algorithm with five state-of-the-art designs. Experimental results convincingly show that our algorithm is powerful in solving MMTSPs.
Liting Xu, Yuyan Han, Xiangxiang Zeng, Gary G. Yen, Hisao Ishibuchi
IEEE Trans. Evol. Comput.4
2024 PEB-DDI: A Task-Specific Dual-View Substructural Learning Framework for Drug-Drug Interaction Prediction
abstract
Adverse drug-drug interactions (DDIs) pose potential risks in polypharmacy due to unknown physicochemical incompatibilities between co-administered drugs. Recent studies have utilized multi-layer graph neural network architectures to model hierarchical molecular substructures of drugs, achieving excellent DDI prediction performance. While extant substructural frameworks effectively encode interactions from atom-level features, they overlook valuable chemical bond representations within molecular graphs. More critically, given the multifaceted nature of DDI prediction tasks involving both known and novel drug combinations, previous methods lack tailored strategies to address these distinct scenarios. The resulting lack of adaptability impedes further improvements to model performance. To tackle these challenges, we propose PEB-DDI, a DDI prediction learning framework with enhanced substructure extraction. First, the information of chemical bonds is integrated and synchronously updated with the atomic nodes. Then, different dual-view strategies are selected based on whether novel drugs are present in the prediction task. Particularly, we constructed Molecular fingerprint-Molecular graph view for transductive task, and Bipartite graph-Molecular graph view for inductive task. Rigorous evaluations on benchmark datasets underscore PEB-DDI's superior performance. Notably, on DrugBank, it achieves an outstanding accuracy rate of 98.18% when predicting previously unknown interactions among approved drugs. Even when faced with novel drugs, PEB-DDI consistently exhibits outstanding generalization capabilities with an accuracy rate of 88.06%, attributing to the proper migrating of molecular basic structure learning.
Xiangzhen Shen, Yuansheng Liu, Bosheng Song, Xiangxiang Zeng
IEEE J. Biomed. Health Informatics5
2024 Dual-View Learning Based on Images and Sequences for Molecular Property Prediction
abstract
The prediction of molecular properties remains a challenging task in the field of drug design and development. Recently, there has been a growing interest in the analysis of biological images. Molecular images, as a novel representation, have proven to be competitive, yet they lack explicit information and detailed semantic richness. Conversely, semantic information in SMILES sequences is explicit but lacks spatial structural details. Therefore, in this study, we focus on and explore the relationship between these two types of representations, proposing a novel multimodal architecture named ISMol. ISMol relies on a cross-attention mechanism to extract information representations of molecules from both images and SMILES strings, thereby predicting molecular properties. Evaluation results on 14 small molecule ADMET datasets indicate that ISMol outperforms machine learning (ML) and deep learning (DL) models based on single-modal representations. In addition, we analyze our method through a large number of experiments to test the superiority, interpretability and generalizability of the method. In summary, ISMol offers a powerful deep learning toolbox for drug discovery in a variety of molecular properties.
Xiang Zhang 0008, Hongxin Xiang, Xixi Yang, Jingxin Dong 0002, Xiangzheng Fu, Xiangxiang Zeng, Keqin Li 0001
IEEE J. Biomed. Health Informatics6
2024 Learning to Denoise Biomedical Knowledge Graph for Robust Molecular Interaction Prediction
abstract
Molecular interaction prediction plays a crucial role in forecasting unknown interactions between molecules, such as drug-target interaction (DTI) and drug-drug interaction (DDI), which are essential in the field of drug discovery and therapeutics. Although previous prediction methods have yielded promising results by leveraging the rich semantics and topological structure of biomedical knowledge graphs (KGs), they have primarily focused on enhancing predictive performance without addressing the presence of inevitable noise and inconsistent semantics. This limitation has hindered the advancement of KG-based prediction methods. To address this limitation, we propose BioKDN (BiomedicalKnowledge GraphDenoisingNetwork) for robust molecular interaction prediction. BioKDN refines the reliable structure of local subgraphs by denoising noisy links in a learnable manner, providing a general module for extracting task-relevant interactions. To enhance the reliability of the refined structure, BioKDN maintains consistent and robust semantics by smoothing relations around the target interaction. By maximizing the mutual information between reliable structure and smoothed relations, BioKDN emphasizes informative semantics to enable precise predictions. Experimental results on real-world datasets show that BioKDN surpasses state-of-the-art models in DTI and DDI prediction tasks, confirming the effectiveness and robustness of BioKDN in denoising unreliable interactions within contaminated KGs.
Tengfei Ma 0002, Yujie Chen 0002, Wen Tao, Dashun Zheng, Xuan Lin, Patrick Pang 0001, Yijun Wang 0002, Longyue Wang, Bosheng Song, Xiangxiang Zeng, Philip S. Yu
IEEE Trans. Knowl. Data Eng.11
2024 Memristive Circuit Implementation of Caenorhabditis Elegans Mechanism for Neuromorphic Computing
abstract
To overcome the energy efficiency bottleneck of the von Neumann architecture and scaling limit of silicon transistors, an emerging but promising solution is neuromorphic computing, a new computing paradigm inspired by how biological neural networks handle the massive amount of information in a parallel and efficient way. Recently, there is a surge of interest in the nematode worm Caenorhabditis elegans (C. elegans), an ideal model organism to probe the mechanisms of biological neural networks. In this article, we propose a neuron model for C. elegans with leaky integrate-and-fire (LIF) dynamics and adjustable integration time. We utilize these neurons to build the C. elegans neural network according to their neural physiology, which comprises: 1) sensory modules; 2) interneuron modules; and 3) motoneuron modules. Leveraging these block designs, we develop a serpentine robot system, which mimics the locomotion behavior of C. elegans upon external stimulus. Moreover, experimental results of C. elegans neurons presented in this article reveals the robustness (1% error w.r.t. 10% random noise) and flexibility of our design in term of parameter setting. The work paves the way for future intelligent systems by mimicking the C. elegans neural system.
Hegan Chen, Qinghui Hong, Chunhua Wang 0001, Xiangxiang Zeng, Jiliang Zhang 0002
IEEE Trans. Neural Networks Learn. Syst.5
2024 WalkGAN: Network Representation Learning With Sequence-Based Generative Adversarial Networks
abstract
Network representation learning, also known as network embedding, aims to learn the low-dimensional representations of vertices while capturing and preserving the network structure. For real-world networks, the edges that represent some important relationships between the vertices of a network may be missed and may result in degenerated performance. The existing methods usually treat missing edges as negative samples, thereby ignoring the true connections between two vertices in a network. To capture the true network structure effectively, we propose a novel network representation learning method called WalkGAN, where random walk scheme and generative adversarial networks (GAN) are incorporated into a network embedding framework. Specifically, WalkGAN leverages GAN to generate the synthetic sequences of the vertices that sufficiently simulate random walk on a network and further learn vertex representations from these vertex sequences. Thus, the unobserved links between the vertices are inferred with high probability instead of treating them as nonexistence. Experimental results on the benchmark network datasets demonstrate that WalkGAN achieves significant performance improvements for vertex classification, link prediction, and visualization tasks.
Taisong Jin, Xixi Yang, Zhengtao Yu 0001, Yongmei Zhang, Feiran Jie, Xiangxiang Zeng, Min Jiang 0005
IEEE Trans. Neural Networks Learn. Syst.7
2024 Geometry-Based Molecular Generation With Deep Constrained Variational Autoencoder
abstract
Finding target molecules with specific chemical properties plays a decisive role in drug development. We proposed GEOM-CVAE, a constrained variational autoencoder based on geometric representation for molecular generation with specific properties, which is protein-context-dependent. In terms of machine learning, it includes continuous feature embedding encoder and molecular generation decoder. Our key contribution is to propose an efficient geometric embedding method, including the spatial structure representations of drug molecule (converting the 3-D coordinates into image) and the geometric graph representations of protein target (modeling the protein surface as a mesh). The 3-D geometric information is vital to successful molecular generation, which is different from previous molecular generative methods based on 1-D or 2-D. Our model framework generates specific molecules in two phases, by first generating special image with molecular 3-D information to learn latent representations and generating molecules with constrained condition based on geometric graph convolution for specific protein and then inputting the generated structural molecules into a parser network for obtaining Simplified Molecular Input Line Entry System (SMILES) strings. Our model achieves competitive performance that implies its potential effectiveness to enable the exploration of the vast chemical space for drug discovery.
Chunyan Li 0002, Junfeng Yao, Wei Wei 0006, Zhangming Niu, Xiangxiang Zeng, Jin Li 0007, Jianmin Wang 0016
IEEE Trans. Neural Networks Learn. Syst.5
2023 LagNet: Deep Lagrangian Mechanics for Plug-and-Play Molecular Representation Learning
abstract
Molecular representation learning is a fundamental problem in the field of drug discovery and molecular science. Whereas incorporating molecular 3D information in the representations of molecule seems beneficial, which is related to computational chemistry with the basic task of predicting stable 3D structures (conformations) of molecules. Existing machine learning methods either rely on 1D and 2D molecular properties or simulate molecular force field to use additional 3D structure information via Hamiltonian network. The former has the disadvantage of ignoring important 3D structure features, while the latter has the disadvantage that existing Hamiltonian neural network must satisfy the “canonial” constraint, which is difficult to be obeyed in many cases. In this paper, we propose a novel plug-and-play architecture LagNet by simulating molecular force field only with parameterized position coordinates, which implements Lagrangian mechanics to learn molecular representation by preserving 3D conformation without obeying any additional restrictions. LagNet is designed to generate known conformations and generalize for unknown ones from molecular SMILES. Implicit positions in LagNet are learned iteratively using discrete-time Lagrangian equations. Experimental results show that LagNet can well learn 3D molecular structure features, and outperforms previous state-of-the-art baselines related molecular representation by a significant margin.
Chunyan Li 0002, Junfeng Yao, Jinsong Su, Zhaoyang Liu 0002, Xiangxiang Zeng
AAAI5
2023 Interpretable multi-view attention network for drug-drug interaction prediction
abstract
Drug-drug interaction (DDI) plays an increasingly crucial role in drug discovery. Predicting potential DDI is also essential for clinical research. Given the high cost and risk of wet-lab experiments, in-silico DDI prediction is an alternative choice. Recently, deep learning methods have been developed for DDI prediction. However, most of existing methods focus on feature extraction from either molecular SMILES sequences or drug interactive networks, ignoring the valuable complementary information that can be derived from these two views. In this paper, we propose a novel interpretable Multi-View Attention network (MVA-DDI) for DDI prediction. MVA-DDI can effectively extracts drug representations from different perspectives to improve DDI prediction. Specifically, for a given drug, we design a transformer-based encoder and a graph convolutional networkbased encoder to learn sequence and graph representations from SMILES sequence and molecular graph, respectively. To fully exploit the complementary information between the sequence and molecular views, an attention mechanism is further adopted to adaptively aggregate the sequence and graph representations by taking the importance of different views into accounts, generating the final drug representations. Comparison experiments demonstrated that our MVA-DDI1model achieved superior performance to state-of-the-art models on DDI prediction.
Xuan Lin, Qi Wen 0002, Yahui Long, Xiangxiang Zeng
BIBM6
2023 PACS: Prediction and analysis of cancer subtypes from multi-omics data based on a multi-head attention mechanism model
abstract
Due to the high heterogeneity and clinical characteristics of cancer, there are significant differences in multi-omic data and clinical characteristics among different cancer subtypes. Therefore, accurate classification of cancer subtypes can help doctors choose the most appropriate treatment options, improve treatment outcomes, and provide more accurate patient survival predictions. In this study, we propose a supervised multi-head attention mechanism model (SMA) to classify cancer subtypes successfully. The attention mechanism and feature sharing module of the SMA model can successfully learn the global and local feature information of multi-omics data. Second, it enriches the parameters of the model by deeply fusing multi-head attention encoders from Siamese through the fusion module. Validated by extensive experiments, the SMA model achieves the highest accuracy, F1 macroscopic, F1 weighted, and accurate classification of cancer subtypes in simulated, single-cell, and cancer multi-omics datasets compared to AE, CNN, and GNN-based models. Therefore, we contribute to future research on multi-omics data using our attention-based approach.
Liangrui Pan, Pinle Qin, Pengfei Rong, Xiangxiang Zeng, Dazheng Liu, Shaoliang Peng
BIBM4
2023 ImmRegInformer: an interpretable Transformer-based method for prioritizing immune-regulatory cancer drivers
abstract
The identification and prioritization of immune-regulatory cancer driver mutations present a promising study for precision immunotherapy of cancer but remain considerable challenges. Here we introduced a novel method ImmRegInformer to systematically explore the regulatory relationship between cancer driver mutations and immune response by leveraging the powerful Transformer model and the lasso-regularised ordinal regression. In particular, our method integrated the mutation co-occurrence information with the self-attention weight to discern the underlying relationships between different driver mutations when regulating the immune cytolytic activity (CYT). Using ImmRegInformer, we identified 250 immune-regulating driver mutations in 8223 pan-cancer samples. They were verified in terms of the mutation frequency and interactions with the cytolytic signature genes. Further, we found the complementary roles of self-attention weight and mutation co-occurrence in prioritizing the driver mutations exhibited dominant associations with CYT. In conclusion, this study underscored the importance of employing deep learning methods like Transformer to unlock hidden insights into the biological complexities of cancer immunity, offering a new avenue in dissecting the immune regulatory mechanism and potential clinical applications. ImmRegInformer is freely available at https://github.com/Liwen-Liberty/ImmregInformer.
Yijun Peng, Wending Pi, Xiangxiang Zeng, Shaoliang Peng
BIBM5
2023 Adaptive Compositional Continual Meta-Learning
abstract
This paper focuses on continual meta-learning, where few-shot tasks are heterogeneous and sequentially available. Recent works use a mixture model for meta-knowledge to deal with the heterogeneity. However, these methods suffer from parameter inefficiency caused by two reasons: (1) the underlying assumption of mutual exclusiveness among mixture components hinders sharing meta-knowledge across heterogeneous tasks. (2) they only allow increasing mixture components and cannot adaptively filter out redundant components. In this paper, we propose an Adaptive Compositional Continual Meta-Learning (ACML) algorithm, which employs a compositional premise to associate a task with a subset of mixture components, allowing meta-knowledge sharing among heterogeneous tasks. Moreover, to adaptively adjust the number of mixture components, we propose a component sparsification method based on evidential theory to filter out redundant components. Experimental results show ACML outperforms strong baselines, showing the effectiveness of our compositional meta-knowledge, and confirming that ACML can adaptively learn meta-knowledge.
Bin Wu 0025, Jinyuan Fang, Xiangxiang Zeng, Shangsong Liang, Qiang Zhang 0026
ICML3
2023 GPMO: Gradient Perturbation-Based Contrastive Learning for Molecule Optimization
abstract
Optimizing molecules with desired properties is a crucial step in de novo drug design. While translation-based methods have achieved initial success, they continue to face the challenge of the “exposure bias” problem. The challenge of preventing the “exposure bias” problem of molecule optimization lies in the need for both positive and negative molecules of contrastive learning. That is because generating positive molecules through data augmentation requires domain-specific knowledge, and randomly sampled negative molecules are easily distinguished from the real molecules. Hence, in this work, we propose a molecule optimization method called GPMO, which leverages a gradient perturbation-based contrastive learning method to prevent the “exposure bias” problem in translation-based molecule optimization. With the assistance of positive and negative molecules, GPMO is able to effectively handle both real and artificial molecules. GPMO is a molecule optimization method that is conditioned on matched molecule pairs for drug discovery. Our empirical studies show that GPMO outperforms the state-of-the- art molecule optimization methods. Furthermore, the negative and positive perturbations improve the robustness of GPMO.
Xixi Yang, Yafeng Deng, Yuansheng Liu, Dong-Sheng Cao 0001, Xiangxiang Zeng
IJCAI6
2023 Totally Dynamic Hypergraph Neural Networks
abstract
Recent dynamic hypergraph neural networks (DHGNNs) are designed to adaptively optimize the hypergraph structure to avoid the dependence on the initial hypergraph structure, thus capturing more hidden information for representation learning. However, most existing DHGNNs cannot adjust the hyperedge number and thus fail to fully explore the underlying hypergraph structure. This paper proposes a new method, namely, totally hypergraph neural network (TDHNN), to adjust the hyperedge number for optimizing the hypergraph structure. Specifically, the proposed method first captures hyperedge feature distribution to obtain dynamical hyperedge features rather than fixed ones, by conducting the sampling from the learned distribution. The hypergraph is then constructed based on the attention coefficients of both sampled hyperedges and nodes. The node features are dynamically updated by designing a simple hypergraph convolution algorithm. Experimental results on real datasets demonstrate the effectiveness of the proposed method, compared to SOTA methods. The source code can be accessed via https://github.com/HHW-zhou/TDHNN.
Peng Zhou 0012, Zongqian Wu, Xiangxiang Zeng, Guoqiu Wen, Junbo Ma, Xiaofeng Zhu 0001
IJCAI3
2023 Dimensionality reduction and visualization of single-cell RNA-seq data with an improved deep variational autoencoder
abstract
Single-cell RNA sequencing (scRNA-seq) is a revolutionary breakthrough that determines the precise gene expressions on individual cells and deciphers cell heterogeneity and subpopulations. However, scRNA-seq data are much noisier than traditional high-throughput RNA-seq data because of technical limitations, leading to many scRNA-seq data studies about dimensionality reduction and visualization remaining at the basic data-stacking stage. In this study, we propose an improved variational autoencoder model (termed DREAM) for dimensionality reduction and a visual analysis of scRNA-seq data. Here, DREAM combines the variational autoencoder and Gaussian mixture model for cell type identification, meanwhile explicitly solving 'dropout' events by introducing the zero-inflated layer to obtain the low-dimensional representation that describes the changes in the original scRNA-seq dataset. Benchmarking comparisons across nine scRNA-seq datasets show that DREAM outperforms four state-of-the-art methods on average. Moreover, we prove that DREAM can accurately capture the expression dynamics of human preimplantation embryonic development. DREAM is implemented in Python, freely available via the GitHub website, https://github.com/Crystal-JJ/DREAM.
Junlin Xu, Yuansheng Liu, Bosheng Song, Xiulan Guo, Xiangxiang Zeng, Quan Zou 0001
Briefings Bioinform.6
2023 DSN-DDI: an accurate and generalized framework for drug-drug interaction prediction by dual-view representation learning
abstract
Drug-drug interaction (DDI) prediction identifies interactions of drug combinations in which the adverse side effects caused by the physicochemical incompatibility have attracted much attention. Previous studies usually model drug information from single or dual views of the whole drug molecules but ignore the detailed interactions among atoms, which leads to incomplete and noisy information and limits the accuracy of DDI prediction. In this work, we propose a novel dual-view drug representation learning network for DDI prediction ('DSN-DDI'), which employs local and global representation learning modules iteratively and learns drug substructures from the single drug ('intra-view') and the drug pair ('inter-view') simultaneously. Comprehensive evaluations demonstrate that DSN-DDI significantly improved performance on DDI prediction for the existing drugs by achieving a relatively improved accuracy of 13.01% and an over 99% accuracy under the transductive setting. More importantly, DSN-DDI achieves a relatively improved accuracy of 7.07% to unseen drugs and shows the usefulness for real-world DDI applications. Finally, DSN-DDI exhibits good transferability on synergistic drug combination prediction and thus can serve as a generalized framework in the drug discovery field.
Bin Shao 0002, Xiangxiang Zeng, Tong Wang 0014, Tie-Yan Liu
Briefings Bioinform.4
2023 Comprehensive evaluation of deep and graph learning on drug-drug interactions prediction
abstract
Recent advances and achievements of artificial intelligence (AI) as well as deep and graph learning models have established their usefulness in biomedical applications, especially in drug-drug interactions (DDIs). DDIs refer to a change in the effect of one drug to the presence of another drug in the human body, which plays an essential role in drug discovery and clinical research. DDIs prediction through traditional clinical trials and experiments is an expensive and time-consuming process. To correctly apply the advanced AI and deep learning, the developer and user meet various challenges such as the availability and encoding of data resources, and the design of computational methods. This review summarizes chemical structure based, network based, natural language processing based and hybrid methods, providing an updated and accessible guide to the broad researchers and development community with different domain knowledge. We introduce widely used molecular representation and describe the theoretical frameworks of graph neural network models for representing molecular structures. We present the advantages and disadvantages of deep and graph learning methods by performing comparative experiments. We discuss the potential technical challenges and highlight future directions of deep and graph learning models for accelerating DDIs prediction.
Xuan Lin, Lichang Dai, Yafang Zhou, Jianyu Shi, Dong-Sheng Cao 0001, Bosheng Song, Philip S. Yu, Xiangxiang Zeng
Briefings Bioinform.12
2023 Sequence Alignment/Map format: a comprehensive review of approaches and applications
abstract
The Sequence Alignment/Map (SAM) format file is the text file used to record alignment information. Alignment is the core of sequencing analysis, and downstream tasks accept mapping results for further processing. Given the rapid development of the sequencing industry today, a comprehensive understanding of the SAM format and related tools is necessary to meet the challenges of data processing and analysis. This paper is devoted to retrieving knowledge in the broad field of SAM. First, the format of SAM is introduced to understand the overall process of the sequencing analysis. Then, existing work is systematically classified in accordance with generation, compression and application, and the involved SAM tools are specifically mined. Lastly, a summary and some thoughts on future directions are provided.
Yuansheng Liu, Xiangzhen Shen, Yongshun Gong, Bosheng Song, Xiangxiang Zeng
Briefings Bioinform.6
2023 Machine learning on protein-protein interaction prediction: models, challenges and trends
abstract
Protein-protein interactions (PPIs) carry out the cellular processes of all living organisms. Experimental methods for PPI detection suffer from high cost and false-positive rate, hence efficient computational methods are highly desirable for facilitating PPI detection. In recent years, benefiting from the enormous amount of protein data produced by advanced high-throughput technologies, machine learning models have been well developed in the field of PPI prediction. In this paper, we present a comprehensive survey of the recently proposed machine learning-based prediction methods. The machine learning models applied in these methods and details of protein data representation are also outlined. To understand the potential improvements in PPI prediction, we discuss the trend in the development of machine learning-based methods. Finally, we highlight potential directions in PPI prediction, such as the use of computationally predicted protein structures to extend the data source for machine learning models. This review is supposed to serve as a companion for further improvements in this field.
Xiaocai Zhang, Yuansheng Liu, Binshuang Zheng, Yanlin Yin, Xiangxiang Zeng
Briefings Bioinform.7
2023 Prediction of multi-relational drug-gene interaction via Dynamic hyperGraph Contrastive Learning
abstract
Drug-gene interaction prediction occupies a crucial position in various areas of drug discovery, such as drug repurposing, lead discovery and off-target detection. Previous studies show good performance, but they are limited to exploring the binding interactions and ignoring the other interaction relationships. Graph neural networks have emerged as promising approaches owing to their powerful capability of modeling correlations under drug-gene bipartite graphs. Despite the widespread adoption of graph neural network-based methods, many of them experience performance degradation in situations where high-quality and sufficient training data are unavailable. Unfortunately, in practical drug discovery scenarios, interaction data are often sparse and noisy, which may lead to unsatisfactory results. To undertake the above challenges, we propose a novel Dynamic hyperGraph Contrastive Learning (DGCL) framework that exploits local and global relationships between drugs and genes. Specifically, graph convolutions are adopted to extract explicit local relations among drugs and genes. Meanwhile, the cooperation of dynamic hypergraph structure learning and hypergraph message passing enables the model to aggregate information in a global region. With flexible global-level messages, a self-augmented contrastive learning component is designed to constrain hypergraph structure learning and enhance the discrimination of drug/gene representations. Experiments conducted on three datasets show that DGCL is superior to eight state-of-the-art methods and notably gains a 7.6% performance improvement on the DGIdb dataset. Further analyses verify the robustness of DGCL for alleviating data sparsity and over-smoothing issues.
Wen Tao, Yuansheng Liu, Xuan Lin, Bosheng Song, Xiangxiang Zeng
Briefings Bioinform.5
2023 Reducing false positive rate of docking-based virtual screening by active learning
abstract
Machine learning-based scoring functions (MLSFs) have become a very favorable alternative to classical scoring functions because of their potential superior screening performance. However, the information of negative data used to construct MLSFs was rarely reported in the literature, and meanwhile the putative inactive molecules recorded in existing databases usually have obvious bias from active molecules. Here we proposed an easy-to-use method named AMLSF that combines active learning using negative molecular selection strategies with MLSF, which can iteratively improve the quality of inactive sets and thus reduce the false positive rate of virtual screening. We chose energy auxiliary terms learning as the MLSF and validated our method on eight targets in the diverse subset of DUD-E. For each target, we screened the IterBioScreen database by AMLSF and compared the screening results with those of the four control models. The results illustrate that the number of active molecules in the top 1000 molecules identified by AMLSF was significantly higher than those identified by the control models. In addition, the free energy calculation results for the top 10 molecules screened out by the AMLSF, null model and control models based on DUD-E also proved that more active molecules can be identified, and the false positive rate can be reduced by AMLSF.
Shao-Hua Shi, Xiangxiang Zeng, Su-You Liu, Zhao-Qian Liu, Yafeng Deng, Aiping Lu, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.4
2023 Chemical structure-aware molecular image representation learning
abstract
Current methods of molecular image-based drug discovery face two major challenges: (1) work effectively in absence of labels, and (2) capture chemical structure from implicitly encoded images. Given that chemical structures are explicitly encoded by molecular graphs (such as nitrogen, benzene rings and double bonds), we leverage self-supervised contrastive learning to transfer chemical knowledge from graphs to images. Specifically, we propose a novel Contrastive Graph-Image Pre-training (CGIP) framework for molecular representation learning, which learns explicit information in graphs and implicit information in images from large-scale unlabeled molecules via carefully designed intra- and inter-modal contrastive learning. We evaluate the performance of CGIP on multiple experimental settings (molecular property prediction, cross-modal retrieval and distribution similarity), and the results show that CGIP can achieve state-of-the-art performance on all 12 benchmark datasets and demonstrate that CGIP transfers chemical knowledge in graphs to molecular images, enabling image encoder to perceive chemical structures in images. We hope this simple and effective framework will inspire people to think about the value of image for molecular representation learning.
Hongxin Xiang, Shuting Jin, Xiangrong Liu, Xiangxiang Zeng
Briefings Bioinform.4
2023 Monodirectional evolutional symport tissue P systems with channel states and cell division
Bosheng Song, Kenli Li 0001, Xiangxiang Zeng, Mario J. Pérez-Jiménez, Claudio Zandron
Sci. China Inf. Sci.3
2023 A general hypergraph learning algorithm for drug multi-task predictions in micro-to-macro biomedical networks
abstract
The powerful combination of large-scale drug-related interaction networks and deep learning provides new opportunities for accelerating the process of drug discovery. However, chemical structures that play an important role in drug properties and high-order relations that involve a greater number of nodes are not tackled in current biomedical networks. In this study, we present a general hypergraph learning framework, which introduces Drug-Substructures relationship into Molecular interaction Networks to construct the micro-to-macro drug centric heterogeneous network (DSMN), and develop a multi-branches HyperGraph learning model, called HGDrug, for Drug multi-task predictions. HGDrug achieves highly accurate and robust predictions on 4 benchmark tasks (drug-drug, drug-target, drug-disease, and drug-side-effect interactions), outperforming 8 state-of-the-art task specific models and 6 general-purpose conventional models. Experiments analysis verifies the effectiveness and rationality of the HGDrug model architecture as well as the multi-branches setup, and demonstrates that HGDrug is able to capture the relations between drugs associated with the same functional groups. In addition, our proposed drug-substructure interaction networks can help improve the performance of existing network models for drug-related prediction tasks.
Shuting Jin, Yinghui Jiang, Leyi Wei, Zhuohang Yu, Xiangxiang Zeng, Xiangrong Liu
PLoS Comput. Biol.8
2023 Tissue P Systems With States in Cells
abstract
Tissue-like P systems with channel states are a type of classical membrane systems in which objects transferred among regions are controlled by states placed in the channels between regions. However, an important biological fact is the existence of a “barrier” to the diffusion of signal molecules, which tend to remain confined to some particular micro-habitat. This feature allows quorum sensing to convey information about the physiological state of spatially separated sub-populations. Therefore, in this article, we design a novel class-variant of P systems namedtissue P systems with states in cells(TSIC P systems). Here, each cell contains one and only one state at any moment (the environment has no state), and objects transferred among regions are controlled by states (or a state) that are placed in the corresponding cells (or a cell). We discuss thecomputability theoryof TSIC P systems by showing that Turing universality is acquired by TSIC P systems, which are worked both in a flat maximal parallelism and in a maximal parallelism. In addition, when cell division is considered in TSIC P systems, then tissue P systems with states in cells and cell division (TSICD P systems) are constructed. The (presumed)computational efficiencyof TSICD P systems is reached by offering a uniform solution to the satisfiability problem.
Bosheng Song, Kenli Li 0001, David Orellana-Martín, Xiangxiang Zeng, Mario J. Pérez-Jiménez
IEEE Trans. Computers4
2023 PointDE: Protein Docking Evaluation Using 3D Point Cloud Neural Network
abstract
Protein-protein interactions (PPIs) play essential roles in many vital movements and the determination of protein complex structure is helpful to discover the mechanism of PPI. Protein-protein docking is being developed to model the structure of the protein. However, there is still a challenge to selecting the near-native decoys generated by protein-protein docking. Here, we propose a docking evaluation method using 3D point cloud neural network named PointDE. PointDE transforms protein structure to the point cloud. Using the state-of-the-art point cloud network architecture and a novel grouping mechanism, PointDE can capture the geometries of the point cloud and learn the interaction information from the protein interface. On public datasets, PointDE surpasses the state-of-the-art method using deep learning. To further explore the ability of our method in different types of protein structures, we developed a new dataset generated by high-quality antibody-antigen complexes. The result in this antibody-antigen dataset shows the strong performance of PointDE, which will be helpful for the understanding of PPI mechanisms.
Xiaoping Min, Xiangxiang Zeng, Shengxiang Ge, Ning-Shao Xia
IEEE ACM Trans. Comput. Biol. Bioinform.5
2023 Effectively Identifying Compound-Protein Interaction Using Graph Neural Representation
abstract
Effectively identifying compound-protein interactions (CPIs) is crucial for new drug design, which is an important step in silico drug discovery. Current machine learning methods for CPI prediction mainly use one-demensional (1D) compound/protein strings and/or the specific descriptors. However, they often ignore the fact that molecules are essentially modeled by the molecular graph. We observe that in real-world scenarios, the topological structure information of the molecular graph usually provides an overview of how the atoms are connected, and the local chemical context reveals the functionality of the protein sequence in CPI. These two types of information are complementary to each other and they are both significant for modeling compound-protein pairs. Motivated by this, we propose an end-to-end deep learning framework named GraphCPI, which captures the structural information of compounds and leverages the chemical context of protein sequences for solving the CPI prediction task. Our framework can integrate any popular graph neural networks for learning compounds, and it combines with a convolutional neural network for embedding sequences. To compare our method with classic and state-of-the-art deep learning methods, we conduct extensive experiments based on several widely-used CPI datasets. The experimental results show the feasibility and competitiveness of our proposed method.
Xuan Lin, Zhe Quan, Zhi-Jie Wang 0009, Xiangxiang Zeng, Philip S. Yu
IEEE ACM Trans. Comput. Biol. Bioinform.5
2023 Modality-DTA: Multimodality Fusion Strategy for Drug-Target Affinity Prediction
abstract
Prediction of the drug-target affinity (DTA) plays an important role in drug discovery. Existing deep learning methods for DTA prediction typically leverage a single modality, namely simplified molecular input line entry specification (SMILES) or amino acid sequence to learn representations. SMILES or amino acid sequences can be encoded into different modalities. Multimodality data provide different kinds of information, with complementary roles for DTA prediction. We propose Modality-DTA, a novel deep learning method for DTA prediction that leverages the multimodality of drugs and targets. A group of backward propagation neural networks is applied to ensure the completeness of the reconstruction process from the latent feature representation to original multimodality data. The tag between the drug and target is used to reduce the noise information in the latent representation from multimodality data. Experiments on three benchmark datasets show that our Modality-DTA outperforms existing methods in all metrics. Modality-DTA reduces the mean square error by 15.7% and improves the area under the precisionrecall curve by 12.74% in the Davis dataset. We further find that the drug modality Morgan fingerprint and the target modality generated by one-hot-encoding play the most significant roles. To the best of our knowledge, Modality-DTA is the first method to explore multimodality for DTA prediction.
Xixi Yang, Zhangming Niu, Yuansheng Liu, Bosheng Song, Weiqiang Lu, Xiangxiang Zeng
IEEE ACM Trans. Comput. Biol. Bioinform.7
2023 Spiking neural P systems with weights and delays on synapses
Bosheng Song, Xiangxiang Zeng
Theor. Comput. Sci.3
2023 A Multi-Population Multi-Objective Evolutionary Algorithm Based on the Contribution of Decision Variables to Objectives for Large-Scale Multi/Many-Objective Optimization
abstract
Most existing multiobjective evolutionary algorithms treat all decision variables as a whole to perform genetic operations and optimize all objectives with one population at the same time. Considering different control attributes, different decision variables have different optimization effects on each objective, so decision variables can be divided into convergence- or diversity-related variables. In this article, we propose a new metric called the optimization degree of the convergence-related decision variable to each objective to calculate the contribution objective of each decision variable. All decision variables are grouped according to their contribution objectives. Then, a multiobjective evolutionary algorithm, namely, decision variable contributing to objectives evolutionary algorithm (DVCOEA), has been proposed. In order to balance the convergence and diversity of the population, the DVCOEA algorithm combines the multipopulation multiobjective framework, where two different optimization strategies are designed to optimize the subpopulation and individuals in the external archive, respectively. Finally, DVCOEA is compared with several state-of-the-art algorithms on a number of benchmark functions. Experimental results show that DVCOEA is a competitive approach for solving large-scale multi/many-objective problems.
Yusuke Nojima, Xiangxiang Zeng
IEEE Trans. Cybern.7
2023 KG-MTL: Knowledge Graph Enhanced Multi-Task Learning for Molecular Interaction
abstract
Molecular interaction prediction is essential in various applications including drug discovery and material science. The problem becomes quite challenging when the interaction is represented by unmapped relationships in molecular networks, namely molecular interaction, because it easily suffers from (i) insufficient labeled data with many false-positive samples, and (ii) ignoring a large number of biological entities with rich information in the knowledge graph. Most of the existing methods cannot properly exploit the information of knowledge graph and molecule graph simultaneously. In this paper, we propose a large-scaleKnowledgeGraph enhancedMulti-TaskLearning model, namely KG-MTL, which extracts the features from both knowledge graph and molecular graph in a synergistic way. Moreover, we design an effectiveShared Unitthat helps the model to jointly preserve the semantic relations of drug entity and the neighbor structures of the compound in both knowledge graph and molecular graph. Extensive experiments on four real-world datasets demonstrate that our proposed KG-MTL outperforms the state-of-the-art methods on two representative molecular interaction prediction tasks: drug-target interaction prediction and compound-protein interaction prediction. The source code of KG-MTL is available athttps://github.com/xzenglab/KG-MTL.
Tengfei Ma 0002, Xuan Lin, Bosheng Song, Philip S. Yu, Xiangxiang Zeng
IEEE Trans. Knowl. Data Eng.5
2022 A heterogeneous graph cross-omics attention model for single-cell representation learning
abstract
Single-cell multi-omics sequencing technologies allow simultaneous measurement of transcriptome and epigenome profiles in the same cell, providing unprecedented opportunities to dissect cell heterogeneity. Despite great efforts, conjoint analysis of single-cell multi-omics data still suffers from sparsity, high dimensionality and binary. In this study, we present a heterogeneous graph cross-omics attention model (scHGA), a computational tool based on a heterogeneous graph neural network combining two attention mechanisms to jointly analyze single-cell multi-omics data based on different protocols data, including SNARE-seq, scMT-seq and sci-CAR. To avoid the cell heterogeneity of single-omics data, scHGA automatically learns a cell association graph to capture neighbor information. The latent representation of aggregated cells generated by hierarchical attention can fuse knowledge across different omics to dissect cellular heterogeneity, providing a better scheme to characterize the features of cells. scHGA is an effective exploration of graph neural networks in single-cell multi-omics analysis, providing new insights into the understanding of single-cell sequencing data.
Yue Liu 0041, Shulin Wang, Wei Zhang 0089, Xiangxiang Zeng, Chee Keong Kwoh 0001
BIBM5
2022 Deep learning in retrosynthesis planning: datasets, models and tools
abstract
In recent years, synthesizing drugs powered by artificial intelligence has brought great convenience to society. Since retrosynthetic analysis occupies an essential position in synthetic chemistry, it has received broad attention from researchers. In this review, we comprehensively summarize the development process of retrosynthesis in the context of deep learning. This review covers all aspects of retrosynthesis, including datasets, models and tools. Specifically, we report representative models from academia, in addition to a detailed description of the available and stable platforms in the industry. We also discuss the disadvantages of the existing models and provide potential future trends, so that more abecedarians will quickly understand and participate in the family of retrosynthesis planning.
Jingxin Dong 0002, Mingyi Zhao, Yuansheng Liu, Yansen Su, Xiangxiang Zeng
Briefings Bioinform.5
2022 Are dropout imputation methods for scRNA-seq effective for scATAC-seq data?
abstract
The tremendous progress of single-cell sequencing technology has given researchers the opportunity to study cell development and differentiation processes at single-cell resolution. Assay of Transposase-Accessible Chromatin by deep sequencing (ATAC-seq) was proposed for genome-wide analysis of chromatin accessibility. Due to technical limitations or other reasons, dropout events are almost a common occurrence for extremely sparse single-cell ATAC-seq data, leading to confusion in downstream analysis (such as clustering). Although considerable progress has been made in the estimation of scRNA-seq data, there is currently no specific method for the inference of dropout events in single-cell ATAC-seq data. In this paper, we select several state-of-the-art scRNA-seq imputation methods (including MAGIC, SAVER, scImpute, deepImpute, PRIME, bayNorm and knn-smoothing) in recent years to infer dropout peaks in scATAC-seq data, and perform a systematic evaluation of these methods through several downstream analyses. Specifically, we benchmarked these methods in terms of correlation with meta-cell, clustering, subpopulations distance analysis, imputation performance for corruption datasets, identification of TF motifs and computation time. The experimental results indicated that most of the imputed peaks increased the correlation with the reference meta-cell, while the performance of different methods on different datasets varied greatly in different downstream analyses, thus should be used with caution. In general, MAGIC performed better than the other methods most consistently across all assessments. Our source code is freely available at https://github.com/yueyueliu/scATAC-master.
Yue Liu 0041, Shu-Lin Wang, Xiangxiang Zeng, Wei Zhang 0089
Briefings Bioinform.4
2022 A weighted bilinear neural collaborative filtering approach for drug repositioning
abstract
Drug repositioning is an efficient and promising strategy for traditional drug discovery and development. Many research efforts are focused on utilizing deep-learning approaches based on a heterogeneous network for modeling complex drug-disease associations. Similar to traditional latent factor models, which directly factorize drug-disease associations, they assume the neighbors are independent of each other in the network and thus tend to be ineffective to capture localized information. In this study, we propose a novel neighborhood and neighborhood interaction-based neural collaborative filtering approach (called DRWBNCF) to infer novel potential drugs for diseases. Specifically, we first construct three networks, including the known drug-disease association network, the drug-drug similarity and disease-disease similarity networks (using the nearest neighbors). To take the advantage of localized information in the three networks, we then design an integration component by proposing a new weighted bilinear graph convolution operation to integrate the information of the known drug-disease association, the drug's and disease's neighborhood and neighborhood interactions into a unified representation. Lastly, we introduce a prediction component, which utilizes the multi-layer perceptron optimized by the α-balanced focal loss function and graph regularization to model the complex drug-disease associations. Benchmarking comparisons on three datasets verified the effectiveness of DRWBNCF for drug repositioning. Importantly, the unknown drug-disease associations predicted by DRWBNCF were validated against clinical trials and three authoritative databases and we listed several new DRWBNCF-predicted potential drugs for breast cancer (e.g. valrubicin and teniposide) and small cell lung cancer (e.g. valrubicin and cytarabine).
Yajie Meng, Changcheng Lu, Min Jin 0002, Junlin Xu, Xiangxiang Zeng, Jialiang Yang
Briefings Bioinform.5
2022 Learning spatial structures of proteins improves protein-protein interaction prediction
abstract
Spatial structures of proteins are closely related to protein functions. Integrating protein structures improves the performance of protein-protein interaction (PPI) prediction. However, the limited quantity of known protein structures restricts the application of structure-based prediction methods. Utilizing the predicted protein structure information is a promising method to improve the performance of sequence-based prediction methods. We propose a novel end-to-end framework, TAGPPI, to predict PPIs using protein sequence alone. TAGPPI extracts multi-dimensional features by employing 1D convolution operation on protein sequences and graph learning method on contact maps constructed from AlphaFold. A contact map contains abundant spatial structure information, which is difficult to obtain from 1D sequence data directly. We further demonstrate that the spatial information learned from contact maps improves the ability of TAGPPI in PPI prediction tasks. We compare the performance of TAGPPI with those of nine state-of-the-art sequence-based methods, and TAGPPI outperforms such methods in all metrics. To the best of our knowledge, this is the first method to use the predicted protein topology structure graph for sequence-based PPI prediction. More importantly, our proposed architecture could be extended to other prediction tasks related to proteins.
Bosheng Song, Xiaoyan Luo, Xiaoli Luo, Yuansheng Liu, Zhangming Niu, Xiangxiang Zeng
Briefings Bioinform.6
2022 Deep learning joint models for extracting entities and relations in biomedical: a survey and comparison
abstract
The rapid development of biomedicine has produced a large number of biomedical written materials. These unstructured text data create serious challenges for biomedical researchers to find information. Biomedical named entity recognition (BioNER) and biomedical relation extraction (BioRE) are the two most fundamental tasks of biomedical text mining. Accurately and efficiently identifying entities and extracting relations have become very important. Methods that perform two tasks separately are called pipeline models, and they have shortcomings such as insufficient interaction, low extraction quality and easy redundancy. To overcome the above shortcomings, many deep learning-based joint name entity recognition and relation extraction models have been proposed, and they have achieved advanced performance. This paper comprehensively summarize deep learning models for joint name entity recognition and relation extraction for biomedicine. The joint BioNER and BioRE models are discussed in the light of the challenges existing in the BioNER and BioRE tasks. Five joint BioNER and BioRE models and one pipeline model are selected for comparative experiments on four biomedical public datasets, and the experimental results are analyzed. Finally, we discuss the opportunities for future development of deep learning-based joint BioNER and BioRE models.
Yansen Su, Minglu Wang, Pengpeng Wang, Chun-Hou Zheng 0001, Yuansheng Liu, Xiangxiang Zeng
Briefings Bioinform.6
2022 preMLI: a pre-trained method to uncover microRNA-lncRNA potential interactions
abstract
The interaction between microribonucleic acid and long non-coding ribonucleic acid plays a very important role in biological processes, and the prediction of the one is of great significance to the study of its mechanism of action. Due to the limitations of traditional biological experiment methods, more and more computational methods are applied to this field. However, the existing methods often have problems, such as inadequate acquisition of potential features of the sequence due to simple coding and the need to manually extract features as input. We propose a deep learning model, preMLI, based on rna2vec pre-training and deep feature mining mechanism. We use rna2vec to train the ribonucleic acid (RNA) dataset and to obtain the RNA word vector representation and then mine the RNA sequence features separately and finally concatenate the two feature vectors as the input of the prediction task. The preMLI performs better than existing methods on benchmark datasets and has cross-species prediction capabilities. Experiments show that both pre-training and deep feature mining mechanisms have a positive impact on the prediction performance of the model. To be more specific, pre-training can provide more accurate word vector representations. The deep feature mining mechanism also improves the prediction performance of the model. Meanwhile, The preMLI only needs RNA sequence as the input of the model and has better cross-species prediction performance than the most advanced prediction models, which have reference value for related research.
Likun Jiang, Shuting Jin, Xiangxiang Zeng, Xiangrong Liu
Briefings Bioinform.4
2022 MLysPRED: graph-based multi-view clustering and multi-dimensional normal distribution resampling techniques to predict multiple lysine sites
abstract
Posttranslational modification of lysine residues, K-PTM, is one of the most popular PTMs. Some lysine residues in proteins can be continuously or cascaded covalently modified, such as acetylation, crotonylation, methylation and succinylation modification. The covalent modification of lysine residues may have some special functions in basic research and drug development. Although many computational methods have been developed to predict lysine PTMs, up to now, the K-PTM prediction methods have been modeled and learned a single class of K-PTM modification. In view of this, this study aims to fill this gap by building a multi-label computational model that can be directly used to predict multiple K-PTMs in proteins. In this study, a multi-label prediction model, MLysPRED, is proposed to identify multiple lysine sites using features generated from human protein sequences. In MLysPRED, three kinds of multi-label sequence encoding algorithms (MLDBPB, MLPSDAAP, MLPSTAAP) are proposed and combined with three encoding strategies (CHHAA, DR and Kmer) to convert preprocessed lysine sequences into effective numerical features. A multidimensional normal distribution oversampling technique and graph-based multi-view clustering under-sampling algorithm were first proposed and incorporated to reduce the proportion of the original training samples, and multi-label nearest neighbor algorithm is used for classification. It is observed that MLysPRED achieved an Aiming of 92.21%, Coverage of 94.98%, Accuracy of 89.63%, Absolute-True of 81.46% and Absolute-False of 0.0682 on the independent datasets. Additionally, comparison of results with five existing predictors also indicated that MLysPRED is very promising and encouraging to predict multiple K-PTMs in proteins. For the convenience of the experimental scientists, 'MLysPRED' has been deployed as a user-friendly web-server at http://47.100.136.41:8181.
Yun Zuo 0001, Xiangxiang Zeng, Qiang Zhang 0026, Xiangrong Liu
Briefings Bioinform.3
2022 On the cross-migrativity between uninorms and overlap (grouping) functions
Kuanyun Zhu, Xiangxiang Zeng, Junsheng Qiao
Fuzzy Sets Syst.2
2022 Rule synchronization for monodirectional tissue-like P systems with channel states
Bosheng Song, Xiangxiang Zeng
Inf. Comput.3
2022 Normal forms for spiking neural P systems and some of its variants
Ivan Cedric H. Macababayao, Francis George Cabarle, Ren Tristan A. de la Cruz, Xiangxiang Zeng
Inf. Sci.4
2022 A Robust Algorithm Based on Link Label Propagation for Identifying Functional Modules From Protein-Protein Interaction Networks
abstract
Identifying functional modules in protein-protein interaction (PPI) networks elucidates cellular organization and mechanism. Various methods have been proposed to identify the functional modules in PPI networks, but most of these methods do not consider the noisy links in PPI networks. They achieve a competitive performance on the PPI networks without noisy links, but the performance of these methods considerably deteriorates in the noisy PPI networks. Furthermore, the noisy links are inevitable in the PPI networks. In this paper, we propose a novel link-driven label propagation algorithm (LLPA) to identify functional modules in PPI networks. The LLPA first find link clusters in PPI networks, and then the functional modules are identified from the link clusters. Two strategies aimed to ensure the robustness of LLPA are proposed. One strategy involves the proposed LLPA updating the link labels in accordance with the designed weight of the link, which can reduce the incidence of noisy links. The other strategy involves the filtration of some noisy labels from the link clusters to further reduce the influence of noisy links. The performance evaluation on three real PPI networks shows that LLPA outperforms other eight state-of-the-art detection algorithms in terms of accuracy and robustness.
Hao Jiang 0023, Fei Zhan, Congtao Wang, Jianfeng Qiu, Yansen Su, Chun-Hou Zheng 0001, Xingyi Zhang 0001, Xiangxiang Zeng
IEEE ACM Trans. Comput. Biol. Bioinform.8
2022 3DMol-Net: Learn 3D Molecular Representation Using Adaptive Graph Convolutional Network Based on Rotation Invariance
abstract
Studying the deep learning-based molecular representation has great significance on predicting molecular property, promoted the development of drug screening and new drug discovery, and improving human well-being for avoiding illnesses. It is essential to learn the characterization of drug for various downstream tasks, such as molecular property prediction. In particular, the 3D structure features of molecules play an important role in biochemical function and activity prediction. The 3D characteristics of molecules largely determine the properties of the drug and the binding characteristics of the target. However, most current methods merely rely on 1D or 2D properties while ignoring the 3D topological structure, thereby degrading the performance of molecular inferring. In this paper, we propose 3DMol-Net to enhance the molecular representation, considering both the topology and rotation invariance (RI) of the 3D molecular structure. Specifically, we construct a molecular graph with soft relations related to the spatial arrangement of the 3D coordinates to learn 3D topology of arbitrary graph structure and employ an adaptive graph convolutional network to predict molecular properties and biochemical activities. Comparing with current graph-based methods, 3DMol-Net demonstrates superior performance in terms of both regression and classification tasks. Further verification of RI and visualization also show better robustness and representation capacity of our model.
Chunyan Li 0002, Wei Wei 0006, Jin Li 0007, Junfeng Yao, Xiangxiang Zeng, Zhihan Lyu
IEEE J. Biomed. Health Informatics5
2022 Monodirectional Evolutional Symport Tissue P Systems With Promoters and Cell Division
abstract
Monodirectional tissue P systems with promoters are natural inspired parallel computing paradigms, where only symport rules are permitted, and with the restriction of “monodirectionality”, objects for two given regions are transferred in one direction. In this article, a novel kind of P systems, monodirectional evolutional symport tissue P systems with promoters (MESTP P systems) is raised, where objects may be revised during the movement between two regions. The computational theory of MESTP P systems that rules are employed in a flat maximally parallel pattern is investigated. We prove that finite natural number sets are created by MESTP P systems applying one cell, at most 1 promoter and all evolutional symport rules having a maximal length 2 or with arbitrary number of cells, promoters and all evolutional symport rules having a maximal length 2. MESTP P systems are Turing universal when two cells, at most 1 promoter and all evolutional symport rules having a maximal length 2 are employed. In addition, with the help of cell division mechanism, monodirectional evolutional symport tissue P systems with promoters and cell division (MESTPD P systems) are employed to solve NP-complete (the SAT) problem, where system uses at most 1 promoter and all evolutional symport rules having a maximal length 3. These results show that MESTP(D) P systems are still computationally powerful even if monodirectionality control mechanism is imposed, thereby developing membrane algorithms for MESTP(D) P systems is theoretically possible as well as potentially exploitable.
Bosheng Song, Kenli Li 0001, Xiangxiang Zeng
IEEE Trans. Parallel Distributed Syst.3
2021 LADstackING: Stacking Ensemble Learning-based Computational Model for Predicting Potential LncRNA-disease Associations
abstract
In recent years, accumulation of researches have proved many diseases that seriously endanger human health originate from mutations or dysfunctions in LncRNA (Long non-encoding RNA). Therefore, it is important to discover the intrinsic associations between the LncRNAs and the diseases. Meantime, accurately identifying potential associations between diseases and LncRNAs remains a highly challenging task. In this paper, we proposed a model based on the stacking ensemble learning framework called LADstackING to predict the potential LncRNA associated disease. LADstackING effectively integrates different types of strong predictive performance models rather than the same type models with weak predictive performance.LADstackING is able to exploit the respective advantages of different base models in its framework and significantly improve the overall predictive performance. The multi-perspective features bring by different base models allow LADstackING remain stable in facing of sparse data sources. Moreover, the overall predictive performance of LADstackING is greatly improved compare to the stat-of-art models. Experimental results and case study result demonstrate that LADstackING performs promising in predicting the potential LncRNA-disease associations.
Jiechen Li, Xiangxiang Zeng, Yong Dou, Fei Xia 0003, Shaoliang Peng
BIBM2
2021 Pm6 A: an Integrated Classification Algorithm for 2021 IEEE International Conference on Bioinformatics and Biomedicine (BIBM) Identifying m6 A Sites
abstract
As a major RNA methylation modification, $\mathrm{N}^{6}_{-}$ methyladenosine (m6A) affects the occurrence and development of various human cancers through a variety of mechanisms. It has been reported that m6A RNA methylation involves different physiological and pathological processes. Therefore, the detection of m6A is helpful to reveal its biological function. Due to the high cost and time-consuming of high-throughput sequencing and the inaccurate sites identified, computational tools are needed to guide the accurate prediction of m6A modified sites and help reduce the costs associated with high-throughput sequencing. In this study, an integrated classification algorithm, Pm6A, is proposed to identify S. cerevisiae m6A sites using features generated from RNA sequences. In Pm6A, six sequence encoding schemes (pseudo dinucleotide composition, dinucleotide-based auto covariance, dinucleotide-based cross covariance, dinucleotide-based auto-cross covariance, mismatch and subsequence) are used for feature extraction, the VIBES ensemble algorithm based on four base classifiers, namely nearest neighbor, support vector machine, discriminant analysis and artificial neural network is used for classification, the optimized forward search algorithm is used to find the optimal parameters of model. The results of 10-fold cross-validation show that the proposed approach achieved better specificity (at Dataset 1) and better accuracy (at Dataset 2) than other methods. It is expected that Pm6A will be a useful tool for predicting m6A sites.
Yun Zuo 0001, Xiangrong Liu, Xiangxiang Zeng
BIBM3
2021 Fast k-NN Graph Construction by GPU based NN-Descent
abstract
NN-Descent is a classic k-NN graph construction approach. It is still widely employed in machine learning, computer vision, and information retrieval tasks due to its efficiency and genericness. However, the current design only works well on CPU. In this paper, NN-Descent has been redesigned to adapt to the GPU architecture. A new graph update strategy called selective update is proposed. It reduces the data exchange between GPU cores and GPU global memory significantly, which is the processing bottleneck under GPU computation architecture. This redesign leads to full exploitation of the parallelism of the GPU hardware. In the meantime, the genericness, as well as the simplicity of NN-Descent, are well-preserved. Moreover, a procedure that allows to k-NN graph to be merged efficiently on GPU is proposed. It makes the construction of high-quality k-NN graphs for out-of-GPU-memory datasets tractable. Our approach is 100-250× faster than the single-thread NN-Descent and is 2.5-5× faster than the existing GPU-based approaches as we tested on million as well as billion scale datasets.
Wanlei Zhao, Xiangxiang Zeng, Jianye Yang 0001
CIKM3
2021 Drug repositioning based on the heterogeneous information fusion graph convolutional network
abstract
In silico reuse of old drugs (also known as drug repositioning) to treat common and rare diseases is increasingly becoming an attractive proposition because it involves the use of de-risked drugs, with potentially lower overall development costs and shorter development timelines. Therefore, there is a pressing need for computational drug repurposing methodologies to facilitate drug discovery. In this study, we propose a new method, called DRHGCN (Drug Repositioning based on the Heterogeneous information fusion Graph Convolutional Network), to discover potential drugs for a certain disease. To make full use of different topology information in different domains (i.e. drug-drug similarity, disease-disease similarity and drug-disease association networks), we first design inter- and intra-domain feature extraction modules by applying graph convolution operations to the networks to learn the embedding of drugs and diseases, instead of simply integrating the three networks into a heterogeneous network. Afterwards, we parallelly fuse the inter- and intra-domain embeddings to obtain the more representative embeddings of drug and disease. Lastly, we introduce a layer attention mechanism to combine embeddings from multiple graph convolution layers for further improving the prediction performance. We find that DRHGCN achieves high performance (the average AUROC is 0.934 and the average AUPR is 0.539) in four benchmark datasets, outperforming the current approaches. Importantly, we conducted molecular docking experiments on DRHGCN-predicted candidate drugs, providing several novel approved drugs for Alzheimer's disease (e.g. benzatropine) and Parkinson's disease (e.g. trihexyphenidyl and haloperidol).
Changcheng Lu, Junlin Xu, Yajie Meng, Peng Wang 0035, Xiangzheng Fu, Xiangxiang Zeng, Yansen Su
Briefings Bioinform.7
2021 ITP-Pred: an interpretable method for predicting, therapeutic peptides with fused features low-dimension representation
abstract
The peptide therapeutics market is providing new opportunities for the biotechnology and pharmaceutical industries. Therefore, identifying therapeutic peptides and exploring their properties are important. Although several studies have proposed different machine learning methods to predict peptides as being therapeutic peptides, most do not explain the decision factors of model in detail. In this work, an Interpretable Therapeutic Peptide Prediction (ITP-Pred) model based on efficient feature fusion was developed. First, we proposed three kinds of feature descriptors based on sequence and physicochemical property encoded, namely amino acid composition (AAC), group AAC and coding autocorrelation, and concatenated them to obtain the feature representation of therapeutic peptide. Then, we input it into the CNN-Bi-directional Long Short-Term Memory (BiLSTM) model to automatically learn recognition of therapeutic peptides. The cross-validation and independent verification experiments results indicated that ITP-Pred has a higher prediction performance on the benchmark dataset than other comparison methods. Finally, we analyzed the output of the model from two aspects: sequence order and physical and chemical properties, mining important features as guidance for the design of better models that can complement existing methods.
Li Wang 0145, Xiangzheng Fu, Chenxing Xia, Xiangxiang Zeng, Quan Zou 0001
Briefings Bioinform.5
2021 Application of deep learning methods in biological networks
abstract
The increase in biological data and the formation of various biomolecule interaction databases enable us to obtain diverse biological networks. These biological networks provide a wealth of raw materials for further understanding of biological systems, the discovery of complex diseases and the search for therapeutic drugs. However, the increase in data also increases the difficulty of biological networks analysis. Therefore, algorithms that can handle large, heterogeneous and complex data are needed to better analyze the data of these network structures and mine their useful information. Deep learning is a branch of machine learning that extracts more abstract features from a larger set of training data. Through the establishment of an artificial neural network with a network hierarchy structure, deep learning can extract and screen the input information layer by layer and has representation learning ability. The improved deep learning algorithm can be used to process complex and heterogeneous graph data structures and is increasingly being applied to the mining of network data information. In this paper, we first introduce the used network data deep learning models. After words, we summarize the application of deep learning on biological networks. Finally, we discuss the future development prospects of this field.
Shuting Jin, Xiangxiang Zeng, Feng Xia 0007, Xiangrong Liu
Briefings Bioinform.2
2021 A spatial-temporal gated attention module for molecular property prediction based on molecular geometry
abstract
MOTIVATION: Geometry-based properties and characteristics of drug molecules play an important role in drug development for virtual screening in computational chemistry. The 3D characteristics of molecules largely determine the properties of the drug and the binding characteristics of the target. However, most of the previous studies focused on 1D or 2D molecular descriptors while ignoring the 3D topological structure, thereby degrading the performance of molecule-related prediction. Because it is very time-consuming to use dynamics to simulate molecular 3D conformer, we aim to use machine learning to represent 3D molecules by using the generated 3D molecular coordinates from the 2D structure. RESULTS: We proposed Drug3D-Net, a novel deep neural network architecture based on the spatial geometric structure of molecules for predicting molecular properties. It is grid-based 3D convolutional neural network with spatial-temporal gated attention module, which can extract the geometric features for molecular prediction tasks in the process of convolution. The effectiveness of Drug3D-Net is verified on the public molecular datasets. Compared with other deep learning methods, Drug3D-Net shows superior performance in predicting molecular properties and biochemical activities. AVAILABILITY AND IMPLEMENTATION: https://github.com/anny0316/Drug3D-Net. SUPPLEMENTARY DATA: Supplementary data are available online at https://academic.oup.com/bib.
Chunyan Li 0002, Jianmin Wang 0016, Zhangming Niu, Junfeng Yao, Xiangxiang Zeng
Briefings Bioinform.5
2021 De novo generation of dual-target ligands using adversarial training and reinforcement learning
abstract
Artificial intelligence, such as deep generative methods, represents a promising solution to de novo design of molecules with the desired properties. However, generating new molecules with biological activities toward two specific targets remains an extremely difficult challenge. In this work, we conceive a novel computational framework, herein called dual-target ligand generative network (DLGN), for the de novo generation of bioactive molecules toward two given objectives. Via adversarial training and reinforcement learning, DLGN treats a sequence-based simplified molecular input line entry system (SMILES) generator as a stochastic policy for exploring chemical spaces. Two discriminators are then used to encourage the generation of molecules that belong to the intersection of two bioactive-compound distributions. In a case study, we employ our methods to design a library of dual-target ligands targeting dopamine receptor D2 and 5-hydroxytryptamine receptor 1A as new antipsychotics. Experimental results demonstrate that the proposed model can generate novel compounds with high similarity to both bioactive datasets in several structure-based metrics. Our model exhibits a performance comparable to that of various state-of-the-art multi-objective molecule generation models. We envision that this framework will become a generally applicable approach for designing dual-target drugs in silico.
Fengqing Lu, Mufei Li, Xiaoping Min, Chunyan Li 0002, Xiangxiang Zeng
Briefings Bioinform.5
2021 Predicting enhancer-promoter interactions by deep learning and matching heuristic
abstract
Enhancer-promoter interactions (EPIs) play an important role in transcriptional regulation. Recently, machine learning-based methods have been widely used in the genome-scale identification of EPIs due to their promising predictive performance. In this paper, we propose a novel method, termed EPI-DLMH, for predicting EPIs with the use of DNA sequences only. EPI-DLMH consists of three major steps. First, a two-layer convolutional neural network is used to learn local features, and an bidirectional gated recurrent unit network is used to capture long-range dependencies on the sequences of promoters and enhancers. Second, an attention mechanism is used for focusing on relatively important features. Finally, a matching heuristic mechanism is introduced for the exploration of the interaction between enhancers and promoters. We use benchmark datasets in evaluating and comparing the proposed method with existing methods. Comparative results show that our model is superior to currently existing models in multiple cell lines. Specifically, we found that the matching heuristic mechanism introduced into the proposed model mainly contributes to the improvement of performance in terms of overall accuracy. Additionally, compared with existing models, our model is more efficient with regard to computational speed.
Xiaoping Min, Congmin Ye, Xiangrong Liu, Xiangxiang Zeng
Briefings Bioinform.4
2021 Deep learning methods for biomedical named entity recognition: a survey and qualitative comparison
abstract
The biomedical literature is growing rapidly, and the extraction of meaningful information from the large amount of literature is increasingly important. Biomedical named entity (BioNE) identification is one of the critical and fundamental tasks in biomedical text mining. Accurate identification of entities in the literature facilitates the performance of other tasks. Given that an end-to-end neural network can automatically extract features, several deep learning-based methods have been proposed for BioNE recognition (BioNER), yielding state-of-the-art performance. In this review, we comprehensively summarize deep learning-based methods for BioNER and datasets used in training and testing. The deep learning methods are classified into four categories: single neural network-based, multitask learning-based, transfer learning-based and hybrid model-based methods. They can be applied to BioNER in multiple domains, and the results are determined by the dataset size and type. Lastly, we discuss the future development and opportunities of BioNER methods.
Bosheng Song, Fen Li, Yuansheng Liu, Xiangxiang Zeng
Briefings Bioinform.4
2021 A novel antibacterial peptide recognition algorithm based on BERT
abstract
As the best substitute for antibiotics, antimicrobial peptides (AMPs) have important research significance. Due to the high cost and difficulty of experimental methods for identifying AMPs, more and more researches are focused on using computational methods to solve this problem. Most of the existing calculation methods can identify AMPs through the sequence itself, but there is still room for improvement in recognition accuracy, and there is a problem that the constructed model cannot be universal in each dataset. The pre-training strategy has been applied to many tasks in natural language processing (NLP) and has achieved gratifying results. It also has great application prospects in the field of AMP recognition and prediction. In this paper, we apply the pre-training strategy to the model training of AMP classifiers and propose a novel recognition algorithm. Our model is constructed based on the BERT model, pre-trained with the protein data from UniProt, and then fine-tuned and evaluated on six AMP datasets with large differences. Our model is superior to the existing methods and achieves the goal of accurate identification of datasets with small sample size. We try different word segmentation methods for peptide chains and prove the influence of pre-training steps and balancing datasets on the recognition effect. We find that pre-training on a large number of diverse AMP data, followed by fine-tuning on new data, is beneficial for capturing both new data's specific features and common features between AMP sequences. Finally, we construct a new AMP dataset, on which we train a general AMP recognition model.
Jianyuan Lin, Lianmin Zhao, Xiangxiang Zeng, Xiangrong Liu
Briefings Bioinform.4
2021 iEnhancer-XG: interpretable sequence-based enhancers and their strength predictor
abstract
MOTIVATION: Enhancers are non-coding DNA fragments with high position variability and free scattering. They play an important role in controlling gene expression. As machine learning has become more widely used in identifying enhancers, a number of bioinformatic tools have been developed. Although several models for identifying enhancers and their strengths have been proposed, their accuracy and efficiency have yet to be improved. RESULTS: We propose a two-layer predictor called 'iEnhancer-XG.' It comprises a one-layer predictor (for identifying enhancers) and a second classifier (for their strength) and uses 'XGBoost' as a base classifier and five feature extraction methods, namely, k-Spectrum Profile, Mismatch k-tuple, Subsequence Profile, Position-specific scoring matrix (PSSM) and Pseudo dinucleotide composition (PseDNC). Each method has an independent output. We place the feature vector matrix into the ensemble learning for fusion. This experiment involves the method of 'SHapley Additive explanations' to provide interpretability for the previous black box machine learning methods and improve their credibility. The accuracies of the ensemble learning method are 0.811 (first layer) and 0.657 (second layer). The rigorous 10-fold cross-validation confirms that the proposed method is significantly better than existing technologies. AVAILABILITY AND IMPLEMENTATION: The source code and dataset for the enhancer predictions have been uploaded to https://github.com/jimmyrate/ienhancer-xg. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xuanbai Ren, Xiangzheng Fu, Mingyu Gao 0004, Xiangxiang Zeng
Bioinform.6
2021 MUFFIN: multi-scale feature fusion for drug-drug interaction prediction
abstract
MOTIVATION: Adverse drug-drug interactions (DDIs) are crucial for drug research and mainly cause morbidity and mortality. Thus, the identification of potential DDIs is essential for doctors, patients and the society. Existing traditional machine learning models rely heavily on handcraft features and lack generalization. Recently, the deep learning approaches that can automatically learn drug features from the molecular graph or drug-related network have improved the ability of computational models to predict unknown DDIs. However, previous works utilized large labeled data and merely considered the structure or sequence information of drugs without considering the relations or topological information between drug and other biomedical objects (e.g. gene, disease and pathway), or considered knowledge graph (KG) without considering the information from the drug molecular structure. RESULTS: Accordingly, to effectively explore the joint effect of drug molecular structure and semantic information of drugs in knowledge graph for DDI prediction, we propose a multi-scale feature fusion deep learning model named MUFFIN. MUFFIN can jointly learn the drug representation based on both the drug-self structure information and the KG with rich bio-medical information. In MUFFIN, we designed a bi-level cross strategy that includes cross- and scalar-level components to fuse multi-modal features well. MUFFIN can alleviate the restriction of limited labeled data on deep learning models by crossing the features learned from large-scale KG and drug molecular graph. We evaluated our approach on three datasets and three different tasks including binary-class, multi-class and multi-label DDI prediction tasks. The results showed that MUFFIN outperformed other state-of-the-art baselines. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/xzenglab/MUFFIN.
Yujie Chen 0002, Tengfei Ma 0002, Xixi Yang, Jianmin Wang 0016, Bosheng Song, Xiangxiang Zeng
Bioinform.6
2021 Minirmd: accurate and fast duplicate removal tool for short reads via multiple minimizers
abstract
SUMMARY: Removing duplicate and near-duplicate reads, generated by high-throughput sequencing technologies, is able to reduce computational resources in downstream applications. Here we develop minirmd, a de novo tool to remove duplicate reads via multiple rounds of clustering using different length of minimizer. Experiments demonstrate that minirmd removes more near-duplicate reads than existing clustering approaches and is faster than existing multi-core tools. To the best of our knowledge, minirmd is the first tool to remove near-duplicates on reverse-complementary strand. AVAILABILITY AND IMPLEMENTATION: https://github.com/yuansliu/minirmd. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yuansheng Liu, Xiaocai Zhang, Quan Zou 0001, Xiangxiang Zeng
Bioinform.4
2021 CarSite-II: an integrated classification algorithm for identifying carbonylated sites based on K-means similarity-based undersampling and synthetic minority oversampling techniques
abstract
BACKGROUND: Carbonylation is a non-enzymatic irreversible protein post-translational modification, and refers to the side chain of amino acid residues being attacked by reactive oxygen species and finally converted into carbonyl products. Studies have shown that protein carbonylation caused by reactive oxygen species is involved in the etiology and pathophysiological processes of aging, neurodegenerative diseases, inflammation, diabetes, amyotrophic lateral sclerosis, Huntington's disease, and tumor. Current experimental approaches used to predict carbonylation sites are expensive, time-consuming, and limited in protein processing abilities. Computational prediction of the carbonylation residue location in protein post-translational modifications enhances the functional characterization of proteins. RESULTS: In this study, an integrated classifier algorithm, CarSite-II, was developed to identify K, P, R, and T carbonylated sites. The resampling method K-means similarity-based undersampling and the synthetic minority oversampling technique (SMOTE-KSU) were incorporated to balance the proportions of K, P, R, and T carbonylated training samples. Next, the integrated classifier system Rotation Forest uses "support vector machine" subclassifications to divide three types of feature spaces into several subsets. CarSite-II gained Matthew's correlation coefficient (MCC) values of 0.2287/0.3125/0.2787/0.2814, False Positive rate values of 0.2628/0.1084/0.1383/0.1313, False Negative rate values of 0.2252/0.0205/0.0976/0.0608 for K/P/R/T carbonylation sites by tenfold cross-validation, respectively. On our independent test dataset, CarSite-II yield MCC values of 0.6358/0.2910/0.4629/0.3685, False Positive rate values of 0.0165/0.0203/0.0188/0.0094, False Negative rate values of 0.1026/0.1875/0.2037/0.3333 for K/P/R/T carbonylation sites. The results show that CarSite-II achieves remarkably better performance than all currently available prediction tools. CONCLUSION: The related results revealed that CarSite-II achieved better performance than the currently available five programs, and revealed the usefulness of the SMOTE-KSU resampling approach and integration algorithm. For the convenience of experimental scientists, the web tool of CarSite-II is available in http://47.100.136.41:8081/.
Yun Zuo 0001, Jianyuan Lin, Xiangxiang Zeng, Quan Zou 0001, Xiangrong Liu
BMC Bioinform.3
2021 Neural-like P systems with plasmids
Francis George Cabarle, Xiangxiang Zeng, Niall Murphy, Tao Song 0001, Alfonso Rodríguez-Patón, Xiangrong Liu
Inf. Comput.2
2021 The computational power of monodirectional tissue P systems with symport rules
Bosheng Song, Shengye Huang, Xiangxiang Zeng
Inf. Comput.3
2021 Monodirectional tissue P systems with channel states
Bosheng Song, Xiangxiang Zeng, Alfonso Rodríguez-Patón
Inf. Sci.2
2021 Monodirectional Tissue P Systems With Promoters
abstract
Tissue P systems with promoters provide nondeterministic parallel bioinspired devices that evolve by the interchange of objects between regions, determined by the existence of some special objects called promoters. However, in cellular biology, the movement of molecules across a membrane is transported from high to low concentration. Inspired by this biological fact, in this article, an interesting type of tissue P systems, called monodirectional tissue P systems with promoters, where communication happens between two regions only in one direction, is considered. Results show that finite sets of numbers are produced by such P systems with one cell, using any length of symport rules or with any number of cells, using a maximal length 1 of symport rules, and working in the maximally parallel mode. Monodirectional tissue P systems are Turing universal with two cells, a maximal length 2, and at most one promoter for each symport rule, and working in the maximally parallel mode or with three cells, a maximal length 1, and at most one promoter for each symport rule, and working in the flat maximally parallel mode. We also prove that monodirectional tissue P systems with two cells, a maximal length 1, and at most one promoter for each symport rule (under certain restrictive conditions) working in the flat maximally parallel mode characterizes regular sets of natural numbers. Besides, the computational efficiency of monodirectional tissue P systems with promoters is analyzed when cell division rules are incorporated. Different uniform solutions to the Boolean satisfiability problem (SAT problem) are provided. These results show that with the restrictive condition of "monodirectionality," monodirectional tissue P systems with promoters are still computationally powerful. With the powerful computational power, developing membrane algorithms for monodirectional tissue P systems with promoters is potentially exploitable.
Bosheng Song, Xiangxiang Zeng, Min Jiang 0005, Mario J. Pérez-Jiménez
IEEE Trans. Cybern.2
2021 A Polar-Metric-Based Evolutionary Algorithm
abstract
Over the past two decades, numerous multi- and many-objective evolutionary algorithms (MOEAs and MaOEAs) have been proposed to solve the multi- and many-objective optimization problems (MOPs and MaOPs), respectively. It is known that the difficulty of maintaining the convergence and diversity performances rapidly grows as the number of objectives increases. This phenomenon is especially evident for the Pareto-dominance-based EAs, because the nondominated sorting often fails to provide enough convergent pressure toward the Pareto front (PF). Therefore, many researchers came up with some non-Pareto-dominance-based EAs, which are based on indicator, decomposition, and so on. In this article, we propose a polar-metric ( p -metric)-based EA (PMEA) for tackling both MOPs and MaOPs. p -metric is a recently proposed performance indicator which adopts a set of uniformly distributed direction vectors. In PMEA, we use a two-phase selection which combines both nondominated sorting and p -metric. Moreover, a modification is proposed to adjust the direction vectors of p -metric dynamically. In the experiments, PMEA is compared with six state-of-the-art EAs in total and is measured by three performance metrics, including p -metric. According to the empirical results, PMEA shows promising performances on most of the test problems, involving both MOPs and MaOPs.
Hang Xu 0003, Wenhua Zeng, Xiangxiang Zeng, Gary G. Yen
IEEE Trans. Cybern.3
2021 Mobility Based Trust Evaluation for Heterogeneous Electric Vehicles Network in Smart Cities
abstract
Smart cities can manage assets and resources efficiently by using different types of electronic data collection sensors, devices and vehicles. However, growing complexity of systems and heterogeneous networking also enlarge the destructive effect of compromised or malicious sensor nodes. In this paper, we introduce electric vehicles to conduct trust evaluation for heterogeneous vehicle network in smart cities. Compared with traditional trust evaluation mechanism, mobility-based trust evaluation owns the advantages of low energy consumption and high evaluation accuracy. Meanwhile, we investigate the problem of minimizing transmission hops of trust evaluation and refers to this as the mobile trust evaluation problem (MTEP). We first formalize the MTEP into an optimization problem and present a heuristic moving strategy of single electric vehicle. Then, we consider the MTEP with multiple electric vehicles. By scheduling the electric vehicles to access the nodes on spanning tree with maximum neighbor distance ratio, the algorithm can improve the efficiency of trust evaluation. In experiments, we compare moving strategy of single electric vehicle and multiple electric vehicles with existing methods respectively. The results demonstrate that the proposed algorithms are able to effectively reduce the entire transmission hops of trust evaluation and thus prolong the life of the network.
Tian Wang 0001, Hao Luo 0012, Xiangxiang Zeng, Zhiyong Yu 0001, Anfeng Liu, Arun Kumar Sangaiah
IEEE Trans. Intell. Transp. Syst.3
2020 A multi-task learning method for analyzing microbiota as cancer immunotherapy signal
abstract
Researches have found that tumor immunotherapy can only work for some patients, and the intestinal microbiota is one of the important factors affecting the responses of patients with cancer to immune checkpoint blockade therapy. It is highly desirable to develop computational methods that can predict whether a patient with cancer will have positive effects on cancer immunotherapy by analyzing intestinal microorganisms of the patient. In this study, a multi-task model is introduced to predict the efficacy of cancer immunotherapy on a patient who suffered from non-small cell lung cancer or renal cell carcinoma. The results demonstrate the multi-task model outperforms several single-task methods. Therefore, we believe the multi-task idea can be used to predict the efficacy of cancer immunotherapy based on the gut microbe, which would be important to cancer patients.
Changzhi Jiang, Yousi Fu, Shuting Jin, Xiangrong Liu, Baishan Fang, Xiangxiang Zeng
BIBM6
2020 A drug information embedding method based on graph convolution neural network
abstract
New drug development is an extremely time-consuming and high-risk process. [1]It has been widely valued by the biomedical industry to fully explore the new uses of existing drugs and reorientate them. [2]How to find drug disease with potential therapeutic relationship from a large number of unproven relationship pairs is the research focus of drug reorientation. With the help of machine learning model, we can improve the enrichment degree of potential drug disease relationship pairs, and reduce the false positive rate of prediction. In the past few years, a series of graph based convolutional network models have been developed to calculate the information latent feature representation of nodes and links. Researchers at home and abroad have done a lot of research on network embedding technology based on biomedical data, and have achieved a series of important research results. Among them, the research methods used can be divided into two categories: one is the traditional machine learning algorithm based on artificial feature extraction, the other is the method based on deep learning. For example, kipf and welling [3]proposed a new graph convolution network (GCN) with parts of existing models, DeepDR [4] and DTINet [5] based on node characteristics and their connections, which can be used for node classification. Aiming at the problem of imbalance of drug information data samples, the invention provides a drug relocation method based on deep learning multi-source heterogeneous network. In order to avoid the limitations of traditional feature extraction methods, such as highly dependent on the experience and knowledge of medical staff, strong subjectivity, consuming a lot of time and energy to complete, and extracting high-quality features with distinguishing features often exists In this paper, with the help of graph convolution encoder model and variational auto encoder neural network, we can automatically learn the characteristics of multi-source and heterogeneous drug low-dimensional network, and complete the drug relocation of drug disease association prediction.
Xiaoyi Feng, Shaoliang Peng, Fei Li 0040, Xiangxiang Zeng, Yunhao Liu 0001
HealthCom5
2020 KGNN: Knowledge Graph Neural Network for Drug-Drug Interaction Prediction
abstract
Drug-drug interaction (DDI) prediction is a challenging problem in pharmacology and clinical application, and effectively identifying potential DDIs during clinical trials is critical for patients and society. Most of existing computational models with AI techniques often concentrate on integrating multiple data sources and combining popular embedding methods together. Yet, researchers pay less attention to the potential correlations between drug and other entities such as targets and genes. Moreover, recent studies also adopted knowledge graph (KG) for DDI prediction. Yet, this line of methods learn node latent embedding directly, but they are limited in obtaining the rich neighborhood information of each entity in the KG. To address the above limitations, we propose an end-to-end framework, called Knowledge Graph Neural Network (KGNN), to resolve the DDI prediction. Our framework can effectively capture drug and its potential neighborhoods by mining their associated relations in KG. To extract both high-order structures and semantic relations of the KG, we learn from the neighborhoods for each entity in the KG as their local receptive, and then integrate neighborhood information with bias from representation of the current entity. This way, the receptive field can be naturally extended to multiple hops away to model high-order topological information and to obtain drugs potential long-distance correlations. We have implemented our method and conducted experiments based on several widely-used datasets. Empirical results show that KGNN outperforms the classic and state-of-the-art models.
Xuan Lin, Zhe Quan, Zhi-Jie Wang 0009, Tengfei Ma 0002, Xiangxiang Zeng
IJCAI5
2020 A novel molecular representation with BiGRU neural networks for learning atom
abstract
Molecular representations play critical roles in researching drug design and properties, and effective methods are beneficial to assisting in the calculation of molecules and solving related problem in drug discovery. In previous years, most of the traditional molecular representations are based on hand-crafted features and rely heavily on biological experimentations, which are often costly and time consuming. However, recent researches achieve promising results using machine learning on various domains. In this article, we present a novel method named Smi2Vec-BiGRU that is designed for learning atoms and solving the single- and multitask binary classification problems in the field of drug discovery, which are the basic and also key problems in this field. Specifically, our approach transforms the molecule data in the SMILES format into a set of sample vectors and then feeds them into the bidirectional gated recurrent unit neural networks for training, which learns low-dimensional vector representations for molecular drug. We conduct extensive experiments on several widely used benchmarks including Tox21, SIDER and ClinTox. The experimental results show that our approach can achieve state-of-the-art performance on these benchmarking datasets, demonstrating the feasibility and competitiveness of our proposed approach.
Xuan Lin, Zhe Quan, Zhi-Jie Wang 0009, Xiangxiang Zeng
Briefings Bioinform.5
2020 Computational methods for identifying the critical nodes in biological networks
abstract
A biological network is complex. A group of critical nodes determines the quality and state of such a network. Increasing studies have shown that diseases and biological networks are closely and mutually related and that certain diseases are often caused by errors occurring in certain nodes in biological networks. Thus, studying biological networks and identifying critical nodes can help determine the key targets in treating diseases. The problem is how to find the critical nodes in a network efficiently and with low cost. Existing experimental methods in identifying critical nodes generally require much time, manpower and money. Accordingly, many scientists are attempting to solve this problem by researching efficient and low-cost computing methods. To facilitate calculations, biological networks are often modeled as several common networks. In this review, we classify biological networks according to the network types used by several kinds of common computational methods and introduce the computational methods used by each type of network.
Xiangrong Liu, Zengyan Hong, Juan Liu 0003, Alfonso Rodríguez-Patón, Quan Zou 0001, Xiangxiang Zeng
Briefings Bioinform.7
2020 Predicting disease-associated circular RNAs using deep forests combined with positive-unlabeled learning methods
abstract
Identification of disease-associated circular RNAs (circRNAs) is of critical importance, especially with the dramatic increase in the amount of circRNAs. However, the availability of experimentally validated disease-associated circRNAs is limited, which restricts the development of effective computational methods. To our knowledge, systematic approaches for the prediction of disease-associated circRNAs are still lacking. In this study, we propose the use of deep forests combined with positive-unlabeled learning methods to predict potential disease-related circRNAs. In particular, a heterogeneous biological network involving 17 961 circRNAs, 469 miRNAs, and 248 diseases was constructed, and then 24 meta-path-based topological features were extracted. We applied 5-fold cross-validation on 15 disease data sets to benchmark the proposed approach and other competitive methods and used Recall@k and PRAUC@k to evaluate their performance. In general, our method performed better than the other methods. In addition, the performance of all methods improved with the accumulation of known positive labels. Our results provided a new framework to investigate the associations between circRNA and disease and might improve our understanding of its functions.
Xiangxiang Zeng, Quan Zou 0001
Briefings Bioinform.1
2020 Sequence clustering in bioinformatics: an empirical study
abstract
Sequence clustering is a basic bioinformatics task that is attracting renewed attention with the development of metagenomics and microbiomics. The latest sequencing techniques have decreased costs and as a result, massive amounts of DNA/RNA sequences are being produced. The challenge is to cluster the sequence data using stable, quick and accurate methods. For microbiome sequencing data, 16S ribosomal RNA operational taxonomic units are typically used. However, there is often a gap between algorithm developers and bioinformatics users. Different software tools can produce diverse results and users can find them difficult to analyze. Understanding the different clustering mechanisms is crucial to understanding the results that they produce. In this review, we selected several popular clustering tools, briefly explained the key computing principles, analyzed their characters and compared them using two independent benchmark datasets. Our aim is to assist bioinformatics users in employing suitable clustering tools effectively to analyze big sequencing data. Related data, codes and software tools were accessible at the link http://lab.malab.cn/∼lg/clustering/.
Quan Zou 0001, Xingpeng Jiang, Xiangrong Liu, Xiangxiang Zeng
Briefings Bioinform.5
2020 StackCPPred: a stacking and pairwise energy content-based prediction of cell-penetrating peptides and their uptake efficiency
abstract
MOTIVATION: Cell-penetrating peptides (CPPs) are a vehicle for transporting into living cells pharmacologically active molecules, such as short interfering RNAs, nanoparticles, plasmid DNAs and small peptides, thus offering great potential as future therapeutics. Existing experimental techniques for identifying CPPs are time-consuming and expensive. Thus, the prediction of CPPs from peptide sequences by using computational methods can be useful to annotate and guide the experimental process quickly. Many machine learning-based methods have recently emerged for identifying CPPs. Although considerable progress has been made, existing methods still have low feature representation capabilities, thereby limiting further performance improvements. RESULTS: We propose a method called StackCPPred, which proposes three feature methods on the basis of the pairwise energy content of the residue as follows: RECM-composition, PseRECM and RECM-DWT. These features are used to train stacking-based machine learning methods to effectively predict CPPs. On the basis of the CPP924 and CPPsite3 datasets with jackknife validation, StackDPPred achieved 94.5% and 78.3% accuracy, which was 2.9% and 5.8% higher than the state-of-the-art CPP predictors, respectively. StackCPPred can be a powerful tool for predicting CPPs and their uptake efficiency, facilitating hypothesis-driven experimental design and accelerating their applications in clinical therapy. AVAILABILITY AND IMPLEMENTATION: Source code and data can be downloaded from https://github.com/Excelsior511/StackCPPred. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiangzheng Fu, Xiangxiang Zeng, Quan Zou 0001
Bioinform.3
2020 Identifying enhancer-promoter interactions with neural network based on pre-trained DNA vectors and attention mechanism
abstract
MOTIVATION: Identification of enhancer-promoter interactions (EPIs) is of great significance to human development. However, experimental methods to identify EPIs cost too much in terms of time, manpower and money. Therefore, more and more research efforts are focused on developing computational methods to solve this problem. Unfortunately, most existing computational methods require a variety of genomic data, which are not always available, especially for a new cell line. Therefore, it limits the large-scale practical application of methods. As an alternative, computational methods using sequences only have great genome-scale application prospects. RESULTS: In this article, we propose a new deep learning method, namely EPIVAN, that enables predicting long-range EPIs using only genomic sequences. To explore the key sequential characteristics, we first use pre-trained DNA vectors to encode enhancers and promoters; afterwards, we use one-dimensional convolution and gated recurrent unit to extract local and global features; lastly, attention mechanism is used to boost the contribution of key features, further improving the performance of EPIVAN. Benchmarking comparisons on six cell lines show that EPIVAN performs better than state-of-the-art predictors. Moreover, we build a general model, which has transfer ability and can be used to predict EPIs in various cell lines. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at: https://github.com/hzy95/EPIVAN.
Zengyan Hong, Xiangxiang Zeng, Leyi Wei, Xiangrong Liu
Bioinform.2
2020 Network-based prediction of drug-target interactions using an arbitrary-order proximity embedded deep forest
abstract
MOTIVATION: Systematic identification of molecular targets among known drugs plays an essential role in drug repurposing and understanding of their unexpected side effects. Computational approaches for prediction of drug-target interactions (DTIs) are highly desired in comparison to traditional experimental assays. Furthermore, recent advances of multiomics technologies and systems biology approaches have generated large-scale heterogeneous, biological networks, which offer unexpected opportunities for network-based identification of new molecular targets among known drugs. RESULTS: In this study, we present a network-based computational framework, termed AOPEDF, an arbitrary-order proximity embedded deep forest approach, for prediction of DTIs. AOPEDF learns a low-dimensional vector representation of features that preserve arbitrary-order proximity from a highly integrated, heterogeneous biological network connecting drugs, targets (proteins) and diseases. In total, we construct a heterogeneous network by uniquely integrating 15 networks covering chemical, genomic, phenotypic and network profiles among drugs, proteins/targets and diseases. Then, we build a cascade deep forest classifier to infer new DTIs. Via systematic performance evaluation, AOPEDF achieves high accuracy in identifying molecular targets among known drugs on two external validation sets collected from DrugCentral [area under the receiver operating characteristic curve (AUROC) = 0.868] and ChEMBL (AUROC = 0.768) databases, outperforming several state-of-the-art methods. In a case study, we showcase that multiple molecular targets predicted by AOPEDF are associated with mechanism-of-action of substance abuse disorder for several marketed drugs (such as aripiprazole, risperidone and haloperidol). AVAILABILITY AND IMPLEMENTATION: Source code and data can be downloaded from https://github.com/ChengF-Lab/AOPEDF.
Xiangxiang Zeng, Siyi Zhu, Yuan Hou, Pengyue Zhang, Lang Li 0001, L. Frank Huang, Stephen J. Lewis, Ruth Nussinov, Feixiong Cheng
Bioinform.1
2020 Deep Collaborative Filtering for Prediction of Disease Genes
abstract
Accurate prioritization of potential disease genes is a fundamental challenge in biomedical research. Various algorithms have been developed to solve such problems. Inductive Matrix Completion (IMC) is one of the most reliable models for its well-established framework and its superior performance in predicting gene-disease associations. However, the IMC method does not hierarchically extract deep features, which might limit the quality of recovery. In this case, the architecture of deep learning, which obtains high-level representations and handles noises and outliers presented in large-scale biological datasets, is introduced into the side information of genes in our Deep Collaborative Filtering (DCF) model. Further, for lack of negative examples, we also exploit Positive-Unlabeled (PU) learning formulation to low-rank matrix completion. Our approach achieves substantially improved performance over other state-of-the-art methods on diseases from the Online Mendelian Inheritance in Man (OMIM) database. Our approach is 10 percent more efficient than standard IMC in detecting a true association, and significantly outperforms other alternatives in terms of the precision-recall metric at the top-k predictions. Moreover, we also validate the disease with no previously known gene associations and newly reported OMIM associations. The experimental results show that DCF is still satisfactory for ranking novel disease phenotypes as well as mining unexplored relationships. The source code and the data are available at https://github.com/xzenglab/DCF.
Xiangxiang Zeng, Yinglai Lin, Yuying He, Linyuan Lu, Xiaoping Min, Alfonso Rodríguez-Patón
IEEE ACM Trans. Comput. Biol. Bioinform.1
2020 Multiobjective Particle Swarm Optimization Based on Network Embedding for Complex Network Community Detection
abstract
Community detection in complex networks is significant to social network analysis. Most of the algorithms take advantage of single-objective optimization methods, which may not be effective for complex networks. Compared with single-objective algorithms, multiobjective evolutionary algorithms can avoid local optimization. However, multiobjective evolutionary algorithms often encounter problems of excessive search space and low efficiency. To solve these issues, this study introduces network embedding into the multiobjective particle swarm algorithm and maps nodes into a low-latitude space, thereby effectively reducing the search space while increasing search efficiency via a consensus propagation strategy. Experimental results demonstrate that a novel effective algorithm based on multiobjective particle swarm optimization (NE-PSO) performs effectively and has competitive performance in comparison with state-of-the-art approaches on synthetic and real-world networks, especially the large-scale ones.
Xiangrong Liu, Yanzi Du, Min Jiang 0005, Xiangxiang Zeng
IEEE Trans. Comput. Soc. Syst.4
2020 A Consensus Community-Based Particle Swarm Optimization for Dynamic Community Detection
abstract
The community detection in dynamic networks is essential for important applications such as social network analysis. Such detection requires simultaneous maximization of the clustering accuracy at the current time step while minimization of the clustering drift between two successive time steps. In most situations, such objectives are often in conflict with each other. This article proposes the concept of consensus community. Knowledge from the previous step is obtained by extracting the intrapopulation consensus communities from the optimal population of the previous step. Subsequently, the intrapopulation consensus communities of the previous step obtained is voted by the population of the current time step during the evolutionary process. A subset of the consensus communities, which receives a high support rate, will be recognized as the interpopulation consensus communities of the previous and current steps. Interpopulation consensus communities are the knowledge that can be transferred from the previous to the current step. The population of the current time step can evolve toward the direction similar to the population in the previous time step by retaining such interpopulation consensus community during the evolutionary process. Community structure is subjected to evaluation, update, and mutation events, which are directed by using interpopulation consensus community information during the evolutionary process. The experimental results over many artificial and real-world dynamic networks illustrate that the proposed method produces more accurate and robust results than those of the state-of-the-art approaches.
Xiangxiang Zeng, Gary G. Yen
IEEE Trans. Cybern.1
2020 A Network Reduction-Based Multiobjective Evolutionary Algorithm for Community Detection in Large-Scale Complex Networks
abstract
Evolutionary algorithms have been demonstrated to be very competitive in the community detection for complex networks. They, however, show poor scalability to large-scale networks due to the exponential increase of search space. In this paper, we suggest a network reduction-based multiobjective evolutionary algorithm for community detection in large-scale networks, where the size of the networks is recursively reduced as the evolution proceeds. In each reduction of the network, the local communities found by the elite individuals in the population are identified as nodes of the reduced network for further evolution, thereby considerably reducing the search space. A local community repairing strategy is also suggested to correct the misidentified nodes after each network reduction during the evolution. Experimental results on synthetic and real-world networks demonstrate the superiority of the proposed algorithm over several state-of-the-art community detection algorithms for large-scale networks, in terms of both computational efficiency and detection performance.
Xingyi Zhang 0001, Kefei Zhou, Hebin Pan, Lei Zhang 0060, Xiangxiang Zeng, Yaochu Jin
IEEE Trans. Cybern.5
2019 A Deep Neural Network for Antimicrobial Peptide Recognition
abstract
With the widespread use of antibiotics, many bacteria have developed resistance. Antimicrobial peptides have broad applications in medicine because of their high antibacterial activity. In this paper, a neural network model is introduced to recognize and detect antimicrobial peptides. Our model consists of an embedded, convolutional, bidirectional LSTM, and full connection layers. The embedded layer is used to code different amino acid residues into different vectors. The convolutional layer and bidirectional LSTM extract peptide amino acid residue sequence information. The full connection layer maps the sequence information linearly to the interval from 0 to 1, as the peptide for the probability of antimicrobial peptides. Training and testing on several different datasets reveal that our model performs better than other proposed models.
Jianyuan Lin, Xiangxiang Zeng, Yun Zuo 0001, Ying Ju 0002, Xiangrong Liu
BIBM2
2019 GraphCPI: Graph Neural Representation Learning for Compound-Protein Interaction
abstract
Accurately predicting compound-protein interactions (CPIs) is of great help to increase the efficiency and reduce costs in drug development. Most of existing machine learning models for CPI prediction often represent compounds and proteins in one-dimensional strings, or use the descriptor-based methods. These models might ignore the fact that molecules are essentially structured by the chemical bond of atoms. However, in real-world scenarios, the topological structure information usually provides an overview of how the atoms are connected, and the local chemical context reveals the functionality of the protein sequence in CPI. These two types of information are complementary to each other and they are both important for modeling compounds and proteins. Motivated by this, this paper suggests an end-to-end deep learning framework called GraphCPI, which captures the structural information of compounds and leverages the chemical context of protein sequences for solving the CPI prediction task. Our framework can integrate any popular graph nerual networks for learning compounds, and it combines with a convolutional neural network for embedding sequences. We conduct extensive experiments based on two benchmark CPI datasets. The experimental results demonstrate that our proposed framework is feasible and also competitive, comparing against classic and state-of-the-art methods.
Zhe Quan, Xuan Lin, Zhi-Jie Wang 0009, Xiangxiang Zeng
BIBM5
2019 Drug Target Interaction Prediction using Multi-task Learning and Co-attention
abstract
Various machine learning models have been proposed as cost-effective means to predict Drug-Target Interactions (DTI). Most existing researches treat DTI prediction either as a classification task (i.e. output negative or positive labels to indicate existence of interaction) or as a regression task (i.e. output numerical values as the strength of interaction). However, classifiers are more prone to higher bias and regression models tend to overfit the training data to generate large variance. In this paper, we explore to balance the bias and variance by a multi-task learning framework. We propose an architecture to both predict accurate values of strength of interaction and decide correct boundary between positive and negative interactions. Furthermore, the two tasks are performed on a shared feature representation, which is learnt using a co-attention mechanism. Comprehensive experiments demonstrate that the proposed method significantly outperforms state-of-the-art methods.
Yuyou Weng, Chen Lin 0001, Xiangxiang Zeng, Yun Liang 0003
BIBM3
2019 deepDR: a network-based deep learning approach to in silico drug repositioning
abstract
MOTIVATION: Traditional drug discovery and development are often time-consuming and high risk. Repurposing/repositioning of approved drugs offers a relatively low-cost and high-efficiency approach toward rapid development of efficacious treatments. The emergence of large-scale, heterogeneous biological networks has offered unprecedented opportunities for developing in silico drug repositioning approaches. However, capturing highly non-linear, heterogeneous network structures by most existing approaches for drug repositioning has been challenging. RESULTS: In this study, we developed a network-based deep-learning approach, termed deepDR, for in silico drug repurposing by integrating 10 networks: one drug-disease, one drug-side-effect, one drug-target and seven drug-drug networks. Specifically, deepDR learns high-level features of drugs from the heterogeneous networks by a multi-modal deep autoencoder. Then the learned low-dimensional representation of drugs together with clinically reported drug-disease pairs are encoded and decoded collectively via a variational autoencoder to infer candidates for approved drugs for which they were not originally approved. We found that deepDR revealed high performance [the area under receiver operating characteristic curve (AUROC) = 0.908], outperforming conventional network-based or machine learning-based approaches. Importantly, deepDR-predicted drug-disease associations were validated by the ClinicalTrials.gov database (AUROC = 0.826) and we showcased several novel deepDR-predicted approved drugs for Alzheimer's disease (e.g. risperidone and aripiprazole) and Parkinson's disease (e.g. methylphenidate and pergolide). AVAILABILITY AND IMPLEMENTATION: Source code and data can be downloaded from https://github.com/ChengF-Lab/deepDR. SUPPLEMENTARY INFORMATION: Supplementary data are available online at Bioinformatics.
Xiangxiang Zeng, Siyi Zhu, Xiangrong Liu, Yadi Zhou, Ruth Nussinov, Feixiong Cheng
Bioinform.1
2019 On solutions and representations of spiking neural P systems with rules on synapses
Francis George Cabarle, Ren Tristan A. de la Cruz, Dionne Peter P. Cailipan, Xiangrong Liu, Xiangxiang Zeng
Inf. Sci.6
2019 A component overlapping attribute clustering (COAC) algorithm for single-cell RNA sequencing data analysis and potential pathobiological implications
abstract
Recent advances in next-generation sequencing and computational technologies have enabled routine analysis of large-scale single-cell ribonucleic acid sequencing (scRNA-seq) data. However, scRNA-seq technologies have suffered from several technical challenges, including low mean expression levels in most genes and higher frequencies of missing data than bulk population sequencing technologies. Identifying functional gene sets and their regulatory networks that link specific cell types to human diseases and therapeutics from scRNA-seq profiles are daunting tasks. In this study, we developed a Component Overlapping Attribute Clustering (COAC) algorithm to perform the localized (cell subpopulation) gene co-expression network analysis from large-scale scRNA-seq profiles. Gene subnetworks that represent specific gene co-expression patterns are inferred from the components of a decomposed matrix of scRNA-seq profiles. We showed that single-cell gene subnetworks identified by COAC from multiple time points within cell phases can be used for cell type identification with high accuracy (83%). In addition, COAC-inferred subnetworks from melanoma patients' scRNA-seq profiles are highly correlated with survival rate from The Cancer Genome Atlas (TCGA). Moreover, the localized gene subnetworks identified by COAC from individual patients' scRNA-seq data can be used as pharmacogenomics biomarkers to predict drug responses (The area under the receiver operating characteristic curves ranges from 0.728 to 0.783) in cancer cell lines from the Genomics of Drug Sensitivity in Cancer (GDSC) database. In summary, COAC offers a powerful tool to identify potential network-based diagnostic and pharmacogenomics biomarkers from large-scale scRNA-seq profiles. COAC is freely available at https://github.com/ChengF-Lab/COAC.
He Peng, Xiangxiang Zeng, Yadi Zhou, Ruth Nussinov, Feixiong Cheng
PLoS Comput. Biol.2
2019 Details in the evaluation of circular RNA detection tools: Reply to Chen and Chuang
abstract
Chia-Ying Chen and Trees-Juen Chuang (referred as CYC & TJC below) recently submitted their comment [1] on our previous paper [2].In their paper, they scrutinized the CircBase [3] candidates that we used and pointed out several weak points of our paper.In summary, they suggested that the positive dataset we derived from CircBase required further evaluation.They also indicated that using all of these candidates as our dataset was not appropriate.They further suggested that three main confounding factors may affect our assessment of circRNA detection tools and that their performances should be re-evaluated.Before we begin to discuss their comment, we will briefly introduce the positive dataset we used.First, as stated in our previous paper, the 14,689 candidates detected in HeLa cells were downloaded from CircBase and reported by the study of Salzman et al. [4].These candidates were not identified with the use of find_circ [5] tool.As described in the study of Salzman et al. [4], all UCSC annotated exons in scrambled order were used to construct a custom database and identify circRNA candidates.Second, in our positive dataset, constant coverage of 10× for the intervening sequence and a minimum of two read pairs (paired-end simulated reads) to cross the back-spliced junction sites were generated for each candidate.Now, we will discuss the three confounding factors they listed in their paper.First, they suggested to remove 1046 candidates with unannotated exon boundaries from the positive dataset, especially candidates without canonical splice signals, such as GT-AG, GC-AG, or AT-AC, for the junctions.As mentioned above, CircBase-deposited circRNA candidates that we used were identified by Salzman et al. [4]; the candidates identified by their method should all match the exon boundaries.The discrepancies may be caused by inconsistent gene annotation files used.Salzman et al. [4] used UCSC known genes [6], whereas CYC & TJC used NCBI RefSeq-identified mRNA annotation files.We manually checked several candidates marked with "junctions with unannotated exon boundaries" in CYC & TJC's Supplemental Dataset S1.The junction sites of these candidates were annotated as exon boundaries in UCSC known genes annotation file (http://hgdownload.soe.ucsc.edu/goldenPath/hg19/database/knownGene.txt.gz).Thus, detection of circRNAs with annotated exon boundaries relies on the gene annotation files used, and novel candidates may be missed because of the incompleteness of the current database [7].For example, Szabo et al. [7] reinforced an annotation-based algorithm with a de novo module and discovered a validated circRNA from the not-fully-annotated RMST gene and several U12 cir-cRNAs produced from unannotated boundaries.Such case was also demonstrated by Xiao-Ou Zhang et al. [8].They detected thousands of novel exons (non-RefSeq, non-Ensembl, or non-UCSC known genes) in circRNAs by using an updated CIRCexplorere2 tool, and several of them were confirmed by Northern blot analysis and Sanger sequencing after RT-PCR [8].Other examples were shown by Salzman et al. [4], they found several noncoding RNA genes expressed
Xiangxiang Zeng, Maozu Guo 0001, Quan Zou 0001
PLoS Comput. Biol.1
2019 Meta-Path Methods for Prioritizing Candidate Disease miRNAs
abstract
MicroRNAs (miRNAs) play critical roles in regulating gene expression at post-transcriptional levels. Numerous experimental studies indicate that alterations and dysregulations in miRNAs are associated with important complex diseases, especially cancers. Predicting potential miRNA-disease association is beneficial not only to explore the pathogenesis of diseases, but also to understand biological processes. In this work, we propose two methods that can effectively predict potential miRNA-disease associations using our reconstructed miRNA and disease similarity networks, which are based on the latest experimental data. We reconstruct a miRNA functional similarity network using the following biological information: the miRNA family information, miRNA cluster information, experimentally valid miRNA-target association and disease-miRNA information. We also reconstruct a disease similarity network using disease functional information and disease semantic information. We present Katz with specific weights and Katz with machine learning, on the comprehensive heterogeneous network. These methods, which achieve corresponding AUC values of 0.897 and 0.919, exhibit performance superior to the existing methods. Comprehensive data networks and reasonable considerations guarantee the high performance of our methods. Contrary to several methods, which cannot work in such situations, the proposed methods also predict associations for diseases without any known related miRNAs. A web service for the download and prediction of relationships between diseases and miRNAs is available at http://lab.malab.cn/soft/MDPredict/.
Xuan Zhang 0010, Quan Zou 0001, Alfonso Rodríguez-Patón, Xiangxiang Zeng
IEEE ACM Trans. Comput. Biol. Bioinform.4
2019 An Evolutionary Algorithm Based on Minkowski Distance for Many-Objective Optimization
abstract
The existing multiobjective evolutionary algorithms (EAs) based on nondominated sorting may encounter serious difficulties in tackling many-objective optimization problems (MaOPs), because the number of nondominated solutions increases exponentially with the number of objectives, leading to a severe loss of selection pressure. To address this problem, some existing many-objective EAs (MaOEAs) adopt Euclidean or Manhattan distance to estimate the convergence of each solution during the environmental selection process. Nevertheless, either Euclidean or Manhattan distance is a special case of Minkowski distance with the order P=2 or P=1 , respectively. Thus, it is natural to adopt Minkowski distance for convergence estimation, in order to cover various types of Pareto fronts (PFs) with different concavity-convexity degrees. In this paper, a Minkowski distance-based EA is proposed to solve MaOPs. In the proposed algorithm, first, the concavity-convexity degree of the approximate PF, denoted by the value of P , is dynamically estimated. Subsequently, the Minkowski distance of order P is used to estimate the convergence of each solution. Finally, the optimal solutions are selected by a comprehensive method, based on both convergence and diversity. In the experiments, the proposed algorithm is compared with five state-of-the-art MaOEAs on some widely used benchmark problems. Moreover, the modified versions for two compared algorithms, integrated with the proposed P -estimation method and the Minkowski distance, are also designed and analyzed. Empirical results show that the proposed algorithm is very competitive against other MaOEAs for solving MaOPs, and two modified compared algorithms are generally more effective than their predecessors.
Hang Xu 0003, Wenhua Zeng, Xiangxiang Zeng, Gary G. Yen
IEEE Trans. Cybern.3
2019 MOEA/HD: A Multiobjective Evolutionary Algorithm Based on Hierarchical Decomposition
abstract
Recently, numerous multiobjective evolutionary algorithms (MOEAs) have been proposed to solve the multiobjective optimization problems (MOPs). One of the most widely studied MOEAs is that based on decomposition (MOEA/D), which decomposes an MOP into a series of scalar optimization subproblems, via a set of uniformly distributed weight vectors. MOEA/D shows excellent performance on most mild MOPs, but may face difficulties on ill MOPs, with complex Pareto fronts, which are pointed, long tailed, disconnected, or degenerate. That is because the weight vectors used in decomposition are all preset and invariant. To overcome it, a new MOEA based on hierarchical decomposition (MOEA/HD) is proposed in this paper. In MOEA/HD, subproblems are layered into different hierarchies, and the search directions of lower-hierarchy subproblems are adaptively adjusted, according to the higher-hierarchy search results. In the experiments, MOEA/HD is compared with four state-of-the-art MOEAs, in terms of two widely used performance metrics. According to the empirical results, MOEA/HD shows promising performance on all the test problems.
Hang Xu 0003, Wenhua Zeng, Xiangxiang Zeng
IEEE Trans. Cybern.4
2018 LncRNA-disease association prediction based on neighborhood information aggregation in neural network
Hongjie Chen 0003, Xun Wang 0010, Xuan Zhang 0010, Xiangxiang Zeng, Tao Song 0001, Alfonso Rodríguez-Patón
BIBM4
2018 Drug Target Interaction Prediction with Non-random Missing Labels
Sheng Ni, Chen Lin 0001, Xiangxiang Zeng, Yun Liang 0003
BIBM3
2018 Prediction of potential disease-associated microRNAs using structural perturbation method
abstract
Motivation: The identification of disease-related microRNAs (miRNAs) is an essential but challenging task in bioinformatics research. Similarity-based link prediction methods are often used to predict potential associations between miRNAs and diseases. In these methods, all unobserved associations are ranked by their similarity scores. Higher score indicates higher probability of existence. However, most previous studies mainly focus on designing advanced methods to improve the prediction accuracy while neglect to investigate the link predictability of the networks that present the miRNAs and diseases associations. In this work, we construct a bilayer network by integrating the miRNA-disease network, the miRNA similarity network and the disease similarity network. We use structural consistency as an indicator to estimate the link predictability of the related networks. On the basis of the indicator, a derivative algorithm, called structural perturbation method (SPM), is applied to predict potential associations between miRNAs and diseases. Results: The link predictability of bilayer network is higher than that of miRNA-disease network, indicating that the prediction of potential miRNAs-diseases associations on bilayer network can achieve higher accuracy than based merely on the miRNA-disease network. A comparison between the SPM and other algorithms reveals the reliable performance of SPM which performed well in a 5-fold cross-validation. We test fifteen networks. The AUC values of SPM are higher than some well-known methods, indicating that SPM could serve as a useful computational method for improving the identification accuracy of miRNA‒disease associations. Moreover, in a case study on breast neoplasm, 80% of the top-20 predicted miRNAs have been manually confirmed by previous experimental studies. Availability and implementation: https://github.com/lecea/SPM-code.git. Supplementary information: Supplementary data are available at Bioinformatics online.
Xiangxiang Zeng, Linyuan Lu, Quan Zou 0001
Bioinform.1
2017 Iteratively collective prediction of disease-gene associations through the incomplete network
abstract
The prediction of links between genes and disease is still one of the biggest challenges in the field of human health. Almost all state-of-the-art studies on the prediction of gene-disease links focuson a single pair of links, ignoring the associations and interactions among different types of links. Moreover, the biological information networks are usually incomplete. In this paper, we study the similarity measure to be used on two different types of nodes, based on the metapaths between them (Wsrm). Then an iterative self-updating approach for link prediction using heterogeneous information network is proposed to fit the incompletion of the network (ISL), which is a semi-supervised learning formula. Using the biological integrated network constructed from OMIM and HumanNet dataset (30,896 nodes and 1,200,166 edges) we applied our framework. The area under the receiver operating characteristic is 0.941, indicating that our approach significantly outperforms the state-of-the-art gene-disease link prediction approaches. Moreover, the sensitivity analysis signifies that our approach is robust. Consequently, our proposed framework demonstrates an efficient and accurate approach for link prediction between genes and diseases. In addition, during iteration, the accuracy of the result gradually increases. The example dataset and the implementation of our approach is avaliable at https://github.com/xymeng16/ISL.
Xiangyi Meng, Quan Zou 0001, Alfonso Rodríguez-Patón, Xiangxiang Zeng
BIBM4
2017 A comprehensive overview and evaluation of circular RNA detection tools
abstract
Circular RNA (circRNA) is mainly generated by the splice donor of a downstream exon joining to an upstream splice acceptor, a phenomenon known as backsplicing. It has been reported that circRNA can function as microRNA (miRNA) sponges, transcriptional regulators, or potential biomarkers. The availability of massive non-polyadenylated transcriptomes data has facilitated the genome-wide identification of thousands of circRNAs. Several circRNA detection tools or pipelines have recently been developed, and it is essential to provide useful guidelines on these pipelines for users, including a comprehensive and unbiased comparison. Here, we provide an improved and easy-to-use circRNA read simulator that can produce mimicking backsplicing reads supporting circRNAs deposited in CircBase. Moreover, we compared the performance of 11 circRNA detection tools on both simulated and real datasets. We assessed their performance regarding metrics such as precision, sensitivity, F1 score, and Area under Curve. It is concluded that no single method dominated on all of these metrics. Among all of the state-of-the-art tools, CIRI, CIRCexplorer, and KNIFE, which achieved better balanced performance between their precision and sensitivity, compared favorably to the other methods.
Xiangxiang Zeng, Maozu Guo 0001, Quan Zou 0001
PLoS Comput. Biol.1
2017 Inferring MicroRNA-Disease Associations by Random Walk on a Heterogeneous Network with Multiple Data Sources
abstract
Since the discovery of the regulatory function of microRNA (miRNA), increased attention has focused on identifying the relationship between miRNA and disease. It has been suggested that computational method are an efficient way to identify potential disease-related miRNAs for further confirmation using biological experiments. In this paper, we first highlighted three limitations commonly associated with previous computational methods. To resolve these limitations, we established disease similarity subnetwork and miRNA similarity subnetwork by integrating multiple data sources, where the disease similarity is composed of disease semantic similarity and disease functional similarity, and the miRNA similarity is calculated using the miRNA-target gene and miRNA-lncRNA (long non-coding RNA) associations. Then, a heterogeneous network was constructed by connecting the disease similarity subnetwork and the miRNA similarity subnetwork using the known miRNA-disease associations. We extended random walk with restart to predict miRNA-disease associations in the heterogeneous network. The leave-one-out cross-validation achieved an average area under the curve (AUC) of 0:8049 across 341 diseases and 476 miRNAs. For five-fold cross-validation, our method achieved an AUC from 0:7970 to 0:9249 for 15 human diseases. Case studies further demonstrated the feasibility of our method to discover potential miRNA-disease associations. An online service for prediction is freely available at http://ifmda.aliapp.com.
Yuansheng Liu, Xiangxiang Zeng, Zengyou He, Quan Zou 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2017 Prediction and Validation of Disease Genes Using HeteSim Scores
abstract
Deciphering the gene disease association is an important goal in biomedical research. In this paper, we use a novel relevance measure, called HeteSim, to prioritize candidate disease genes. Two methods based on heterogeneous networks constructed using protein-protein interaction, gene-phenotype associations, and phenotype-phenotype similarity, are presented. In HeteSim_MultiPath (HSMP), HeteSim scores of different paths are combined with a constant that dampens the contributions of longer paths. In HeteSim_SVM (HSSVM), HeteSim scores are combined with a machine learning method. The 3-fold experiments show that our non-machine learning method HSMP performs better than the existing non-machine learning methods, our machine learning method HSSVM obtains similar accuracy with the best existing machine learning method CATAPULT. From the analysis of the top 10 predicted genes for different diseases, we found that HSSVM avoid the disadvantage of the existing machine learning based methods, which always predict similar genes for different diseases. The data sets and Matlab code for the two methods are freely available for download at http://lab.malab.cn/data/HeteSim/index.jsp.
Xiangxiang Zeng, Yuanlu Liao, Yuansheng Liu, Quan Zou 0001
IEEE ACM Trans. Comput. Biol. Bioinform.1
2016 Latent factor model with heterogeneous similarity regularization for predicting gene-disease associations
abstract
The correct prediction of human genes related to diseases has been a challenge in biological research. Considering extensive gene-disease data verified by biological experiments, we can apply computational methods to perform correct predictions with reduced time and expenses. On the basis of a previously designed latent factorization model (LFM), which performs well in recommender systems, we propose a latent factor model with heterogeneous similarity regularization (LFMHSR) to predict disease-related genes. Various types of data, including those of humans and other related species, are used in this method. First, model I with an average heterogeneous regularization is proposed on the basis of a typical LFM. Second, model II with personal heterogeneous regularization is developed to improve the deficiency of the previous model. Data on other nonhuman species and vector space similarity or Pearson correlation coefficient metrics are also utilized in our method. Results reveal that the performance of LFMHSR is 7% more efficient than that of other existing approaches. Therefore, our proposed approach can be employed to predict novel diseases or genes with no known associations.
Xiangxiang Zeng, Ningxiang Ding, Quan Zou 0001
BIBM1
2016 HPTree: Reconstructing phylogenetic trees for ultra-large unaligned DNA sequences via NJ model and Hadoop
abstract
Constructing phylogenetic tree for ultra-large sequences (eg. Files more than 1GB) is quite difficult, especially for the unaligned DNA sequences. It is meaningless and impracticable to do multiple sequence alignment for large diverse DNA sequences. We try to do clustering firstly for the mounts of DNA sequences, and divide them into several clusters. Then each cluster is aligned and phylogenetic analysed in parallel. Hadoop, which is the most popular parallel platform in cloud computing, is employed for this process. Our software tool HPTree can handle the >1GB DNA sequence file or more than 1,000,000 DNA sequences in few hours. Users could try HPTree in the cloud computing platform (eg. Amazon) or their own clusters for the big data phylogenetic tree reconstruction. No super machine or large memory is required. HPTree could benefit the users who focus on population evolution or long common genes (eg. 16s rRNA) evolution. The software tool along with its codes and datasets are accessible at http://lab.malab.cn/soft/HPtree/.
Quan Zou 0001, Shixiang Wan, Xiangxiang Zeng
BIBM3
2016 Integrative approaches for predicting microRNA function and prioritizing disease-related microRNA using biological interaction networks
abstract
MicroRNAs (miRNA) play critical roles in regulating gene expressions at the posttranscriptional levels. The prediction of disease-related miRNA is vital to the further investigation of miRNA's involvement in the pathogenesis of disease. In previous years, biological experimentation is the main method used to identify whether miRNA was associated with a given disease. With increasing biological information and the appearance of new miRNAs every year, experimental identification of disease-related miRNAs poses considerable difficulties (e.g. time-consumption and high cost). Because of the limitations of experimental methods in determining the relationship between miRNAs and diseases, computational methods have been proposed. A key to predict potential disease-related miRNA based on networks is the calculation of similarity among diseases and miRNA over the networks. Different strategies lead to different results. In this review, we summarize the existing computational approaches and present the confronted difficulties that help understand the research status. We also discuss the principles, efficiency and differences among these methods. The comprehensive comparison and discussion elucidated in this work provide constructive insights into the matter.
Xiangxiang Zeng, Xuan Zhang 0010, Quan Zou 0001
Briefings Bioinform.1
2016 Investment behavior prediction in heterogeneous information network
Xiangxiang Zeng, Stephen C. H. Leung, Ziyu Lin, Xiangrong Liu
Neurocomputing1
2016 Computing with viruses
Xu Chen 0020, Mario J. Pérez-Jiménez, Luis Valencia-Cabrera, Beizhan Wang, Xiangxiang Zeng
Theor. Comput. Sci.5
2015 Short-term optimal hydrothermal scheduling problem considering power flow constraint
abstract
Short-term optimal hydrothermal scheduling problem is one of the most popular research issues in power systems optimization. A novel mathematical model of the short-term optimal hydrothermal scheduling is proposed in this paper. This model aims at minimizing the total fuel cost of the thermal generating units while satisfying the various constraints such as power balance, water balance, transmission network and other system's constraints. A modified differential evolution algorithm is also introduced to solve the short-term optimal hydrothermal scheduling problem. In the proposed approach, an operation of migration and a self-adaptive mechanism are presented to improve the searching efficiency. Moreover, four constraint handling rules are proposed to handle the complex constraints of short-term optimal hydrothermal scheduling problem. An IEEE nine buses test system is applied to verify the proposed mathematic model and algorithm. The numerical results show the feasibility and efficiency of the proposed approach to the short-term optimal hydrothermal scheduling problem.
Jingrui Zhang, Shuang Lin, Xiangxiang Zeng, Qinghui Tang
CEC3
2015 A Stable Matching-Based Selection and Memory Enhanced MOEA/D for Evolutionary Dynamic Multiobjective Optimization
abstract
In the real world, dynamic changes may occur during multi-objective optimization. In those situations, it is vital to track the time-varying Pareto optimal set over time. This paper is to integrate a memory-enhanced multi-objective evolutionary algorithm based on decomposition (denoted by dMOEA/D-M) with a simple and effective stable matching (STM) model (denoted by dMOEA/D-STM). MOEA/D is an effective algorithm for optimizing static multi-objective problems. For adapting to the dynamic changes, firstly, an improved environment detector is presented. Then, memory and matching skills is designed to address the difficulties of re-initialization. The STM model, which originates from economics, guides the re-initialization in dMOEA/D-STM. Empirical experiments prove the effectiveness of the memory strategy and STM model.
Xiaofeng Chen 0001, Xiangxiang Zeng
ICTAI3
2015 A decision-making framework for precision marketing
Zhen You, Yain-Whar Si, Xiangxiang Zeng, Stephen C. H. Leung
Expert Syst. Appl.4
2015 Identification of cytokine via an improved genetic algorithm
Xiangxiang Zeng, Sisi Yuan, Xianxian Huang, Quan Zou 0001
Frontiers Comput. Sci.1
2015 The power of time-free tissue P systems: Attacking NP-complete problems
Xiangrong Liu, Juan Suo, Stephen C. H. Leung, Juan Liu 0003, Xiangxiang Zeng
Neurocomputing5
2015 Asynchronous spiking neural P systems with rules on synapses
Tao Song 0001, Quan Zou 0001, Xiangrong Liu, Xiangxiang Zeng
Neurocomputing4
2015 Asynchronous Spiking Neural P Systems with Anti-Spikes
Tao Song 0001, Xiangrong Liu, Xiangxiang Zeng
Neural Process. Lett.3
2014 nDNA-prot: identification of DNA-binding proteins based on unbalanced classification
abstract
BACKGROUND: DNA-binding proteins are vital for the study of cellular processes. In recent genome engineering studies, the identification of proteins with certain functions has become increasingly important and needs to be performed rapidly and efficiently. In previous years, several approaches have been developed to improve the identification of DNA-binding proteins. However, the currently available resources are insufficient to accurately identify these proteins. Because of this, the previous research has been limited by the relatively unbalanced accuracy rate and the low identification success of the current methods. RESULTS: In this paper, we explored the practicality of modelling DNA binding identification and simultaneously employed an ensemble classifier, and a new predictor (nDNA-Prot) was designed. The presented framework is comprised of two stages: a 188-dimension feature extraction method to obtain the protein structure and an ensemble classifier designated as imDC. Experiments using different datasets showed that our method is more successful than the traditional methods in identifying DNA-binding proteins. The identification was conducted using a feature that selected the minimum Redundancy and Maximum Relevance (mRMR). An accuracy rate of 95.80% and an Area Under the Curve (AUC) value of 0.986 were obtained in a cross validation. A test dataset was tested in our method and resulted in an 86% accuracy, versus a 76% using iDNA-Prot and a 68% accuracy using DNA-Prot. CONCLUSIONS: Our method can help to accurately identify DNA-binding proteins, and the web server is accessible at http://datamining.xmu.edu.cn/~songli/nDNA. In addition, we also predicted possible DNA-binding protein sequences in all of the sequences from the UniProtKB/Swiss-Prot database.
Xiangxiang Zeng, Quan Zou 0001
BMC Bioinform.3
2014 Small universal simple spiking neural P systems with weights
Xiangxiang Zeng, Linqiang Pan, Mario J. Pérez-Jiménez
Sci. China Inf. Sci.1
2014 Weighted Spiking Neural P Systems with Rules on Synapses
abstract
Spiking neural P systems (SN P systems, for short) with rules on synapses are a new variant of SN P systems, where the spiking and forgetting rules are placed on synapses instead of in neurons. Recent studies illustrated that this variant of SN P sys
Xingyi Zhang 0001, Xiangxiang Zeng, Linqiang Pan
Fundam. Informaticae2
2014 On languages generated by spiking neural P systems with weights
Xiangxiang Zeng, Lei Xu 0002, Xiangrong Liu, Linqiang Pan
Inf. Sci.1
2014 Spiking Neural P Systems with Thresholds
abstract
Spiking neural P systems with weights are a new class of distributed and parallel computing models inspired by spiking neurons. In such models, a neuron fires when its potential equals a given value (called a threshold). In this work, spiking neural P systems with thresholds (SNPT systems) are introduced, where a neuron fires not only when its potential equals the threshold but also when its potential is higher than the threshold. Two types of SNPT systems are investigated. In the first one, we consider that the firing of a neuron consumes part of the potential (the amount of potential consumed depends on the rule to be applied). In the second one, once a neuron fires, its potential vanishes (i.e., it is reset to zero). The computation power of the two types of SNPT systems is investigated. We prove that the systems of the former type can compute all Turing computable sets of numbers and the systems of the latter type characterize the family of semilinear sets of numbers. The results show that the firing mechanism of neurons has a crucial influence on the computation power of the SNPT systems, which also answers an open problem formulated in Wang, Hoogeboom, Pan, Păun, and Pérez-Jiménez ( 2010 ).
Xiangxiang Zeng, Xingyi Zhang 0001, Tao Song 0001, Linqiang Pan
Neural Comput.1
2014 On Some Classes of Sequential Spiking Neural P Systems
abstract
Spiking neural P systems (SN P systems) are a class of distributed parallel computing devices inspired by the way neurons communicate by means of spikes; neurons work in parallel in the sense that each neuron that can fire should fire, but the work in each neuron is sequential in the sense that at most one rule can be applied at each computation step. In this work, with biological inspiration, we consider SN P systems with the restriction that at each step, one of the neurons (i.e., sequential mode) or all neurons (i.e., pseudo-sequential mode) with the maximum (or minimum) number of spikes among the neurons that are active (can spike) will fire. If an active neuron has more than one enabled rule, it nondeterministically chooses one of the enabled rules to be applied, and the chosen rule is applied in an exhaustive manner (a kind of local parallelism): the rule is used as many times as possible. This strategy makes the system sequential or pseudo-sequential from the global view of the whole network and locally parallel at the level of neurons. We obtain four types of SN P systems: maximum/minimum spike number induced sequential/pseudo-sequential SN P systems with exhaustive use of rules. We prove that SN P systems of these four types are all Turing universal as number-generating computation devices. These results illustrate that the restriction of sequentiality may have little effect on the computation power of SN P systems.
Xingyi Zhang 0001, Xiangxiang Zeng, Bin Luo 0001, Linqiang Pan
Neural Comput.2
2012 A uniform solution to the independent set problem through tissue P systems with cell separation
Xingyi Zhang 0001, Xiangxiang Zeng, Bin Luo 0001, Zheng Zhang 0001
Frontiers Comput. Sci.2
2012 Spiking Neural P Systems with Weighted Synapses
Linqiang Pan, Xiangxiang Zeng, Xingyi Zhang 0001
Neural Process. Lett.2
2011 Time-Free Spiking Neural P Systems
abstract
Different biological processes take different times to be completed, which can also be influenced by many environmental factors. In this work, a realistic definition of nonsynchronized spiking neural P systems (SN P systems, for short) is considered: during the work of an SN P system, the execution times of spiking rules cannot be known exactly (i.e., they are arbitrary). In order to establish robust systems against the environmental factors, a special class of SN P systems, called time-free SN P systems, is introduced, which always produce the same computation result independent of the execution times of the rules. The universality of time-free SN P systems is investigated. It is proved that these P systems with extended rules (several spikes can be produced by a rule) are equivalent to register machines. However, if the number of spikes present in the system is bounded, then the power of time-free SN P systems falls, and in this case, a characterization of semilinear sets of natural numbers is obtained.
Linqiang Pan, Xiangxiang Zeng, Xingyi Zhang 0001
Neural Comput.2
2010 Deterministic solutions to QSAT and Q3SAT by spiking neural P systems with pre-computed resources
Tseren-Onolt Ishdorj, Alberto Leporati, Linqiang Pan, Xiangxiang Zeng, Xingyi Zhang 0001
Theor. Comput. Sci.4
2009 Homogeneous Spiking Neural P Systems
abstract
Spiking neural P systems are a class of distributed parallel computing models inspired from the way the neurons communicate with each other by means of electrical impulses (called "spikes"). In this paper, we consider a restricted variant of spiking neural P systems, called homogeneous spiking neural P systems, where each neuron has the same set of rules. The universality of homogeneous spiking neural P systems is investigated. One of universality results is that it is sufficient for homogeneous spiking neural P system to have only one neuron that behaves nondeterministically in order to achieve Turing completeness.
Xiangxiang Zeng, Xingyi Zhang 0001, Linqiang Pan
Fundam. Informaticae1
2009 On languages generated by asynchronous spiking neural P systems
Xingyi Zhang 0001, Xiangxiang Zeng, Linqiang Pan
Theor. Comput. Sci.2
2008 Smaller Universal Spiking Neural P Systems
Xingyi Zhang 0001, Xiangxiang Zeng, Linqiang Pan
Fundam. Informaticae2
2008 On string languages generated by spiking neural P systems with exhaustive use of rules
Xingyi Zhang 0001, Xiangxiang Zeng, Linqiang Pan
Nat. Comput.2