Bosheng Song

dblp:16/10049 · DBLP profile ↗
← Back
59ranked-venue papers
21as first author
42since 2021 · last 2026
0000-0002-1479-5399ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Theory of computation · 21 · 10 first-author · 7 since 2021Artificial intelligence and machine learning · 15 · 3 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 4 first-author · 15 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 4 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 A systematic review of molecular representation learning foundation models
abstract
Molecular representation learning (MRL) is afoundation in leveraging computational methods for drug discovery, enabling the transformation of molecular structure and properties into numerical vectors. These vectors serve as input for machine learning models and facilitate the prediction and analysis of molecular attributes, functions, and reactions. The advent of foundation models has introduced both new opportunities and challenges to MRL. These models have improved generalizability and migration in scarce data. Through pretraining and fine-tuning, foundation models can be adapted to various domains. Their robust encoding and generative abilities also allow the transformation of molecular data into more expressive forms. This paper provides a detailed review of current mainstream molecular descriptors and datasets, focusing primarily on the representation of small molecules while excluding larger molecules such as proteins and peptides. It classifies foundation models into two primary categories based on the form of input: unimodal-based and multimodal-based models. For each category, representative models are identified and their advantages and disadvantages evaluated. Moreover, we systematically summarize four core pretraining strategies for MRL foundation models, analyzing their task designs, applicable scenarios, and impacts on downstream performance. In addition, the application of molecular representation foundation models in drug discovery and development is discussed, together with the current status of model interpretability. The paper concludes with insights into the future directions of MRL foundation models.
Bosheng Song, Yuansheng Liu, Sisi Yuan, Xia Zhen
Briefings Bioinform.1
2026 Empowering chemical structures with biological insights for scalable phenotypic virtual screening
abstract
MOTIVATION: The scalable identification of bioactive compounds is essential for contemporary drug discovery. This process faces a key trade-off: structural screening offers scalability but lacks biological context, whereas high-content phenotypic profiling provides deep biological insights but is resource-intensive. The primary challenge is to extract robust biological signals from noisy data and encode them into representations that do not require biological data at inference. RESULTS: This study presents DECODE (DEcomposing Cellular Observations of Drug Effects), a framework that bridges this gap by empowering chemical representations with intrinsic biological semantics to enable structure-based in silico biological profiling. DECODE leverages limited paired transcriptomic and morphological data as supervisory signals during training, enabling the extraction of a measurement-invariant biological fingerprint from chemical structures and explicit filtering of modality-specific variation. Across held-out retrieval, scaffold-split and UMAP-clustering virtual-screening benchmarks, DECODE improves functional retrieval and early active-compound prioritization over baselines. AVAILABILITY AND IMPLEMENTATION: The codes and datasets of DECODE are available at https://github.com/lian-xiao/DECODE.
Xiaoqing Lian, Pengsen Ma, Tengfeng Ma, Zhong-Hao Ren, Xibao Cai, Zhixiang Cheng, Bosheng Song, Sisi Yuan, Chen Lin 0001
Bioinform.7
2026 Neuropeptide-regulated P system for semi-supervised segmentation of multiple brain metastases on MRIs
Jie Xue 0001, Kexu Zhao, Xiyu Liu 0001, Bosheng Song, Shulei Chang, Guanzhong Gong, Dengwang Li
Eng. Appl. Artif. Intell.4
2026 Learnable dendrite neural P systems and applications in survival prediction of glioblastoma patients
Xiu Yin, Xiyu Liu 0001, Shulei Chang, Bosheng Song, Guanzhong Gong, Jiaxing Yin, Dengwang Li, Jie Xue 0001
Neural Networks4
2025 S²DN: Learning to Denoise Unconvincing Knowledge for Inductive Knowledge Graph Completion
abstract
Inductive Knowledge Graph Completion (KGC) aims to infer missing facts between newly emerged entities within knowledge graphs (KGs), posing a significant challenge. While recent studies have shown promising results in inferring such entities through knowledge subgraph reasoning, they suffer from (i) the semantic inconsistencies of similar relations, and (ii) noisy interactions inherent in KGs due to the presence of unconvincing knowledge for emerging entities. To address these challenges, we propose a Semantic Structure-aware Denoising Network (S2DN) for inductive KGC. Our goal is to learn adaptable general semantics and reliable structures to distill consistent semantic knowledge while preserving reliable interactions within KGs. Specifically, we introduce a semantic smoothing module over the enclosing subgraphs to retain the universal semantic knowledge of relations. We incorporate a structure refining module to filter out unreliable interactions and offer additional knowledge, retaining robust structure surrounding target links. Extensive experiments conducted on three benchmark KGs demonstrate that S2DN surpasses the performance of state-of-the-art models. These results demonstrate the effectiveness of S2DN in preserving semantic consistency and enhancing the robustness of filtering out unreliable interactions in contaminated KGs.
Tengfei Ma 0002, Yujie Chen 0002, Xuan Lin, Bosheng Song, Xiangxiang Zeng
AAAI5
2025 Multi-Objective Molecular Design Through Learning Latent Pareto Set
abstract
Molecular design inherently involves the optimization of multiple conflicting objectives, such as enhancing bio-activity and ensuring synthesizability. Evaluating these objectives often requires resource-intensive computations or physical experiments. Current molecular design methodologies typically approximate the Pareto set using a limited number of molecules. In this paper, we present an innovative approach, called Multi-Objective Molecular Design through Learning Latent Pareto Set (MLPS). MLPS initially utilizes an encoder-decoder model to seamlessly transform the discrete chemical space into a continuous latent space. We then employ local Bayesian optimization models to efficiently search for local optimal solutions (i.e., molecules) within predefined trust regions. Using surrogate objective values derived from these local models, we train a global Pareto set learning model to understand the mapping between direction vectors (called “preferences”) in the objective space and the entire Pareto set in the continuous latent space. Both the global Pareto set learning model and local Bayesian optimization models collaborate to discover high-quality solutions and adapt the trust regions dynamically. Our work is an effective endeavor towards learning the Pareto set for multi-objective molecular design, providing decision-makers with the capability to fine-tune their preferences and thoroughly explore the Pareto set. Experimental results demonstrate that MLPS achieves state-of-the-art performance across various multi-objective scenarios, encompassing diverse objective types and varying numbers of objectives. The effectiveness of MLPS was further validated through real-world challenges in discovering antifungal peptides with low toxicity and high activity.
Xuanbai Ren, Yuansheng Liu, Bosheng Song, Xiangxiang Zeng, Hisao Ishibuchi
AAAI6
2025 Towards Synergistic Path-based Explanations for Knowledge Graph Completion: Exploration and Evaluation
abstract
Knowledge graph completion (KGC) aims to alleviate the inherent incompleteness of knowledge graphs (KGs), a crucial task for numerous applications such as recommendation systems and drug repurposing. The success of knowledge graph embedding (KGE) models provokes the question about the explainability: ``\textit{Which the patterns of the input KG are most determinant to the prediction}?'' Particularly, path-based explainers prevail in existing methods because of their strong capability for human understanding. In this paper, based on the observation that a fact is usually determined by the synergy of multiple reasoning chains, we propose a novel explainable framework, dubbed KGExplainer, to explore synergistic pathways. KGExplainer is a model-agnostic approach that employs a perturbation-based greedy search algorithm to identify the most crucial synergistic paths as explanations within the local structure of target predictions. To evaluate the quality of these explanations, KGExplainer distills an evaluator from the target KGE model, allowing for the examination of their fidelity. We experimentally demonstrate that the distilled evaluator has comparable predictive performance to the target KGE. Experimental results on benchmark datasets demonstrate the effectiveness of KGExplainer, achieving a human evaluation accuracy of 83.3\% and showing promising improvements in explainability. Code is available at \url{https://github.com/xiaomingaaa/KGExplainer}
Tengfei Ma 0002, Xiang Song 0003, Wen Tao, Mufei Li, Jiani Zhang 0003, Xiaoqin Pan, Yijun Wang 0002, Bosheng Song, Xiangxiang Zeng
ICLR8
2025 Self-supervised Blending Structural Context of Visual Molecules for Robust Drug Interaction Prediction
abstract
Identifying drug-drug interactions (DDIs) is critical for ensuring drug safety and advancing drug development, a topic that has garnered significant research interest. While existing methods have made considerable progress, approaches relying solely on known DDIs face a key challenge when applied to drugs with limited data: insufficient exploration of the space of unlabeled pairwise drugs. To address these issues, we innovatively introduce S$^2$VM, a Self-supervised Visual pretraining framework for pair-wise Molecules, to fully fuse structural representations and explore the space of drug pairs for DDI prediction. S$^2$VM incorporates the explicit structure and correlations of visual molecules, such as the positional relationships and connectivity between functional substructures. Specifically, we blend the visual fragments of drug pairs into a unified input for joint encoding and then recover molecule-specific visual information for each drug individually. This approach integrates fine-grained structural representations from unlabeled drug pair data. By using visual fragments as anchors, S$^2$VM effectively captures the spatial information of local molecular components within visual molecules, resulting in more comprehensive embeddings of drug pairs. Experimental results show that S$^2$VM achieves state-of-the-art performance on widely used benchmarks, with Macro-F1 score improvements of 4.21% and 3.31%, respectively. Further extensive results and theoretical analysis demonstrate the effectiveness of S$^2$VM for both few-shot and novel drugs.
Tengfei Ma 0002, Yongsheng Zang, Yujie Chen 0002, Xuanbai Ren, Bosheng Song, Hongxin Xiang, Xiangxiang Zeng
NeurIPS6
2025 The computational properties of P systems with mutative membrane structures
Bosheng Song, Chuanlong Hu, David Orellana-Martín, Antonio Ramírez-de-Arellano, Mario J. Pérez-Jiménez, Xiangxiang Zeng
Inf. Comput.1
2025 Semantic substructure guided multiple objective molecular generation with discrete diffusion probabilistic model
Shugao Chen, Bosheng Song, Ruizhe Chen, Guifei Zhou, Sisi Yuan
Neurocomputing2
2024 SSR-DTA: Substructure-aware multi-layer graph neural networks for drug-target binding affinity prediction
Yuansheng Liu, Xinyan Xia, Yongshun Gong, Bosheng Song, Xiangxiang Zeng
Artif. Intell. Medicine4
2024 Attribute-guided prototype network for few-shot molecular property prediction
abstract
The molecular property prediction (MPP) plays a crucial role in the drug discovery process, providing valuable insights for molecule evaluation and screening. Although deep learning has achieved numerous advances in this area, its success often depends on the availability of substantial labeled data. The few-shot MPP is a more challenging scenario, which aims to identify unseen property with only few available molecules. In this paper, we propose an attribute-guided prototype network (APN) to address the challenge. APN first introduces an molecular attribute extractor, which can not only extract three different types of fingerprint attributes (single fingerprint attributes, dual fingerprint attributes, triplet fingerprint attributes) by considering seven circular-based, five path-based, and two substructure-based fingerprints, but also automatically extract deep attributes from self-supervised learning methods. Furthermore, APN designs the Attribute-Guided Dual-channel Attention module to learn the relationship between the molecular graphs and attributes and refine the local and global representation of the molecules. Compared with existing works, APN leverages high-level human-defined attributes and helps the model to explicitly generalize knowledge in molecular graphs. Experiments on benchmark datasets show that APN can achieve state-of-the-art performance in most cases and demonstrate that the attributes are effective for improving few-shot MPP performance. In addition, the strong generalization ability of APN is verified by conducting experiments on data from different domains.
Linlin Hou, Hongxin Xiang, Xiangxiang Zeng, Dong-Sheng Cao 0001, Bosheng Song
Briefings Bioinform.6
2024 Deep synergetic spiking neural P systems for the overall survival time prediction of glioblastoma patients
Xiu Yin, Xiyu Liu 0001, Jinpeng Dai, Bosheng Song, Chunqiu Xia, Dengwang Li, Jie Xue 0001
Expert Syst. Appl.4
2024 Dynamic threshold spiking neural P systems with weights and multiple channels
Bosheng Song, Yuansheng Liu, Xiangxiang Zeng, Shengye Huang
Theor. Comput. Sci.2
2024 PEB-DDI: A Task-Specific Dual-View Substructural Learning Framework for Drug-Drug Interaction Prediction
abstract
Adverse drug-drug interactions (DDIs) pose potential risks in polypharmacy due to unknown physicochemical incompatibilities between co-administered drugs. Recent studies have utilized multi-layer graph neural network architectures to model hierarchical molecular substructures of drugs, achieving excellent DDI prediction performance. While extant substructural frameworks effectively encode interactions from atom-level features, they overlook valuable chemical bond representations within molecular graphs. More critically, given the multifaceted nature of DDI prediction tasks involving both known and novel drug combinations, previous methods lack tailored strategies to address these distinct scenarios. The resulting lack of adaptability impedes further improvements to model performance. To tackle these challenges, we propose PEB-DDI, a DDI prediction learning framework with enhanced substructure extraction. First, the information of chemical bonds is integrated and synchronously updated with the atomic nodes. Then, different dual-view strategies are selected based on whether novel drugs are present in the prediction task. Particularly, we constructed Molecular fingerprint-Molecular graph view for transductive task, and Bipartite graph-Molecular graph view for inductive task. Rigorous evaluations on benchmark datasets underscore PEB-DDI's superior performance. Notably, on DrugBank, it achieves an outstanding accuracy rate of 98.18% when predicting previously unknown interactions among approved drugs. Even when faced with novel drugs, PEB-DDI consistently exhibits outstanding generalization capabilities with an accuracy rate of 88.06%, attributing to the proper migrating of molecular basic structure learning.
Xiangzhen Shen, Yuansheng Liu, Bosheng Song, Xiangxiang Zeng
IEEE J. Biomed. Health Informatics4
2024 Learning to Denoise Biomedical Knowledge Graph for Robust Molecular Interaction Prediction
abstract
Molecular interaction prediction plays a crucial role in forecasting unknown interactions between molecules, such as drug-target interaction (DTI) and drug-drug interaction (DDI), which are essential in the field of drug discovery and therapeutics. Although previous prediction methods have yielded promising results by leveraging the rich semantics and topological structure of biomedical knowledge graphs (KGs), they have primarily focused on enhancing predictive performance without addressing the presence of inevitable noise and inconsistent semantics. This limitation has hindered the advancement of KG-based prediction methods. To address this limitation, we propose BioKDN (BiomedicalKnowledge GraphDenoisingNetwork) for robust molecular interaction prediction. BioKDN refines the reliable structure of local subgraphs by denoising noisy links in a learnable manner, providing a general module for extracting task-relevant interactions. To enhance the reliability of the refined structure, BioKDN maintains consistent and robust semantics by smoothing relations around the target interaction. By maximizing the mutual information between reliable structure and smoothed relations, BioKDN emphasizes informative semantics to enable precise predictions. Experimental results on real-world datasets show that BioKDN surpasses state-of-the-art models in DTI and DDI prediction tasks, confirming the effectiveness and robustness of BioKDN in denoising unreliable interactions within contaminated KGs.
Tengfei Ma 0002, Yujie Chen 0002, Wen Tao, Dashun Zheng, Xuan Lin, Patrick Pang 0001, Yijun Wang 0002, Longyue Wang, Bosheng Song, Xiangxiang Zeng, Philip S. Yu
IEEE Trans. Knowl. Data Eng.10
2023 Dimensionality reduction and visualization of single-cell RNA-seq data with an improved deep variational autoencoder
abstract
Single-cell RNA sequencing (scRNA-seq) is a revolutionary breakthrough that determines the precise gene expressions on individual cells and deciphers cell heterogeneity and subpopulations. However, scRNA-seq data are much noisier than traditional high-throughput RNA-seq data because of technical limitations, leading to many scRNA-seq data studies about dimensionality reduction and visualization remaining at the basic data-stacking stage. In this study, we propose an improved variational autoencoder model (termed DREAM) for dimensionality reduction and a visual analysis of scRNA-seq data. Here, DREAM combines the variational autoencoder and Gaussian mixture model for cell type identification, meanwhile explicitly solving 'dropout' events by introducing the zero-inflated layer to obtain the low-dimensional representation that describes the changes in the original scRNA-seq dataset. Benchmarking comparisons across nine scRNA-seq datasets show that DREAM outperforms four state-of-the-art methods on average. Moreover, we prove that DREAM can accurately capture the expression dynamics of human preimplantation embryonic development. DREAM is implemented in Python, freely available via the GitHub website, https://github.com/Crystal-JJ/DREAM.
Junlin Xu, Yuansheng Liu, Bosheng Song, Xiulan Guo, Xiangxiang Zeng, Quan Zou 0001
Briefings Bioinform.4
2023 Comprehensive evaluation of deep and graph learning on drug-drug interactions prediction
abstract
Recent advances and achievements of artificial intelligence (AI) as well as deep and graph learning models have established their usefulness in biomedical applications, especially in drug-drug interactions (DDIs). DDIs refer to a change in the effect of one drug to the presence of another drug in the human body, which plays an essential role in drug discovery and clinical research. DDIs prediction through traditional clinical trials and experiments is an expensive and time-consuming process. To correctly apply the advanced AI and deep learning, the developer and user meet various challenges such as the availability and encoding of data resources, and the design of computational methods. This review summarizes chemical structure based, network based, natural language processing based and hybrid methods, providing an updated and accessible guide to the broad researchers and development community with different domain knowledge. We introduce widely used molecular representation and describe the theoretical frameworks of graph neural network models for representing molecular structures. We present the advantages and disadvantages of deep and graph learning methods by performing comparative experiments. We discuss the potential technical challenges and highlight future directions of deep and graph learning models for accelerating DDIs prediction.
Xuan Lin, Lichang Dai, Yafang Zhou, Jianyu Shi, Dong-Sheng Cao 0001, Bosheng Song, Philip S. Yu, Xiangxiang Zeng
Briefings Bioinform.10
2023 Sequence Alignment/Map format: a comprehensive review of approaches and applications
abstract
The Sequence Alignment/Map (SAM) format file is the text file used to record alignment information. Alignment is the core of sequencing analysis, and downstream tasks accept mapping results for further processing. Given the rapid development of the sequencing industry today, a comprehensive understanding of the SAM format and related tools is necessary to meet the challenges of data processing and analysis. This paper is devoted to retrieving knowledge in the broad field of SAM. First, the format of SAM is introduced to understand the overall process of the sequencing analysis. Then, existing work is systematically classified in accordance with generation, compression and application, and the involved SAM tools are specifically mined. Lastly, a summary and some thoughts on future directions are provided.
Yuansheng Liu, Xiangzhen Shen, Yongshun Gong, Bosheng Song, Xiangxiang Zeng
Briefings Bioinform.5
2023 Prediction of multi-relational drug-gene interaction via Dynamic hyperGraph Contrastive Learning
abstract
Drug-gene interaction prediction occupies a crucial position in various areas of drug discovery, such as drug repurposing, lead discovery and off-target detection. Previous studies show good performance, but they are limited to exploring the binding interactions and ignoring the other interaction relationships. Graph neural networks have emerged as promising approaches owing to their powerful capability of modeling correlations under drug-gene bipartite graphs. Despite the widespread adoption of graph neural network-based methods, many of them experience performance degradation in situations where high-quality and sufficient training data are unavailable. Unfortunately, in practical drug discovery scenarios, interaction data are often sparse and noisy, which may lead to unsatisfactory results. To undertake the above challenges, we propose a novel Dynamic hyperGraph Contrastive Learning (DGCL) framework that exploits local and global relationships between drugs and genes. Specifically, graph convolutions are adopted to extract explicit local relations among drugs and genes. Meanwhile, the cooperation of dynamic hypergraph structure learning and hypergraph message passing enables the model to aggregate information in a global region. With flexible global-level messages, a self-augmented contrastive learning component is designed to constrain hypergraph structure learning and enhance the discrimination of drug/gene representations. Experiments conducted on three datasets show that DGCL is superior to eight state-of-the-art methods and notably gains a 7.6% performance improvement on the DGIdb dataset. Further analyses verify the robustness of DGCL for alleviating data sparsity and over-smoothing issues.
Wen Tao, Yuansheng Liu, Xuan Lin, Bosheng Song, Xiangxiang Zeng
Briefings Bioinform.4
2023 Monodirectional evolutional symport tissue P systems with channel states and cell division
Bosheng Song, Kenli Li 0001, Xiangxiang Zeng, Mario J. Pérez-Jiménez, Claudio Zandron
Sci. China Inf. Sci.1
2023 Spiking neural P system with synaptic vesicles and applications in multiple brain metastasis segmentation
Jie Xue 0001, Deting Kong, Liwen Ren, Bosheng Song, Xiyu Liu 0001, Guanzhong Gong, Dengwang Li
Inf. Sci.4
2023 On the computational efficiency of tissue P systems with evolutional symport/antiport rules
abstract
Tissue P systems with evolutional symport/antiport rules are a variant of tissue P systems, where objects are communicated through regions by symport/antiport rules, and objects may evolve during this process. It is known that such systems are able to solve NP problems in a polynomial time (and exponential space), when cell division is allowed. In this work, we continue to investigate the computational complexity aspects for tissue P systems with evolutional symport/antiport rules. We prove that problems beyond NP can also be solved. In particular, we show that deterministic systems of this type are able to solve all problems in the complexity class PP. Moreover, if non-deterministic systems are considered, then all problems in the class PSPACE can be solved.
Linqiang Pan, Bosheng Song, Claudio Zandron
Knowl. Based Syst.2
2023 Tissue P systems with evolutional communication rules with two objects in the left-hand side
abstract
Abstract In the framework of Membrane Computing, several efficient solutions to computationally hard problems have been given. To find new borderlines between families of P systems that can solve them and the ones that cannot is an important task to tackle the P versus NP problem. Adding syntactic and/or semantic ingredients can mean passing from non-efficiency to presumed efficiency. Here, we try to get narrow frontiers, setting the stage to adapt efficient solutions from a family of P systems to another one. In order to do that, a solution to the problem is given by means of a family of tissue P systems with evolutional symport/antiport rules and cell separation with the restriction that both the left-hand side and the right-hand side of the rules have at most two objects; that is, with recognizer P systems from $${\mathcal {TSEC}}(2, 2)$$ TSEC ( 2 , 2 ) . This result improves a previous one, when 3 objects could be used in the left-hand side of the evolutional communication rules
David Orellana-Martín, Luis Valencia-Cabrera, Bosheng Song, Linqiang Pan, Mario J. Pérez-Jiménez
Nat. Comput.3
2023 Hybrid neural-like P systems with evolutionary channels for multiple brain metastases segmentation
abstract
Neural-like P systems are membrane computing models inspired by natural computing. Spiking neurral (SN) P systems, a kind of neural-like P systems, are viewed as third-generation neural network models. Although real neurons have complex structures, classical SN P systems simplify the structures and corresponding mechanisms to stationary two-dimensional graphs and lack related evolution mechanisms on spikes and channels, which limits the real applications of these models. In this paper, we propose a new hybrid SN P system with evolutionary channels (HN P systems), including three new types of rules for dynamically generating or removing one-one and one-many/many-one channels with related evolutions of spikes on the hybrid neuron structures. Two dynamic regulatory factors are also presented on rules to help guide the optimization of the HN P systems automatically. Based on the new P system, a multiple brain metastases (BMs) segmentation model is developed. The experimental results indicate that the proposed models outperform the state-of-the-art methods on the BMs, which have large variations in sizes, positions and shapes, and low contrast with their surroundings. Performances on the head and neck segmentation dataset also verifies the effectiveness of the HN P system.
Jie Xue 0001, Xiyu Liu 0001, Bosheng Song, Pu Huang 0001, Qiong An, Guanzhong Gong, Dengwang Li
Pattern Recognit.6
2023 Tissue P Systems With States in Cells
abstract
Tissue-like P systems with channel states are a type of classical membrane systems in which objects transferred among regions are controlled by states placed in the channels between regions. However, an important biological fact is the existence of a “barrier” to the diffusion of signal molecules, which tend to remain confined to some particular micro-habitat. This feature allows quorum sensing to convey information about the physiological state of spatially separated sub-populations. Therefore, in this article, we design a novel class-variant of P systems namedtissue P systems with states in cells(TSIC P systems). Here, each cell contains one and only one state at any moment (the environment has no state), and objects transferred among regions are controlled by states (or a state) that are placed in the corresponding cells (or a cell). We discuss thecomputability theoryof TSIC P systems by showing that Turing universality is acquired by TSIC P systems, which are worked both in a flat maximal parallelism and in a maximal parallelism. In addition, when cell division is considered in TSIC P systems, then tissue P systems with states in cells and cell division (TSICD P systems) are constructed. The (presumed)computational efficiencyof TSICD P systems is reached by offering a uniform solution to the satisfiability problem.
Bosheng Song, Kenli Li 0001, David Orellana-Martín, Xiangxiang Zeng, Mario J. Pérez-Jiménez
IEEE Trans. Computers1
2023 Modality-DTA: Multimodality Fusion Strategy for Drug-Target Affinity Prediction
abstract
Prediction of the drug-target affinity (DTA) plays an important role in drug discovery. Existing deep learning methods for DTA prediction typically leverage a single modality, namely simplified molecular input line entry specification (SMILES) or amino acid sequence to learn representations. SMILES or amino acid sequences can be encoded into different modalities. Multimodality data provide different kinds of information, with complementary roles for DTA prediction. We propose Modality-DTA, a novel deep learning method for DTA prediction that leverages the multimodality of drugs and targets. A group of backward propagation neural networks is applied to ensure the completeness of the reconstruction process from the latent feature representation to original multimodality data. The tag between the drug and target is used to reduce the noise information in the latent representation from multimodality data. Experiments on three benchmark datasets show that our Modality-DTA outperforms existing methods in all metrics. Modality-DTA reduces the mean square error by 15.7% and improves the area under the precisionrecall curve by 12.74% in the Davis dataset. We further find that the drug modality Morgan fingerprint and the target modality generated by one-hot-encoding play the most significant roles. To the best of our knowledge, Modality-DTA is the first method to explore multimodality for DTA prediction.
Xixi Yang, Zhangming Niu, Yuansheng Liu, Bosheng Song, Weiqiang Lu, Xiangxiang Zeng
IEEE ACM Trans. Comput. Biol. Bioinform.4
2023 Spiking neural P systems with weights and delays on synapses
Bosheng Song, Xiangxiang Zeng
Theor. Comput. Sci.2
2023 KG-MTL: Knowledge Graph Enhanced Multi-Task Learning for Molecular Interaction
abstract
Molecular interaction prediction is essential in various applications including drug discovery and material science. The problem becomes quite challenging when the interaction is represented by unmapped relationships in molecular networks, namely molecular interaction, because it easily suffers from (i) insufficient labeled data with many false-positive samples, and (ii) ignoring a large number of biological entities with rich information in the knowledge graph. Most of the existing methods cannot properly exploit the information of knowledge graph and molecule graph simultaneously. In this paper, we propose a large-scaleKnowledgeGraph enhancedMulti-TaskLearning model, namely KG-MTL, which extracts the features from both knowledge graph and molecular graph in a synergistic way. Moreover, we design an effectiveShared Unitthat helps the model to jointly preserve the semantic relations of drug entity and the neighbor structures of the compound in both knowledge graph and molecular graph. Extensive experiments on four real-world datasets demonstrate that our proposed KG-MTL outperforms the state-of-the-art methods on two representative molecular interaction prediction tasks: drug-target interaction prediction and compound-protein interaction prediction. The source code of KG-MTL is available athttps://github.com/xzenglab/KG-MTL.
Tengfei Ma 0002, Xuan Lin, Bosheng Song, Philip S. Yu, Xiangxiang Zeng
IEEE Trans. Knowl. Data Eng.3
2023 Hypergraph-Based Numerical Neural-Like P Systems for Medical Image Segmentation
abstract
Neural-like P systems are membrane computing models inspired by natural computing and are viewed as third-generation neural network models. Although real neurons have complex structures, classical neural-like P systems simplify the structures and corresponding mechanisms to two-dimensional graphs or tree-based firing and forgetting communications, which limit the real applications of these models. In this paper, we propose a hypergraph-based numerical neural-like (HNN) P system containing five types of neurons to describe the high-order correlations among neuron structures. Three new kinds of communication mechanisms among neurons are also proposed to address numerical variables and functions. Based on the new neural-like P system, a tumor/organ segmentation model for medical images is developed. The experimental results indicate that the proposed models outperform the state-of-the-art methods based on two hippocampal datasets and a multiple brain metastases dataset, thus verifying the effectiveness of the HNN P system in correctly segmenting tumors/organs.
Jie Xue 0001, Liwen Ren, Bosheng Song, Xiyu Liu 0001, Guanzhong Gong, Dengwang Li
IEEE Trans. Parallel Distributed Syst.3
2022 Learning spatial structures of proteins improves protein-protein interaction prediction
abstract
Spatial structures of proteins are closely related to protein functions. Integrating protein structures improves the performance of protein-protein interaction (PPI) prediction. However, the limited quantity of known protein structures restricts the application of structure-based prediction methods. Utilizing the predicted protein structure information is a promising method to improve the performance of sequence-based prediction methods. We propose a novel end-to-end framework, TAGPPI, to predict PPIs using protein sequence alone. TAGPPI extracts multi-dimensional features by employing 1D convolution operation on protein sequences and graph learning method on contact maps constructed from AlphaFold. A contact map contains abundant spatial structure information, which is difficult to obtain from 1D sequence data directly. We further demonstrate that the spatial information learned from contact maps improves the ability of TAGPPI in PPI prediction tasks. We compare the performance of TAGPPI with those of nine state-of-the-art sequence-based methods, and TAGPPI outperforms such methods in all metrics. To the best of our knowledge, this is the first method to use the predicted protein topology structure graph for sequence-based PPI prediction. More importantly, our proposed architecture could be extended to other prediction tasks related to proteins.
Bosheng Song, Xiaoyan Luo, Xiaoli Luo, Yuansheng Liu, Zhangming Niu, Xiangxiang Zeng
Briefings Bioinform.1
2022 Predicting ncRNA-protein interactions based on dual graph convolutional network and pairwise learning
abstract
Noncoding RNAs (ncRNAs) have recently attracted considerable attention due to their key roles in biology. The ncRNA-proteins interaction (NPI) is often explored to reveal some biological activities that ncRNA may affect, such as biological traits, diseases, etc. Traditional experimental methods can accomplish this work but are often labor-intensive and expensive. Machine learning and deep learning methods have achieved great success by exploiting sufficient sequence or structure information. Graph Neural Network (GNN)-based methods consider the topology in ncRNA-protein graphs and perform well on tasks like NPI prediction. Based on GNN, some pairwise constraint methods have been developed to apply on homogeneous networks, but not used for NPI prediction on heterogeneous networks. In this paper, we construct a pairwise constrained NPI predictor based on dual Graph Convolutional Network (GCN) called NPI-DGCN. To our knowledge, our method is the first to train a heterogeneous graph-based model using a pairwise learning strategy. Instead of binary classification, we use a rank layer to calculate the score of an ncRNA-protein pair. Moreover, our model is the first to predict NPIs on the ncRNA-protein bipartite graph rather than the homogeneous graph. We transform the original ncRNA-protein bipartite graph into two homogenous graphs on which to explore second-order implicit relationships. At the same time, we model direct interactions between two homogenous graphs to explore explicit relationships. Experimental results on the four standard datasets indicate that our method achieves competitive performance with other state-of-the-art methods. And the model is available at https://github.com/zhuoninnin1992/NPIPredict.
Linlin Zhuo, Bosheng Song, Yuansheng Liu, Xiangzheng Fu
Briefings Bioinform.2
2022 Rule synchronization for monodirectional tissue-like P systems with channel states
Bosheng Song, Xiangxiang Zeng
Inf. Comput.2
2022 Monodirectional Evolutional Symport Tissue P Systems With Promoters and Cell Division
abstract
Monodirectional tissue P systems with promoters are natural inspired parallel computing paradigms, where only symport rules are permitted, and with the restriction of “monodirectionality”, objects for two given regions are transferred in one direction. In this article, a novel kind of P systems, monodirectional evolutional symport tissue P systems with promoters (MESTP P systems) is raised, where objects may be revised during the movement between two regions. The computational theory of MESTP P systems that rules are employed in a flat maximally parallel pattern is investigated. We prove that finite natural number sets are created by MESTP P systems applying one cell, at most 1 promoter and all evolutional symport rules having a maximal length 2 or with arbitrary number of cells, promoters and all evolutional symport rules having a maximal length 2. MESTP P systems are Turing universal when two cells, at most 1 promoter and all evolutional symport rules having a maximal length 2 are employed. In addition, with the help of cell division mechanism, monodirectional evolutional symport tissue P systems with promoters and cell division (MESTPD P systems) are employed to solve NP-complete (the SAT) problem, where system uses at most 1 promoter and all evolutional symport rules having a maximal length 3. These results show that MESTP(D) P systems are still computationally powerful even if monodirectionality control mechanism is imposed, thereby developing membrane algorithms for MESTP(D) P systems is theoretically possible as well as potentially exploitable.
Bosheng Song, Kenli Li 0001, Xiangxiang Zeng
IEEE Trans. Parallel Distributed Syst.1
2021 Molecular design in drug discovery: a comprehensive review of deep generative models
abstract
Deep generative models have been an upsurge in the deep learning community since they were proposed. These models are designed for generating new synthetic data including images, videos and texts by fitting the data approximate distributions. In the last few years, deep generative models have shown superior performance in drug discovery especially de novo molecular design. In this study, deep generative models are reviewed to witness the recent advances of de novo molecular design for drug discovery. In addition, we divide those models into two categories based on molecular representations in silico. Then these two classical types of models are reported in detail and discussed about both pros and cons. We also indicate the current challenges in deep generative models for de novo molecular design. De novo molecular design automatically is promising but a long road to be explored.
Yongshun Gong, Yuansheng Liu, Bosheng Song, Quan Zou 0001
Briefings Bioinform.4
2021 Deep learning methods for biomedical named entity recognition: a survey and qualitative comparison
abstract
The biomedical literature is growing rapidly, and the extraction of meaningful information from the large amount of literature is increasingly important. Biomedical named entity (BioNE) identification is one of the critical and fundamental tasks in biomedical text mining. Accurate identification of entities in the literature facilitates the performance of other tasks. Given that an end-to-end neural network can automatically extract features, several deep learning-based methods have been proposed for BioNE recognition (BioNER), yielding state-of-the-art performance. In this review, we comprehensively summarize deep learning-based methods for BioNER and datasets used in training and testing. The deep learning methods are classified into four categories: single neural network-based, multitask learning-based, transfer learning-based and hybrid model-based methods. They can be applied to BioNER in multiple domains, and the results are determined by the dataset size and type. Lastly, we discuss the future development and opportunities of BioNER methods.
Bosheng Song, Fen Li, Yuansheng Liu, Xiangxiang Zeng
Briefings Bioinform.1
2021 MUFFIN: multi-scale feature fusion for drug-drug interaction prediction
abstract
MOTIVATION: Adverse drug-drug interactions (DDIs) are crucial for drug research and mainly cause morbidity and mortality. Thus, the identification of potential DDIs is essential for doctors, patients and the society. Existing traditional machine learning models rely heavily on handcraft features and lack generalization. Recently, the deep learning approaches that can automatically learn drug features from the molecular graph or drug-related network have improved the ability of computational models to predict unknown DDIs. However, previous works utilized large labeled data and merely considered the structure or sequence information of drugs without considering the relations or topological information between drug and other biomedical objects (e.g. gene, disease and pathway), or considered knowledge graph (KG) without considering the information from the drug molecular structure. RESULTS: Accordingly, to effectively explore the joint effect of drug molecular structure and semantic information of drugs in knowledge graph for DDI prediction, we propose a multi-scale feature fusion deep learning model named MUFFIN. MUFFIN can jointly learn the drug representation based on both the drug-self structure information and the KG with rich bio-medical information. In MUFFIN, we designed a bi-level cross strategy that includes cross- and scalar-level components to fuse multi-modal features well. MUFFIN can alleviate the restriction of limited labeled data on deep learning models by crossing the features learned from large-scale KG and drug molecular graph. We evaluated our approach on three datasets and three different tasks including binary-class, multi-class and multi-label DDI prediction tasks. The results showed that MUFFIN outperformed other state-of-the-art baselines. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/xzenglab/MUFFIN.
Yujie Chen 0002, Tengfei Ma 0002, Xixi Yang, Jianmin Wang 0016, Bosheng Song, Xiangxiang Zeng
Bioinform.5
2021 Time-free Solution to Independent Set Problem using P Systems with Active Membranes
abstract
Membrane computing is a branch of natural computingwhich abstracts fromthe structure and the functioning of living cells. The computation models obtained in the field of membrane computing are usually called P systems. P systems have been used to solve computationally hard problems efficiently on the assumption that the execution of each rule is completed in exactly one time-unit (a global clock is assumed for timing and synchronizing the execution of rules). However, in biological reality, different biological processes take different times to be completed, which can also be influenced by many environmental factors. In this work, with this biological reality, we give a time-free solution to independent set problemusing P systems with active membranes, which solve the problem independent of the execution time of the involved rules.
Bosheng Song
Fundam. Informaticae2
2021 The computational power of monodirectional tissue P systems with symport rules
Bosheng Song, Shengye Huang, Xiangxiang Zeng
Inf. Comput.1
2021 Rule synchronization for tissue P systems
Bosheng Song, Linqiang Pan
Inf. Comput.1
2021 Monodirectional tissue P systems with channel states
Bosheng Song, Xiangxiang Zeng, Alfonso Rodríguez-Patón
Inf. Sci.1
2021 Monodirectional Tissue P Systems With Promoters
abstract
Tissue P systems with promoters provide nondeterministic parallel bioinspired devices that evolve by the interchange of objects between regions, determined by the existence of some special objects called promoters. However, in cellular biology, the movement of molecules across a membrane is transported from high to low concentration. Inspired by this biological fact, in this article, an interesting type of tissue P systems, called monodirectional tissue P systems with promoters, where communication happens between two regions only in one direction, is considered. Results show that finite sets of numbers are produced by such P systems with one cell, using any length of symport rules or with any number of cells, using a maximal length 1 of symport rules, and working in the maximally parallel mode. Monodirectional tissue P systems are Turing universal with two cells, a maximal length 2, and at most one promoter for each symport rule, and working in the maximally parallel mode or with three cells, a maximal length 1, and at most one promoter for each symport rule, and working in the flat maximally parallel mode. We also prove that monodirectional tissue P systems with two cells, a maximal length 1, and at most one promoter for each symport rule (under certain restrictive conditions) working in the flat maximally parallel mode characterizes regular sets of natural numbers. Besides, the computational efficiency of monodirectional tissue P systems with promoters is analyzed when cell division rules are incorporated. Different uniform solutions to the Boolean satisfiability problem (SAT problem) are provided. These results show that with the restrictive condition of "monodirectionality," monodirectional tissue P systems with promoters are still computationally powerful. With the powerful computational power, developing membrane algorithms for monodirectional tissue P systems with promoters is potentially exploitable.
Bosheng Song, Xiangxiang Zeng, Min Jiang 0005, Mario J. Pérez-Jiménez
IEEE Trans. Cybern.1
2020 P Systems with Rule Production and Removal
abstract
P systems are a class of parallel computational models inspired by the structure and functioning of living cells, where all the evolution rules used in a system are initially set up and keep unchanged during a computation. In this work, inspired by the fact that chemical reactions in a cell can be affected by both the contents of the cell and the environmental conditions, we introduce a variant of P systems, called P systems with rule production and removal (abbreviated as RPR P systems), where rules in a system are dynamically changed during a computation, that is, at any computation step new rules can be produced and some existing rules can be removed. The computational power of RPR P systems and catalytic RPR P systems is investigated. Specifically, it is proved that catalytic RPR P systems with one catalyst and one membrane are Turing universal; for purely catalytic RPR P systems, one membrane and two catalysts are enough for reaching Turing universality. Moreover, a uniform solution to the SAT problem is provided by using RPR P systems with membrane division. It is known that standard catalytic P systems with one catalyst and one membrane are not Turing universal. These results imply that rule production and removal is a powerful feature for the computational power of P systems.
Linqiang Pan, Bosheng Song
Fundam. Informaticae2
2020 Cell-like P systems with evolutional symport/antiport rules and membrane creation
Bosheng Song, Kenli Li 0001, David Orellana-Martín, Luis Valencia-Cabrera, Mario J. Pérez-Jiménez
Inf. Comput.1
2020 Time-freeness and clock-freeness and related concepts in P systems
abstract
In the majority of models of P systems , rules are applied at the ticks of a global clock and their products are introduced into the system for the following step. In timed P systems, different integer durations are statically assigned to rules; time-free P systems are P systems yielding the same languages independently of these durations. In clock-free P systems, durations are real and are assigned to individual rule applications; thus, different applications of the same rule may last for a different amount of time. In this paper, we formalise timed, time-free, and clock-free P system within a framework for generalised parallel rewriting. We then explore the relationship between these variants of semantics. We show that clock-free P systems cannot efficiently solve intractable problems. Moreover, we consider un-timed systems where we collect the results using arbitrary timing functions as well as un-clocked P systems where we take the union over all possible per-instance rule durations. Finally, we also introduce and study mode-free P systems, whose results do not depend on the choice of a mode within a fixed family of modes, and compare mode-freeness with clock-freeness.
Artiom Alhazov, Rudolf Freund, Sergiu Ivanov 0001, Linqiang Pan, Bosheng Song
Theor. Comput. Sci.5
2020 P systems with symport/antiport rules: When do the surroundings matter?
David Orellana-Martín, Miguel A. Martínez-del-Amor, Luis Valencia-Cabrera, Bosheng Song, Linqiang Pan, Mario J. Pérez-Jiménez
Theor. Comput. Sci.4
2020 Cell-like P systems with polarizations and minimal rules
Linqiang Pan, David Orellana-Martín, Bosheng Song, Mario J. Pérez-Jiménez
Theor. Comput. Sci.3
2018 Language generating alphabetic flat splicing P systems
Linqiang Pan, Bosheng Song, Atulya K. Nagar, K. G. Subramanian 0001
Theor. Comput. Sci.2
2017 An efficient time-free solution to QSAT problem using P systems with proteins on membranes
Bosheng Song, Mario J. Pérez-Jiménez, Linqiang Pan
Inf. Comput.1
2017 Tissue-like P systems with evolutional symport/antiport rules
Bosheng Song, Cheng Zhang 0017, Linqiang Pan
Inf. Sci.1
2017 A time-free uniform solution to subset sum problem by tissue P systems with cell division
abstract
Tissue P systems are a class of bio-inspired computing models motivated by biochemical interactions between cells in a tissue-like arrangement. Tissue P systems with cell division offer a theoretical device to generate an exponentially growing structure in order to solve computationally hard problems efficiently with the assumption that there exists a global clock to mark the time for the system, the execution of each rule is completed in exactly one time unit. Actually, the execution time of different biochemical reactions in cells depends on many uncertain factors. In this work, with this biological inspiration, we remove the restriction on the execution time of each rule, and the computational efficiency of tissue P systems with cell division is investigated. Specifically, we solve subset sum problem by tissue P systems with cell division in a time-free manner in the sense that the correctness of the solution to the problem does not depend on the execution time of the involved rules.
Bosheng Song, Tao Song 0001, Linqiang Pan
Math. Struct. Comput. Sci.1
2016 Tissue P Systems with Protein on Cells
abstract
Tissue P systems are a class of distributed parallel computing devices inspired by biochemical interactions between cells in a tissue-like arrangement, where objects can be exchanged by means of communication channels. In this work, inspired by the biological facts that the movement of most objects through communication channels is controlled by proteins and proteins can move through lipid bilayers between cells (if these cells are fused), we present a new class of variant tissue P systems, called tissue P systems with protein on cells, where multisets of objects (maybe empty), together with proteins between cells are exchanged. The computational power of such P systems is studied. Specifically, an efficient (uniform) solution to the SAT problem by using such P systems with cell division is presented. We also prove that any Turing computable set of numbers can be generated by a tissue P system with protein on cells. Both of these two results are obtained by such P systems with communication rules of length at most 4 (the length of a communication rule is the total number of objects and proteins involved in that rule).
Bosheng Song, Linqiang Pan, Mario J. Pérez-Jiménez
Fundam. Informaticae1
2016 An efficient time-free solution to SAT problem by P systems with proteins on membranes
Bosheng Song, Mario J. Pérez-Jiménez, Linqiang Pan
J. Comput. Syst. Sci.1
2016 Flat maximal parallelism in P systems with promoters
Linqiang Pan, Gheorghe Paun, Bosheng Song
Theor. Comput. Sci.3
2016 The computational power of tissue-like P systems with promoters
Bosheng Song, Linqiang Pan
Theor. Comput. Sci.1
2015 A P_Lingua Based Simulator for P Systems with Symport/Antiport Rules
abstract
Inspired by mitosis process and membrane fission processes, cell-like P systems with symport/antiport rules and membrane division rules or membrane separation rules have been introduced, respectively. These computation systems have two key features: the ability to have infinite copies of some objects (within an active environment) and to generate an exponential workspace in polynomial time. In this work, we extend the P-Lingua framework for simulating that kind of P systems taking into account these two features. Consequently, a new simulator has been developed and included in pLinguaCore library. The functioning of the simulator has been checked by simulating efficient solutions to SAT problem using a family of cell-like P systems with symport/antiport rules and membrane division rules or membrane separation rules. The corresponding MeCoSim based application is also provided.
Luis F. Macías-Ramos, Luis Valencia-Cabrera, Bosheng Song, Tao Song 0001, Linqiang Pan, Mario J. Pérez-Jiménez
Fundam. Informaticae3
2015 Time-free solution to SAT problem by P systems with active membranes and standard cell division rules
Bosheng Song, Tao Song 0001, Linqiang Pan
Nat. Comput.1
2015 Computational efficiency and universality of timed P systems with membrane creation
Bosheng Song, Mario J. Pérez-Jiménez, Linqiang Pan
Soft Comput.1
2015 Computational efficiency and universality of timed P systems with active membranes
Bosheng Song, Linqiang Pan
Theor. Comput. Sci.1