Xiaochen Bo

dblp:94/3428 · DBLP profile ↗
← Back
65ranked-venue papers
3as first author
40since 2021 · last 2026
0000-0003-1911-7922ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 51 · 2 first-author · 30 since 2021Artificial intelligence and machine learning · 9 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An Edge-Pursuit Ising Machine with Programmable Local Fields and Adaptive Annealing
Bocheng Xu, Zihan Wu 0005, Xiyuan Tang, Xiaochen Bo, Yuan Wang 0001
ISCAS5
2026 MuFaDDG: a sequence-based multiscale feature fusion framework for protein stability changes prediction
abstract
MOTIVATION: Predicting the thermodynamic stability of proteins upon single-point mutations is a pivotal step in both protein engineering and medicine. In the study of predicting protein thermodynamic stability, various computational methods, whether they extract features at the local-level or global-level, exhibit their respective advantages and limitations. To leverage the advantages of both features, we developed MuFaDDG, a novel sequence-based method that integrated multiscale feature fusion for improved prediction of protein stability changes (ΔΔG). RESULTS: MuFaDDG achieves comparable performance on the S669 benchmark, demonstrating strong capabilities in stabilizing mutations. Notably, it shows a significant advantage in the ACC metric, with values of 0.75, 0.88, and 0.81 on the direct, reverse, and overall datasets of the CAGI5 Challenge's Frataxin, respectively. Furthermore, our method outperforms leading sequence-based approaches including THPLM, DDGemb, DDGun, and INPS-Seq on protein Myoglobin stability prediction. Additionally, MuFaDDG demonstrates exceptional predictive performance with higher PCC and ACC on the protein ThreeFoil, which is uncurated by FireProtDB and ProThermDB databases. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/PengjiaMa23/MuFaDDG.
Jianting Gong, Pengjia Ma, Zilin Ren, Zhiguo Fu, Xiaochen Bo
Bioinform.8
2026 MMPCS: multi-view molecular pretraining based on consistency information and specific information
abstract
MOTIVATION: The goal of molecular representation learning is to automate the extraction of molecular features, a critical task in cheminformatics and drug discovery. While pretraining models using multiple views like SMILES, 2D graphs, and 3D conformations have advanced the field, integrating them effectively to produce superior representations remains a challenge. RESULTS: To bridge this gap, we propose a novel multi-view molecular pretraining method termed MMPCS, which explicitly factorizes representations into consistency and specific information. Our approach utilizes the Graph Isomorphism Network and the RoBERTa model to encode 2D molecular topological graphs and SMILES sequences, respectively. Each resulting molecular embedding is decomposed into a shared consistency component and a view-specific remainder. An autoencoder then aligns the consistency information across views. The combined consistency and view-specific representations serve as input for downstream tasks, enabling precise and task-aware predictions. When benchmarked against 16 state-of-the-art molecular pretraining methods, MMPCS achieved the highest average performance across both classification and regression tasks for molecular property prediction. It also delivered outstanding results in predicting drug-target binding affinity and cancer drug response, demonstrating its robustness and broad applicability. Additionally, a case study on the SARS-CoV-2 Omicron variant highlights the potential of MMPCS in facilitating drug repurposing efforts. AVAILABILITY AND IMPLEMENTATION: The source code and datasets supporting this study are publicly available at GitHub (https://github.com/xmubiocode/MMPCS) and Zenodo (https://doi.org/10.5281/zenodo.18182748).
Chenyang Xie, Yingying Song, Xiaochen Bo, Zhongnan Zhang
Bioinform.4
2026 ALC-DRKG: an active learning-based framework for dynamic knowledge graph construction for drug repositioning
Shuaibin Liu, Huiyan Xu, Peng Zan, Xiaochen Bo
Inf. Process. Manag.6
2026 MOCT: A Multi-Class Oblique Tree Algorithm for Synergistic Drug Combination Prediction
abstract
Machine learning has been successfully applied to drug combination prediction in recent years. However, in some situations, the class imbalance problem still shows highly negative impacts on the modeling process, which cannot be directly handled by traditional methods. In addition, the interpretability of models is another key point for biological and medical experts. In this study, a clustering-based oblique decision tree (MOCT) algorithm is proposed to extract interpretable knowledge for the multi-class datasets. It firstly clusters samples of different classes, and then a proper feature subspace is generated to split data and forms a nonleaf node. Unlike traditional decision trees, our MOCT only grows one none-leaf node in each layer to generate a concise tree structure. Datasets of drug combinations were collected from three cell lines with three classes (Additive, Antagonism, and Synergy) in experiments, and the results show that our MOCT algorithm is superior to other methods with better interpretability.
Zhikai Lin, Lianlian Wu, Kunhong Liu 0001, Yong Xu 0009, Xiaochen Bo
IEEE Trans. Comput. Biol. Bioinform.7
2025 Geometric Diffusion Model Based on Stochastic Differential Equations for 3D Molecular Generation
abstract
In recent years, denoising diffusion models have exhibited exceptional performance in molecular generation tasks. However, existing approaches still face two major challenges. First, most diffusion-based methods rely on discretization operations and assume a Gaussian distribution, which often results in ambiguous or meaningless intermediate values when generating atom and bond types. Additionally, discretization introduces information loss, further compromising the quality of generated molecules. Second, current methods fail to incorporate constraints on bond angles and atomic valency, potentially leading to molecules that lack geometric consistency and chemical validity. To address these limitations, this paper introduces GeoSDE, a geometric diffusion model based on stochastic differential equations. Unlike conventional approaches, GeoSDE employs a continuous diffusion model that directly generates molecules in continuous space, effectively eliminating discretization-related issues and enhancing the stability and chemical validity of the generated structures. Moreover, GeoSDE explicitly incorporates bond angle and atomic valency constraints during training, providing richer molecular representations and ensuring that the generated molecules adhere more closely to real-world chemical principles and spatial structures. Extensive experiments on the GEOM-QM9 and GEOM-Drugs datasets demonstrate that GeoSDE surpasses existing state-of-the-art models across multiple metrics, particularly in terms of stability, validity, and uniqueness, highlighting its effectiveness in molecular generation tasks.
Xinyi Guan, Xiaochen Bo, Zhongnan Zhang
IJCNN3
2025 DeepPFP: a multi-task-aware architecture for protein function prediction
abstract
Deriving protein function from protein sequences poses a significant challenge due to the intricate relationship between sequence and function. Deep learning has made remarkable strides in predicting sequence-function relationships. However, models tailored for specific tasks or protein types encounter difficulties when using transfer learning across domains. This is attributed to the fact that protein function relies heavily on structural characteristics rather than mere sequence information. Consequently, there is a pressing need for a model capable of capturing shared features among diverse sequence-function mapping tasks to address the generalization issue. In this study, we explore the potential of Model-Agnostic Meta-Learning combined with a protein language model called Evolutionary Scale Modeling to tackle this challenge. Our approach involves training the architecture on five out-domain deep mutational scanning (DMS) datasets and evaluating its performance across four key dimensions. Our findings demonstrate that the proposed architecture exhibits satisfactory performance in terms of generalization and employs an effective few-shot learning strategy. To explain further, Compared to the best results, the Pearson's correlation coefficient (PCC) in the final stage increased by ~0.31%. Furthermore, we leverage the trained architecture to predict binding affinity scores of the DMS dataset of SARS-CoV-2 using transfer learning. Notably, training on a subset of the Ube4b dataset with 500 samples resulted in a notable improvement of 0.11 in the PCC. These results underscore the potential of our conceptual architecture as a promising methodology for multi-task protein function prediction.
Zilin Ren, Jinghong Sun, Yongbing Chen, Xiaochen Bo, Jiguo Xue, Jingyang Gao
Briefings Bioinform.5
2025 IECata: interpretable bilinear attention network and evidential deep learning improve the catalytic efficiency prediction of enzymes
abstract
Enzyme catalytic efficiency (kcat/Km) is a key parameter for identifying high-activity enzymes. Recently, deep learning techniques have demonstrated the potential for fast and accurate kcat/Km prediction. However, three challenges remain: (i) the limited size of the available kcat/Km dataset hinders the development of deep learning models; (ii) the model predictions lack reliable confidence estimates; and (iii) models lack interpretable insights into enzyme-catalyzed reactions. To address these challenges, we proposed IECata, a kcat/Km prediction model that provides uncertainty estimation and interpretability. IECata collected a dataset of 11 815 kcat/Km entries from the BRENDA and SABIO-RK databases, along with an out-of-domain test dataset of 806 entries from the literature. By introducing evidential deep learning, IECata provides uncertainty estimates for kcat/Km predictions. Moreover, it uses a bilinear attention mechanism to focus on learning crucial local interactions to interpret the key residues and substrate atoms in enzyme-catalyzed reactions. Testing results indicate that the prediction performance of IECata exceeds that of state-of-the-art benchmark models. More importantly, it provides a reliable confidence assessment for these predictions. Case studies further highlight that the incorporation of uncertainty in screening for highly active enzymes can effectively increase the hit ratio, thereby improving the efficiency of experimental validation and accelerating directed enzyme evolution. To facilitate researchers' use of IECata, we have developed an online prediction platform: http://mathtc.nscc-tj.cn/cataai/.
Yanpeng Zhao, Zhijiang Yang, Ge Yao, Penggang Han, Peng Zan, Xiukun Wan, Xiaochen Bo
Briefings Bioinform.10
2025 Multi-task multi-view and iterative error-correcting random forest for acute toxicity prediction
Lianlian Wu, Guangyi Lin, Jiayu Zou, Bowei Yan, Kunhong Liu 0001, Xiaochen Bo
Expert Syst. Appl.8
2025 Learning generic and specific prompts with contrastive constraints for multi-task visual scene understanding
Haohao Hu, Peng Zan, Xiaochen Bo
Neurocomputing8
2025 A feature pair-based neural network embedded decision tree for synergistic drug combination prediction
Jiayu Zou, Lianlian Wu, Kunhong Liu 0001, Yong Xu 0009, Xiaochen Bo
Pattern Recognit.6
2024 Exploring Optimal Transport-Based Multi-Grained Alignments for Text-Molecule Retrieval
abstract
The field of bioinformatics has seen significant progress, making the cross-modal text-molecule retrieval task increasingly vital. This task focuses on accurately retrieving molecule structures based on textual descriptions, by effectively aligning textual descriptions and molecules to assist researchers in identifying suitable molecular candidates. However, many existing approaches overlook the details inherent in molecule substructures. In this work, we introduce the Optimal Transport-based Multi-grained Alignments model (ORMA), a novel approach that facilitates multi-grained alignments between textual descriptions and molecules. Our model features a text encoder and a molecule encoder. The text encoder processes textual descriptions to generate both token-level and sentence-level representations, while molecules are modeled as hierarchical heterogeneous graphs, encompassing atom, motif, and molecule nodes to extract representations at these three levels. A key innovation in ORMA is the application of Optimal Transport (OT) to align tokens with motifs, creating multi-token representations that integrate multiple token alignments with their corresponding motifs. Additionally, we employ contrastive learning to refine cross-modal alignments at three distinct scales: token-atom, multitoken-motif, and sentence-molecule, ensuring that the similarities between correctly matched text-molecule pairs are maximized while those of unmatched pairs are minimized. To our knowledge, this is the first attempt to explore alignments at both the motif and multi-token levels. Experimental results on the ChEBI-20 and PCdes datasets demonstrate that ORMA significantly outperforms existing state-of-the-art (SOTA) models. Specifically, in text-molecule retrieval on ChEBI-20, our model achieves a Hits@1 score of 66.5%, surpassing the SOTA model AMAN by 17.1%. Similarly, in molecule-text retrieval, ORMA secures a Hits@1 score of 61.6%, outperforming AMAN by 15.0%.
Zijun Min, Bingshuai Liu, Jinsong Su, Xiaochen Bo
BIBM7
2024 Towards Cross-Modal Text-Molecule Retrieval with Better Modality Alignment
abstract
Cross-modal text-molecule retrieval model aims to learn a shared feature space of the text and molecule modalities for accurate similarity calculation, which facilitates the rapid screening of molecules with specific properties and activities in drug design. However, previous works have two main defects. First, they are inadequate in capturing modality-shared features considering the significant gap between text sequences and molecule graphs. Second, they mainly rely on contrastive learning and adversarial training for cross-modality alignment, both of which mainly focus on the first-order similarity, ignoring the second-order similarity that can capture more structural information in the embedding space. To address these issues, we propose a novel cross-modal text-molecule retrieval model with two-fold improvements. Specifically, on the top of two modality-specific encoders, we stack a memory bank based feature projector that contain learnable memory vectors to extract modality-shared features better. More importantly, during the model training, we calculate four kinds of similarity distributions (text-to-text, text-to-molecule, molecule-to-molecule, and molecule-to-text similarity distributions) for each instance, and then minimize the distance between these similarity distributions (namely second-order similarity losses) to enhance cross-modal alignment. Experimental results and analysis strongly demonstrate the effectiveness of our model. Particularly, our model achieves SOTA performance, outperforming the previously-reported best result by 6.4%.
Wanru Zhuang, Yujie Lin 0003, Chunyan Li 0002, Jinsong Su, Xiaochen Bo
BIBM8
2024 CPDT: A Novel Cluster-based Paired Decision Tree for Identifying Biomedical Entity Interactions
abstract
For the interaction prediction task in the biomedical field, most machine learning algorithms overlook the relationships between entities within a pair by treating their features independently. To address this issue, this paper proposes a novel Cluster-based Paired Decision Tree model (CPDT), which pairs synonymous features of entity pairs to form paired feature spaces for simultaneous processing. It employs an adaptive grid-based clustering algorithm to partition these spaces in an axis-parallel manner, constructing interpretable decision boundaries. Moreover, the clustering algorithm leverages the probability density function to accommodate various data distributions in paired feature spaces, enhancing the effectiveness of sample partitioning. Experimental results demonstrate that CPDT performs well in two interaction prediction tasks: Drug Combination and Synthetic Lethality predictions. Furthermore, CPDT yields simple and interpretable decision rules that uncover potential patterns in biomedical interaction prediction. It also identifies molecules with medical significance, suggesting promising applications in the biomedical domain.
Jiayu Zou, Lianlian Wu, Weiping Lin, Kunhong Liu 0001, Yong Xu 0009, Xiaochen Bo
BIBM7
2024 Stratifying TAD boundaries pinpoints focal genomic regions of regulation, damage, and repair
abstract
Advances in chromatin mapping have exposed the complex chromatin hierarchical organization in mammals, including topologically associating domains (TADs) and their substructures, yet the functional implications of this hierarchy in gene regulation and disease progression are not fully elucidated. Our study delves into the phenomenon of shared TAD boundaries, which are pivotal in maintaining the hierarchical chromatin structure and regulating gene activity. By integrating high-resolution Hi-C data, chromatin accessibility, and DNA double-strand breaks (DSBs) data from various cell lines, we systematically explore the complex regulatory landscape at high-level TAD boundaries. Our findings indicate that these boundaries are not only key architectural elements but also vibrant hubs, enriched with functionally crucial genes and complex transcription factor binding site-clustered regions. Moreover, they exhibit a pronounced enrichment of DSBs, suggesting a nuanced interplay between transcriptional regulation and genomic stability. Our research provides novel insights into the intricate relationship between the 3D genome structure, gene regulation, and DNA repair mechanisms, highlighting the role of shared TAD boundaries in maintaining genomic integrity and resilience against perturbations. The implications of our findings extend to understanding the complexities of genomic diseases and open new avenues for therapeutic interventions targeting the structural and functional integrity of TAD boundaries.
Bijia Chen, Zhangyi Ouyang, Jingxuan Xu, Hejiang Guo, Xuemei Bai, Mengge Tian, Hao Li 0035, Xiaochen Bo, Hebing Chen
Briefings Bioinform.13
2024 ReadCurrent: a VDCNN-based tool for fast and accurate nanopore selective sequencing
abstract
Nanopore selective sequencing allows the targeted sequencing of DNA of interest using computational approaches rather than experimental methods such as targeted multiplex polymerase chain reaction or hybridization capture. Compared to sequence-alignment strategies, deep learning (DL) models for classifying target and nontarget DNA provide large speed advantages. However, the relatively low accuracy of these DL-based tools hinders their application in nanopore selective sequencing. Here, we present a DL-based tool named ReadCurrent for nanopore selective sequencing, which takes electric currents as inputs. ReadCurrent employs a modified very deep convolutional neural network (VDCNN) architecture, enabling significantly lower computational costs for training and quicker inference compared to conventional VDCNN. We evaluated the performance of ReadCurrent across 10 nanopore sequencing datasets spanning human, yeasts, bacteria, and viruses. We observed that ReadCurrent achieved a mean accuracy of 98.57% for classification, outperforming four other DL-based selective sequencing methods. In experimental validation that selectively sequenced microbial DNA from human DNA, ReadCurrent achieved an enrichment ratio of 2.85, which was higher than the 2.7 ratio achieved by MinKNOW using the sequence-alignment strategy. In summary, ReadCurrent can rapidly classify target and nontarget DNA with high accuracy, providing an alternative in the toolbox for nanopore selective sequencing. ReadCurrent is available at https://github.com/Ming-Ni-Group/ReadCurrent.
Kechen Fan, Jiarong Zhang, Zihan Xie, Daguang Jiang, Xiaochen Bo, Shenghui Shi
Briefings Bioinform.6
2024 NASTRA: accurate analysis of short tandem repeat markers by nanopore sequencing with repeat-structure-aware algorithm
abstract
Short-tandem repeats (STRs) are the type of genetic markers extensively utilized in biomedical and forensic applications. Due to sequencing noise in nanopore sequencing, accurate analysis methods are lacking. We developed NASTRA, an innovative tool for Nanopore Autosomal Short Tandem Repeat Analysis, which overcomes traditional database-based methods' limitations and provides a precise germline analysis of STR genetic markers without the need for allele sequence reference. Demonstrating high accuracy in cell line authentication testing and paternity testing, NASTRA significantly surpasses existing methods in both speed and accuracy. This advancement makes it a promising solution for rapid cell line authentication and kinship testing, highlighting the potential of nanopore sequencing for in-field applications.
Zilin Ren, Jiarong Zhang, Jiguo Xue, Xiaochen Bo, Jiangwei Yan
Briefings Bioinform.7
2024 scDAC: deep adaptive clustering of single-cell transcriptomic data with coupled autoencoder and Dirichlet process mixture model
abstract
MOTIVATION: Clustering analysis for single-cell RNA sequencing (scRNA-seq) data is an important step in revealing cellular heterogeneity. Many clustering methods have been proposed to discover heterogenous cell types from scRNA-seq data. However, adaptive clustering with accurate cluster number reflecting intrinsic biology nature from large-scale scRNA-seq data remains quite challenging. RESULTS: Here, we propose a single-cell Deep Adaptive Clustering (scDAC) model by coupling the Autoencoder (AE) and the Dirichlet Process Mixture Model (DPMM). By jointly optimizing the model parameters of AE and DPMM, scDAC achieves adaptive clustering with accurate cluster numbers on scRNA-seq data. We verify the performance of scDAC on five subsampled datasets with different numbers of cell types and compare it with 15 widely used clustering methods across nine scRNA-seq datasets. Our results demonstrate that scDAC can adaptively find accurate numbers of cell types or subtypes and outperforms other methods. Moreover, the performance of scDAC is robust to hyperparameter changes. AVAILABILITY AND IMPLEMENTATION: The scDAC is implemented in Python. The source code is available at https://github.com/labomics/scDAC.
Sijing An, Jinhui Shi, Runyan Liu, Yaowen Chen, Shuofeng Hu, Xinyu Xia 0002, Guohua Dong, Xiaochen Bo, Xiaomin Ying
Bioinform.9
2024 Multi-view uncertainty deep forest: An innovative deep forest equipped with uncertainty estimation for drug-induced liver injury prediction
Yuqi Wen, Yong Xu 0009, Kunhong Liu 0001, Xiaochen Bo
Inf. Sci.6
2023 A Multi-View Learning-Based Bayesian Ruleset Extraction Algorithm For Accurate Hepatotoxicity Prediction
abstract
The interpretable machine learning method is important in drug discovery. Unlike traditional ensemble learning methods, this paper proposes an interpretable algorithm based on Bayesian rule extraction to obtain reliable and explainable results for hepatotoxicity prediction. To extract information from different types of omics data, our algorithm employs a multi-view learning strategy to enhance performance. Specifically, a random forest is trained in each view, and then the Bayesian rule extraction algorithm is designed to select an optimal rule subset, controlling the size and accuracy of the ruleset through probabilities. These rule sets are integrated through multi-view voting to get the final decisions. The performance of our algorithm is tested on the hepatotoxicity dataset, demonstrating that compared to traditional machine learning algorithms and rule-based algorithms, our approach maintains excellent performance while achieving high interpretability in most cases. Our python source code and the related Supplementary Materials are available at: github.com/MLDMXM2017/MV-BRS.
Lianlian Wu, Yong Xu 0009, Kunhong Liu 0001, Xiaochen Bo
BIBM6
2023 MoSCHG: Multi-omics Single-cell Classification based on Heterogeneous Graphs and Supervised Contrastive Learning
abstract
Single-cell classification based on single-omics data is often constrained by the one-sidedness of the data. With the advancement of single-cell sequencing technology, it has become possible to classify single cells using multi-omics data. However, integration and classification of multi-omics data are still challenging. In this study, we propose a model named MoSCHG. In this model, we first construct a heterogeneous bipartite graph based on the data of each omics, where the two types of nodes represent cells and their features (e.g., genes, chromatin) respectively, and the edge weights represent the relationship between cells and features; then GCN is applied with a residual mechanism to learn node embeddings in each bipartite graph and cell embeddings from different graphs are aligned based on supervised contrastive learning; finally, the aligned multi-omics cell embeddings are concatenated and the classification task is completed. Experimental results on three real datasets show that the proposed MoSCHG model outperforms the current state-of-the-art algorithms in classification performance, and through ablation studies, we validate the effectiveness of each module in MoSCHG.
Xinjian Chen 0007, Chenyang Xie, Xiaochen Bo, Zhongnan Zhang
BIBM5
2023 Drug-target and Drug-disease Association Prediction based on Drug-target-disease Network and Multi-task Learning
abstract
Traditional drug-target and drug-disease associations prediction tasks have been performed independently, without fully exploiting the relationships between drugs and various other entities, leading to inaccurate predictions. With the emergence of large-scale heterogeneous biological networks, multi-task learning can effectively enhance the accuracy of association prediction based on the associations between entities. In this study, we propose a multi-task learning framework named DTD-MTL to predict drug-target and drug-disease associations simultaneously. Firstly, it utilizes a multi-layer relational graph convolutional network (RGCN) to learn the features of each node in the drug-target-disease network. Subsequently, it obtains the initial feature of an edge by concatenating the features of the two nodes on the same edge. To coordinate different prediction tasks, drug features are shared among different tasks. Afterwards, the autoencoder (AE) is used to extract features from different types of edges. In order to make the learned edge features more suitable for different prediction tasks, the distance covariance (DC) is utilized to eliminate the specificity between different types of edges, thereby leveraging the relationships between different tasks more effectively. Finally, the drug-target and drug-disease associations predictions are achieved based on the edge features extracted by the AE. Experimental results on a widely-used dataset show that DTD-MTL outperforms the state-of-the-art methods in the prediction task of drug-target and drug-disease associations.
Binyu Wang, Hongyan Ye, Lianlian Wu, Xiaochen Bo, Zhongnan Zhang
BIBM6
2023 HSGCL-DTA: Hybrid-scale Graph Contrastive Learning based Drug-Target Binding Affinity Prediction
abstract
Drug-target binding affinity (DTA) is a critical criterion for drug screening. Accurate affinity prediction will significantly cut the cost of new drug development and accelerate the drug discovery process. However, most existing approaches frequently utilize sequence or structure information without incorporating any additional information. At the same time, they encode drugs and targets separately, ignoring the important existing drug-target relationships. In this study, we propose a novel DTA prediction approach, named HSGCL-DTA, which is based on hybrid-scale graph contrastive learning. To completely capture the global information and discriminative properties of the heterogeneous graphs, HSGCL-DTA divides the drug-target affinity graph into two subgraphs with stronger and weaker affinities respectively, and the node embeddings of the two subgraphs are obtained based on node-graph level contrastive learning. Afterwards, graph convolutional network (GCN) is used to encode the molecular graph of drugs and targets, and the node embeddings in the strong affinity subgraph are fused with molecule graph embeddings to fully utilize the distinct information in two different views. Another node-node level contrastive learning is performed between the affinity graph and molecular graphs, thereby filtering out task-independent noise that only appears in one graph. The final drug-target embeddings are put into a multilayer perceptron (MLP) for affinity prediction. Experiments on two widely-used datasets have shown that HSGCL-DTA achieves better prediction performance and generalization than the state-of-the-art DTA prediction methods.
Hongyan Ye, Yingying Song, Binyu Wang, Lianlian Wu, Xiaochen Bo, Zhongnan Zhang
ICTAI6
2023 GADRP: graph convolutional networks and autoencoders for cancer drug response prediction
abstract
Drug response prediction in cancer cell lines is of great significance in personalized medicine. In this study, we propose GADRP, a cancer drug response prediction model based on graph convolutional networks (GCNs) and autoencoders (AEs). We first use a stacked deep AE to extract low-dimensional representations from cell line features, and then construct a sparse drug cell line pair (DCP) network incorporating drug, cell line, and DCP similarity information. Later, initial residual and layer attention-based GCN (ILGCN) that can alleviate over-smoothing problem is utilized to learn DCP features. And finally, fully connected network is employed to make prediction. Benchmarking results demonstrate that GADRP can significantly improve prediction performance on all metrics compared with baselines on five datasets. Particularly, experiments of predictions of unknown DCP responses, drug-cancer tissue associations, and drug-pathway associations illustrate the predictive power of GADRP. All results highlight the effectiveness of GADRP in predicting drug responses, and its potential value in guiding anti-cancer drug selection.
Chong Dai, Yuqi Wen, Wenjuan Liu, Xiaochen Bo, Shaoliang Peng
Briefings Bioinform.7
2023 COMMO: a web server for the identification and analysis of consensus gene modules across multiple methods
abstract
SUMMARY: A variety of computational methods have been developed to identify functionally related gene modules from genome-wide gene expression profiles. Integrating the results of these methods to identify consensus modules is a promising approach to produce more accurate and robust results. In this application note, we introduce COMMO, the first web server to identify and analyze consensus gene functionally related gene modules from different module detection methods. First, COMMO implements eight state-of-the-art module detection methods and two consensus clustering algorithms. Second, COMMO provides users with mRNA and protein expression data for 33 cancer types from three public databases. Users can also upload their own data for module detection. Third, users can perform functional enrichment and two types of survival analyses on the observed gene modules. Finally, COMMO provides interactive, customizable visualizations and exportable results. With its extensive analysis and interactive capabilities, COMMO offers a user-friendly solution for conducting module-based precision medicine research. AVAILABILITY AND IMPLEMENTATION: COMMO web is available at https://commo.ncpsb.org.cn/, with the source code available on GitHub: https://github.com/Song-xinyu/COMMO/tree/master.
Mingfei Han 0001, Xinyu Song 0002, Xiaochen Bo
Bioinform.5
2023 EDST: a decision stump based ensemble algorithm for synergistic drug combination prediction
abstract
INTRODUCTION: There are countless possibilities for drug combinations, which makes it expensive and time-consuming to rely solely on clinical trials to determine the effects of each possible drug combination. In order to screen out the most effective drug combinations more quickly, scholars began to apply machine learning to drug combination prediction. However, most of them are of low interpretability. Consequently, even though they can sometimes produce high prediction accuracy, experts in the medical and biological fields can still not fully rely on their judgments because of the lack of knowledge about the decision-making process. RELATED WORK: Decision trees and their ensemble algorithms are considered to be suitable methods for pharmaceutical applications due to their excellent performance and good interpretability. We review existing decision trees or decision tree ensemble algorithms in the medical field and point out their shortcomings. METHOD: This study proposes a decision stump (DS)-based solution to extract interpretable knowledge from data sets. In this method, a set of DSs is first generated to selectively form a decision tree (DST). Different from the traditional decision tree, our algorithm not only enables a partial exchange of information between base classifiers by introducing a stump exchange method but also uses a modified Gini index to evaluate stump performance so that the generation of each node is evaluated by a global view to maintain high generalization ability. Furthermore, these trees are combined to construct an ensemble of DST (EDST). EXPERIMENT: The two-drug combination data sets are collected from two cell lines with three classes (additive, antagonistic and synergistic effects) to test our method. Experimental results show that both our DST and EDST perform better than other methods. Besides, the rules generated by our methods are more compact and more accurate than other rule-based algorithms. Finally, we also analyze the extracted knowledge by the model in the field of bioinformatics. CONCLUSION: The novel decision tree ensemble model can effectively predict the effect of drug combination datasets and easily obtain the decision-making process.
Lianlian Wu, Kunhong Liu 0001, Yong Xu 0009, Xiaochen Bo
BMC Bioinform.6
2023 Effective low-light image enhancement with multiscale and context learning network
Bin Jiang 0006, Xiaochen Bo, Chao Yang 0015
Multim. Tools Appl.3
2022 A Multi-View Learning-Based Rule Extraction Algorithm For Accurate Hepatotoxicity Prediction
abstract
Hepatotoxicity prediction is key to diseases with the high mortality rate. However, most of the algorithms used by now are black box in nature and lack of clear interpretability. This paper proposes a genetic algorithm-based interpretable algorithm based on rules extracted from a random forest. To take advantages from different types of omics data and molecular representations gathered from various datasets, our algorithm utilizes multiple distinct features to form a multi-view learning strategy. In detail, the genetic algorithm is designed to select optimal rules from each view, which are then used to form the ensemble of multi-view rule sets. The experiments are carried out to verify the performance of our algorithm on the hepatotoxicity data. The results confirm that our algorithm can gain high accuracy in most cases with more compact and shorter rules, compared with the original random forest or other rule-based algorithms. Our python source code and the related Supplementary Materials are available at: github.com/MLDMXM2017/MVR-GA.
Yuting Zhong, Bowei Yan, Kunhong Liu 0001, Yong Xu 0009, Xiaochen Bo
BIBM6
2022 An enhanced cascade-based deep forest model for drug combination prediction
abstract
Combination therapy has shown an obvious curative effect on complex diseases, whereas the search space of drug combinations is too large to be validated experimentally even with high-throughput screens. With the increase of the number of drugs, artificial intelligence techniques, especially machine learning methods, have become applicable for the discovery of synergistic drug combinations to significantly reduce the experimental workload. In this study, in order to predict novel synergistic drug combinations in various cancer cell lines, the cell line-specific drug-induced gene expression profile (GP) is added as a new feature type to capture the cellular response of drugs and reveal the biological mechanism of synergistic effect. Then, an enhanced cascade-based deep forest regressor (EC-DFR) is innovatively presented to apply the new small-scale drug combination dataset involving chemical, physical and biological (GP) properties of drugs and cells. Verified by the dataset, EC-DFR outperforms two state-of-the-art deep neural network-based methods and several advanced classical machine learning algorithms. Biological experimental validation performed subsequently on a set of previously untested drug combinations further confirms the performance of EC-DFR. What is more prominent is that EC-DFR can distinguish the most important features, making it more interpretable. By evaluating the contribution of each feature type, GP feature contributes 82.40%, showing the cellular responses of drugs may play crucial roles in synergism prediction. The analysis based on the top contributing genes in GP further demonstrates some potential relationships between the transcriptomic levels of key genes under drug regulation and the synergism of drug combinations.
Weiping Lin, Lianlian Wu, Yuqi Wen, Bowei Yan, Chong Dai, Kunhong Liu 0001, Xiaochen Bo
Briefings Bioinform.9
2022 DTI-HETA: prediction of drug-target interactions based on GCN and GAT on heterogeneous graph
abstract
Drug-target interaction (DTI) prediction plays an important role in drug repositioning, drug discovery and drug design. However, due to the large size of the chemical and genomic spaces and the complex interactions between drugs and targets, experimental identification of DTIs is costly and time-consuming. In recent years, the emerging graph neural network (GNN) has been applied to DTI prediction because DTIs can be represented effectively using graphs. However, some of these methods are only based on homogeneous graphs, and some consist of two decoupled steps that cannot be trained jointly. To further explore GNN-based DTI prediction by integrating heterogeneous graph information, this study regards DTI prediction as a link prediction problem and proposes an end-to-end model based on HETerogeneous graph with Attention mechanism (DTI-HETA). In this model, a heterogeneous graph is first constructed based on the drug-drug and target-target similarity matrices and the DTI matrix. Then, the graph convolutional neural network is utilized to obtain the embedded representation of the drugs and targets. To highlight the contribution of different neighborhood nodes to the central node in aggregating the graph convolution information, a graph attention mechanism is introduced into the node embedding process. Afterward, an inner product decoder is applied to predict DTIs. To evaluate the performance of DTI-HETA, experiments are conducted on two datasets. The experimental results show that our model is superior to the state-of-the-art methods. Also, the identification of novel DTIs indicates that DTI-HETA can serve as a powerful tool for integrating heterogeneous graph information to predict DTIs.
Kanghao Shao, Yuqi Wen, Zhongnan Zhang, Xiaochen Bo
Briefings Bioinform.6
2022 Computational methods, databases and tools for synthetic lethality prediction
abstract
Synthetic lethality (SL) occurs between two genes when the inactivation of either gene alone has no effect on cell survival but the inactivation of both genes results in cell death. SL-based therapy has become one of the most promising targeted cancer therapies in the last decade as PARP inhibitors achieve great success in the clinic. The key point to exploiting SL-based cancer therapy is the identification of robust SL pairs. Although many wet-lab-based methods have been developed to screen SL pairs, known SL pairs are less than 0.1% of all potential pairs due to large number of human gene combinations. Computational prediction methods complement wet-lab-based methods to effectively reduce the search space of SL pairs. In this paper, we review the recent applications of computational methods and commonly used databases for SL prediction. First, we introduce the concept of SL and its screening methods. Second, various SL-related data resources are summarized. Then, computational methods including statistical-based methods, network-based methods, classical machine learning methods and deep learning methods for SL prediction are summarized. In particular, we elaborate on the negative sampling methods applied in these models. Next, representative tools for SL prediction are introduced. Finally, the challenges and future work for SL prediction are discussed.
Junshan Han, Yanpeng Zhao, Caiyun Zhao, Bowei Yan, Chong Dai, Lianlian Wu, Yuqi Wen, Dongjin Leng, Zhongming Wang, Xiaoxi Yang, Xiaochen Bo
Briefings Bioinform.15
2022 Machine learning methods, databases and tools for drug combination prediction
abstract
Combination therapy has shown an obvious efficacy on complex diseases and can greatly reduce the development of drug resistance. However, even with high-throughput screens, experimental methods are insufficient to explore novel drug combinations. In order to reduce the search space of drug combinations, there is an urgent need to develop more efficient computational methods to predict novel drug combinations. In recent decades, more and more machine learning (ML) algorithms have been applied to improve the predictive performance. The object of this study is to introduce and discuss the recent applications of ML methods and the widely used databases in drug combination prediction. In this study, we first describe the concept and controversy of synergism between drug combinations. Then, we investigate various publicly available data resources and tools for prediction tasks. Next, ML methods including classic ML and deep learning methods applied in drug combination prediction are introduced. Finally, we summarize the challenges to ML methods in prediction tasks and provide a discussion on future work.
Lianlian Wu, Yuqi Wen, Dongjin Leng, Chong Dai, Zhongming Wang, Bowei Yan, Xiaochen Bo
Briefings Bioinform.12
2022 Systematic optimization of host-directed therapeutic targets and preclinical validation of repositioned antiviral drugs
abstract
Inhibition of host protein functions using established drugs produces a promising antiviral effect with excellent safety profiles, decreased incidence of resistant variants and favorable balance of costs and risks. Genomic methods have produced a large number of robust host factors, providing candidates for identification of antiviral drug targets. However, there is a lack of global perspectives and systematic prioritization of known virus-targeted host proteins (VTHPs) and drug targets. There is also a need for host-directed repositioned antivirals. Here, we integrated 6140 VTHPs and grouped viral infection modes from a new perspective of enriched pathways of VTHPs. Clarifying the superiority of nonessential membrane and hub VTHPs as potential ideal targets for repositioned antivirals, we proposed 543 candidate VTHPs. We then presented a large-scale drug-virus network (DVN) based on matching these VTHPs and drug targets. We predicted possible indications for 703 approved drugs against 35 viruses and explored their potential as broad-spectrum antivirals. In vitro and in vivo tests validated the efficacy of bosutinib, maraviroc and dextromethorphan against human herpesvirus 1 (HHV-1), hepatitis B virus (HBV) and influenza A virus (IAV). Their drug synergy with clinically used antivirals was evaluated and confirmed. The results proved that low-dose dextromethorphan is better than high-dose in both single and combined treatments. This study provides a comprehensive landscape and optimization strategy for druggable VTHPs, constructing an innovative and potent pipeline to discover novel antiviral host proteins and repositioned drugs, which may facilitate their delivery to clinical application in translational medicine to combat fatal and spreading viral infections.
Dafei Xie, Lianlian Wu, Pingkun Zhou, Xunlong Shi, Xiaochen Bo
Briefings Bioinform.10
2021 Drug-target interaction prediction based on nonnegative and self-representative matrix factorization
abstract
Drug-target interaction prediction is an important research field in computer-aided drug discovery. The data involved in drug-target interaction prediction are characterized by noise, high dimensionality, and sparseness, which leads to poor prediction performance of traditional machine learning methods. Matrix factorization methods are often used to predict unknown or missing data, and can deal with data with the above characteristics. Therefore, a drug-target interaction prediction model based on non-negative and self-representative matrix factorization is proposed in this study. The proposed model performs matrix factorization based on the topological structure of the drug-target interaction data, and focuses on capturing the internal structural information of the drug-target data for representation learning. At the same time, it introduces nonnegative and non-trivial solution constraints to optimize the representation learning results, and integrates the graphs regularization method to optimize the low-dimensional key latent factor matrix, and finally realizes the prediction of drug-target interactions. Experimental results show that the model effectively mines the structural information of drug-target interactions, and is superior to other benchmark methods in the prediction performance.
Yihua Ye, Zhongnan Zhang, Yuqi Wen, Xiaochen Bo
BIBM6
2021 A Metagraph-Based Model for Predicting Drug-Target Interaction on Heterogeneous Network
Peng Ke, Yuqi Wen, Zhongnan Zhang, Xiaochen Bo
ICANN (1)5
2021 Spatial density of open chromatin: an effective metric for the functional characterization of topologically associated domains
abstract
Topologically associated domains (TADs) are spatial and functional units of metazoan chromatin structure. Interpretation of the interplay between regulatory factors and chromatin structure within TADs is crucial to understand the spatial and temporal regulation of gene expression. However, a computational metric for the sensitive characterization of TAD regulatory landscape is lacking. Here, we present the spatial density of open chromatin (SDOC) metric as a quantitative measurement of intra-TAD chromatin state and structure. SDOC sensitively reflects epigenetic properties and gene transcriptional activity in TADs. During mouse T-cell development, we found that TADs with decreased SDOC are enriched in repressed developmental genes, and the joint effect of SDOC-decreasing and TAD clustering corresponds to the highest level of gene repression. In addition, we revealed a pervasive preference for TADs with similar SDOC to interact with each other, which may reflect the principle of chromatin organization.
Hao Li 0035, Hao Hong, Guifang Du, Xin Huang 0005, Yu Sun 0070, Junting Wang 0003, Hebing Chen, Xiaochen Bo
Briefings Bioinform.13
2021 Multi-dimensional data integration algorithm based on random walk with restart
abstract
BACKGROUND: The accumulation of various multi-omics data and computational approaches for data integration can accelerate the development of precision medicine. However, the algorithm development for multi-omics data integration remains a pressing challenge. RESULTS: Here, we propose a multi-omics data integration algorithm based on random walk with restart (RWR) on multiplex network. We call the resulting methodology Random Walk with Restart for multi-dimensional data Fusion (RWRF). RWRF uses similarity network of samples as the basis for integration. It constructs the similarity network for each data type and then connects corresponding samples of multiple similarity networks to create a multiplex sample network. By applying RWR on the multiplex network, RWRF uses stationary probability distribution to fuse similarity networks. We applied RWRF to The Cancer Genome Atlas (TCGA) data to identify subtypes in different cancer data sets. Three types of data (mRNA expression, DNA methylation, and microRNA expression data) are integrated and network clustering is conducted. Experiment results show that RWRF performs better than single data type analysis and previous integrative methods. CONCLUSIONS: RWRF provides powerful support to users to decipher the cancer molecular subtypes, thus may benefit precision treatment of specific patients in clinical practice.
Yuqi Wen, Xinyu Song 0002, Bowei Yan, Xiaoxi Yang, Lianlian Wu, Dongjin Leng, Xiaochen Bo
BMC Bioinform.8
2021 Synthetic Lethal Interactions Prediction Based on Multiple Similarity Measures Fusion
Lianlian Wu, Yuqi Wen, Xiaoxi Yang, Bowei Yan, Xiaochen Bo
J. Comput. Sci. Technol.6
2021 COMSUC: A web server for the identification of consensus molecular subtypes of cancer based on multiple methods and multi-omics data
abstract
Extensive amounts of multi-omics data and multiple cancer subtyping methods have been developed rapidly, and generate discrepant clustering results, which poses challenges for cancer molecular subtype research. Thus, the development of methods for the identification of cancer consensus molecular subtypes is essential. The lack of intuitive and easy-to-use analytical tools has posed a barrier. Here, we report on the development of the COnsensus Molecular SUbtype of Cancer (COMSUC) web server. With COMSUC, users can explore consensus molecular subtypes of more than 30 cancers based on eight clustering methods, five types of omics data from public reference datasets or users' private data, and three consensus clustering methods. The web server provides interactive and modifiable visualization, and publishable output of analysis results. Researchers can also exchange consensus subtype results with collaborators via project IDs. COMSUC is now publicly and freely available with no login requirement at http://comsuc.bioinforai.tech/ (IP address: http://59.110.25.27/). For a video summary of this web server, see S1 Video and S1 File.
Xinyu Song 0002, Xiaoxi Yang, Jijun Yu, Yuqi Wen, Lianlian Wu, Bowei Yan, Jiannan Feng, Xiaochen Bo
PLoS Comput. Biol.9
2021 NegStacking: Drug-Target Interaction Prediction Based on Ensemble Learning and Logistic Regression
abstract
Drug-target interactions (DTIs) identification is an important issue of drug research, and many methods proposed to predict potential DTIs based on machine learning treat it as a binary classification problem. However, the number of known interacting drug-target pairs (positive samples) is far less than that of non-interacting pairs (negative samples). Most methods do not utilize these large numbers of negative samples sufficiently, which limits their prediction performance. To address this problem, we proposed a stacking framework named NegStacking. First, it uses sampling to obtain multiple completely different negative sample sets. Then, each weak learner is trained with a different negative sample set and the same positive sample set, and the logistic regression (LR) is used as a meta-learner to adaptively combine these weak learners. Moreover, in the training process, feature subspacing and hyperparameter perturbation are applied to increase ensemble diversity. Finally, the trained model could be used to predict new samples. We compared NegStacking with other methods, and the experimental results show that our model is superior. NegStacking can improve the performance of predictive DTIs, and it has broad application prospects for improving the drug discovery process. The source code and datasets are available at https://github.com/Open-ss/NegStacking.
Zhongnan Zhang, Xiaochen Bo
IEEE ACM Trans. Comput. Biol. Bioinform.4
2020 DTIGCCN: Prediction of drug-target interactions based on GCN and CNN
abstract
Drug-target interaction (DTI) prediction plays an important role in drug repositioning, drug discovery, and drug design. In recent years, some DTI prediction methods based on machine learning have been proposed. They usually extract features from chemical genomics data. However, these methods are easy to extract redundant information that is not fully related with the prediction task and ignore the latent relationship between drug and target. This paper presents a new DTI prediction model named DTIGCCN. The model uses a spectral-based graph convolutional network (GCN) to extract features from drug and target expression profiles respectively, and a convolutional neural network (CNN) to extract latent associations between drug and target. Finally, the extracted features are concatenated together and fed into an effective classifier for prediction. The advantage of DTIGCCN is that the extracted features are more refined and targeted and the correlation between drug and target is fully applied to the prediction. Experimental results show that our model is superior to the conventional DTI prediction methods based on feature extraction and provides a new idea and method for DTI prediction.
Kanghao Shao, Zhongnan Zhang, Xiaochen Bo
ICTAI4
2020 New insights on human essential genes based on integrated analysis and the construction of the HEGIAP web-based platform
abstract
Essential genes are those whose loss of function compromises organism viability or results in profound loss of fitness. Recent gene-editing technologies have provided new opportunities to characterize essential genes. Here, we present an integrated analysis that comprehensively and systematically elucidates the genetic and regulatory characteristics of human essential genes. First, we found that essential genes act as 'hubs' in protein-protein interaction networks, chromatin structure and epigenetic modification. Second, essential genes represent conserved biological processes across species, although gene essentiality changes differently among species. Third, essential genes are important for cell development due to their discriminate transcription activity in embryo development and oncogenesis. In addition, we developed an interactive web server, the Human Essential Genes Interactive Analysis Platform (http://sysomics.com/HEGIAP/), which integrates abundant analytical tools to enable global, multidimensional interpretation of gene essentiality. Our study provides new insights that improve the understanding of human essential genes.
Hebing Chen, Ruijiang Li, Chenghui Zhao, Hao Hong, Xin Huang 0005, Hao Li 0035, Xiaochen Bo
Briefings Bioinform.10
2020 Domain-adversarial multi-task framework for novel therapeutic property prediction of compounds
abstract
MOTIVATION: With the rapid development of high-throughput technologies, parallel acquisition of large-scale drug-informatics data provides significant opportunities to improve pharmaceutical research and development. One important application is the purpose prediction of small-molecule compounds with the objective of specifying the therapeutic properties of extensive purpose-unknown compounds and repurposing the novel therapeutic properties of FDA-approved drugs. Such a problem is extremely challenging because compound attributes include heterogeneous data with various feature patterns, such as drug fingerprints, drug physicochemical properties and drug perturbation gene expressions. Moreover, there is a complex non-linear dependency among heterogeneous data. In this study, we propose a novel domain-adversarial multi-task framework for integrating shared knowledge from multiple domains. The framework first uses an adversarial strategy to learn target representations and then models non-linear dependency among several domains. RESULTS: Experiments on two real-world datasets illustrate that our approach achieves an obvious improvement over competitive baselines. The novel therapeutic properties of purpose-unknown compounds that we predicted have been widely reported or brought to clinics. Furthermore, our framework can integrate various attributes beyond the three domains examined herein and can be applied in industry for screening significant numbers of small-molecule drug candidates. AVAILABILITY AND IMPLEMENTATION: The source code and datasets are available at https://github.com/JohnnyY8/DAMT-Model. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lingwei Xie, Zhongnan Zhang, Kunhui Lin, Xiaochen Bo, Boyuan Feng, Kun Wan 0001, Yufei Ding 0001
Bioinform.5
2020 DeepHiC: A generative adversarial network for enhancing Hi-C data resolution
abstract
Hi-C is commonly used to study three-dimensional genome organization. However, due to the high sequencing cost and technical constraints, the resolution of most Hi-C datasets is coarse, resulting in a loss of information and biological interpretability. Here we develop DeepHiC, a generative adversarial network, to predict high-resolution Hi-C contact maps from low-coverage sequencing data. We demonstrated that DeepHiC is capable of reproducing high-resolution Hi-C data from as few as 1% downsampled reads. Empowered by adversarial training, our method can restore fine-grained details similar to those in high-resolution Hi-C matrices, boosting accuracy in chromatin loops identification and TADs detection, and outperforms the state-of-the-art methods in accuracy of prediction. Finally, application of DeepHiC to Hi-C data on mouse embryonic development can facilitate chromatin loop detection. We develop a web-based tool (DeepHiC, http://sysomics.com/deephic) that allows researchers to enhance their own Hi-C data with just a few clicks.
Hao Hong, Hao Li 0035, Guifang Du, Yu Sun 0070, Cheng Quan, Chenghui Zhao, Ruijiang Li, Xiaoyao Yin, Yangchen Huang, Hebing Chen, Xiaochen Bo
PLoS Comput. Biol.15
2019 A survey and evaluation of Web-based tools/databases for variant analysis of TCGA data
abstract
The Cancer Genome Atlas (TCGA) is a publicly funded project that aims to catalog and discover major cancer-causing genomic alterations with the goal of creating a comprehensive 'atlas' of cancer genomic profiles. The availability of this genome-wide information provides an unprecedented opportunity to expand our knowledge of tumourigenesis. Computational analytics and mining are frequently used as effective tools for exploring this byzantine series of biological and biomedical data. However, some of the more advanced computational tools are often difficult to understand or use, thereby limiting their application by scientists who do not have a strong computational background. Hence, it is of great importance to build user-friendly interfaces that allow both computational scientists and life scientists without a computational background to gain greater biological and medical insights. To that end, this survey was designed to systematically present available Web-based tools and facilitate the use TCGA data for cancer research.
Hao Li 0035, Ruijiang Li, Hebing Chen, Xiaochen Bo
Briefings Bioinform.7
2019 Stable H3K4me3 is associated with transcription initiation during early embryo development
abstract
MOTIVATION: During development of the mammalian embryo, histone modification H3K4me3 plays an important role in regulating gene expression and exhibits extensive reprograming on the parental genomes. In addition to these dramatic epigenetic changes, certain unchanging regulatory elements are also essential for embryonic development. RESULTS: Using large-scale H3K4me3 chromatin immunoprecipitation sequencing data, we identified a form of H3K4me3 that was present during all eight stages of the mouse embryo before implantation. This 'stable H3K4me3' was highly accessible and much longer than normal H3K4me3. Moreover, most of the stable H3K4me3 was in the promoter region and was enriched in higher chromatin architecture. Using in-depth analysis, we demonstrated that stable H3K4me3 was related to higher gene expression levels and transcriptional initiation during embryonic development. Furthermore, stable H3K4me3 was much more active in blood tumor cells than in normal blood cells, suggesting a potential mechanism of cancer progression. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xin Huang 0005, Ruijiang Li, Hao Hong, Chenghui Zhao, Pingkun Zhou, Hebing Chen, Xiaochen Bo, Hao Li 0035
Bioinform.10
2018 Prediction of DTIs for high-dimensional and class-imbalanced data based on CGAN
Zhongnan Zhang, Xiaochen Bo
BIBM4
2018 DTI-RCNN: New Efficient Hybrid Neural Network Model to Predict Drug-Target Interactions
Xiaoping Zheng, Xinyu Song 0002, Zhongnan Zhang, Xiaochen Bo
ICANN (1)5
2017 Drug - target interaction prediction with a deep-learning-based model
abstract
Drug-target interaction identification is of highly importance in drug research and development. The traditional experimental paradigm is costly, while the previous in silico prediction paradigm remains a challenge because of diversified data production platforms and data scarcity. In this paper, we modeled drug-target interaction prediction as a binary classification task based on transcriptome data of drug stimulation and gene knockout from LINCS project and developed a framework with a deep-learning-based model to predict potential interactions. The evaluation results showed that not only did our framework fit data with better accuracy than other classical methods, but predicted more credible drug-target interactions. What's more, the prediction has high percentage of overlap interactions across other platforms.
Lingwei Xie, Zhongnan Zhang, Xiaochen Bo, Xinyu Song 0002
BIBM4
2017 Exploring spatially adjacent TFBS-clustered regions with Hi-C data
abstract
MOTIVATION: Transcription factor binding sites (TFBSs) are clustered in the human genome, forming the TFBS-clustered regions that regulate gene transcription, which requires dynamic chromatin configurations between promoters and distal regulatory elements. Here, we propose a regulatory model called spatially adjacent TFBS-clustered regions (SATs), in which TFBS-clustered regions are connected by spatial proximity as identified by high-resolution Hi-C data. RESULTS: TFBS-clustered regions forming SATs appeared less frequently in gene promoters than did isolated TFBS-clustered regions, whereas SATs as a whole appeared more frequently. These observations indicate that multiple distal TFBS-clustered regions combined to form SATs to regulate genes. Further examination confirmed that a substantial portion of genes regulated by SATs were located between the paired TFBS-clustered regions instead of the downstream. We reconstructed the chromosomal conformation of the H1 human embryonic stem cell line using the ShRec3D algorithm and proposed the SAT regulatory model. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hebing Chen, Hao Li 0035, Xiaochen Bo
Bioinform.6
2017 BiRen: predicting enhancers with a deep-learning-based model using the DNA sequence alone
abstract
MOTIVATION: Enhancer elements are noncoding stretches of DNA that play key roles in controlling gene expression programmes. Despite major efforts to develop accurate enhancer prediction methods, identifying enhancer sequences continues to be a challenge in the annotation of mammalian genomes. One of the major issues is the lack of large, sufficiently comprehensive and experimentally validated enhancers for humans or other species. Thus, the development of computational methods based on limited experimentally validated enhancers and deciphering the transcriptional regulatory code encoded in the enhancer sequences is urgent. RESULTS: We present a deep-learning-based hybrid architecture, BiRen, which predicts enhancers using the DNA sequence alone. Our results demonstrate that BiRen can learn common enhancer patterns directly from the DNA sequence and exhibits superior accuracy, robustness and generalizability in enhancer prediction relative to other state-of-the-art enhancer predictors based on sequence characteristics. Our BiRen will enable researchers to acquire a deeper understanding of the regulatory code of enhancer sequences. AVAILABILITY AND IMPLEMENTATION: Our BiRen method can be freely accessed at https://github.com/wenjiegroup/BiRen . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bite Yang, Feng Liu 0054, Zhangyi Ouyang, Ziwei Xie, Xiaochen Bo, Wenjie Shu 0001
Bioinform.6
2017 NFPscanner: a webtool for knowledge-based deciphering of biomedical networks
abstract
BACKGROUND: Many biological pathways have been created to represent different types of knowledge, such as genetic interactions, metabolic reactions, and gene-regulating and physical-binding relationships. Biologists are using a wide range of omics data to elaborately construct various context-specific differential molecular networks. However, they cannot easily gain insight into unfamiliar gene networks with the tools that are currently available for pathways resource and network analysis. They would benefit from the development of a standardized tool to compare functions of multiple biological networks quantitatively and promptly. RESULTS: To address this challenge, we developed NFPscanner, a web server for deciphering gene networks with pathway associations. Adapted from a recently reported knowledge-based framework called network fingerprint, NFPscanner integrates the annotated pathways of 7 databases, 4 algorithms, and 2 graphical visualization modules into a webtool. It implements 3 types of network analysis: Fingerprint: Deciphering gene networks and highlighting inherent pathway modules Alignment: Discovering functional associations by finding optimized node mapping between 2 gene networks Enrichment: Calculating and visualizing gene ontology (GO) and pathway enrichment for genes in networks Users can upload gene networks to NFPscanner through the web interface and then interactively explore the networks' functions. CONCLUSIONS: NFPscanner is open-source software for non-commercial use, freely accessible at http://biotech.bmi.ac.cn/nfs .
Wenjian Xu, Ziwei Xie, Haochen He, Hao Hong, Xiaochen Bo, Fei Li 0004
BMC Bioinform.7
2016 De novo identification of replication-timing domains in the human genome by deep learning
abstract
MOTIVATION: The de novo identification of the initiation and termination zones-regions that replicate earlier or later than their upstream and downstream neighbours, respectively-remains a key challenge in DNA replication. RESULTS: Building on advances in deep learning, we developed a novel hybrid architecture combining a pre-trained, deep neural network and a hidden Markov model (DNN-HMM) for the de novo identification of replication domains using replication timing profiles. Our results demonstrate that DNN-HMM can significantly outperform strong, discriminatively trained Gaussian mixture model-HMM (GMM-HMM) systems and other six reported methods that can be applied to this challenge. We applied our trained DNN-HMM to identify distinct replication domain types, namely the early replication domain (ERD), the down transition zone (DTZ), the late replication domain (LRD) and the up transition zone (UTZ), using newly replicated DNA sequencing (Repli-Seq) data across 15 human cells. A subsequent integrative analysis revealed that these replication domains harbour unique genomic and epigenetic patterns, transcriptional activity and higher-order chromosomal structure. Our findings support the 'replication-domain' model, which states (1) that ERDs and LRDs, connected by UTZs and DTZs, are spatially compartmentalized structural and functional units of higher-order chromosomal structure, (2) that the adjacent DTZ-UTZ pairs form chromatin loops and (3) that intra-interactions within ERDs and LRDs tend to be short-range and long-range, respectively. Our model reveals an important chromatin organizational principle of the human genome and represents a critical step towards understanding the mechanisms regulating replication timing. AVAILABILITY AND IMPLEMENTATION: Our DNN-HMM method and three additional algorithms can be freely accessed at https://github.com/wenjiegroup/DNN-HMM The replication domain regions identified in this study are available in GEO under the accession ID GSE53984. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Feng Liu 0054, Hao Li 0035, Pingkun Zhou, Xiaochen Bo, Wenjie Shu 0001
Bioinform.5
2016 RevEcoR: an R package for the reverse ecology analysis of microbiomes
abstract
BACKGROUND: All species live in complex ecosystems. The structure and complexity of a microbial community reflects not only diversity and function, but also the environment in which it occurs. However, traditional ecological methods can only be applied on a small scale and for relatively well-understood biological systems. Recently, a graph-theory-based algorithm called the reverse ecology approach has been developed that can analyze the metabolic networks of all the species in a microbial community, and predict the metabolic interface between species and their environment. RESULTS: Here, we present RevEcoR, an R package and a Shiny Web application that implements the reverse ecology algorithm for determining microbe-microbe interactions in microbial communities. This software allows users to obtain large-scale ecological insights into species' ecology directly from high-throughput metagenomic data. The software has great potential for facilitating the study of microbiomes. CONCLUSIONS: RevEcoR is open source software for the study of microbial community ecology. The RevEcoR R package is freely available under the GNU General Public License v. 2.0 at http://cran.r-project.org/web/packages/RevEcoR/ with the vignette and typical usage examples, and the interactive Shiny web application is available at http://yiluheihei.shinyapps.io/shiny-RevEcoR , or can be installed locally with the source code accessed from https://github.com/yiluheihei/shiny-RevEcoR .
Xiaofei Zheng, Fei Li 0004, Xiaochen Bo
BMC Bioinform.5
2015 New insights into the landscape relationships of host response to bacterial pathogens
abstract
Modern understanding of microbiology largely lays foundation in the biological characterization of microorganisms. However, the landscape relationships of host transcriptional response (HTR) to different bacterial pathogens have not yet been systematically explored. Here, we established the first generation of HTR network (HTRN) according to the HTR similarities among 21 different human pathogenic bacterial species by integrating 258 pairs of host cellular gene expression profiles upon infections. Further, the network was dissected into five bacterial communities of more consensus internal HTR. Interestingly, analysis of signature genes across different communities revealed that distinct community signatures (CS) present differential gene expression patterns. Functional annotation suggested a common feature of host cell response to bacterial infections that specific functional gene clusters (BPs and/or signaling pathways) were preferentially elicited or subverted by community bacterial pathogens. Notably, community signatures (especially key associators participating dissimilar functional profiles) were highly enriched of GWAS disease-related genes, which associated bacterial infections with common and specific non-infectious human disease(s). About 40% of the associations were confirmed by literature investigation that further indicated possible/potential association directionality. Our characterization and analysis were the first to feature differential community HTRs upon bacterial pathogen infections and suggested new perspective of understanding infection-disease associations and underlying pathogenesis.
Xiaoyao Yin, Xiaochen Bo, Cong Niu, Naiyang Guan, Zhigang Luo
IJCNN4
2014 ExpTreeDB: Web-based query and visualization of manually annotated gene expression profiling experiments of human and mouse from GEO
abstract
MOTIVATION: Numerous public microarray datasets are valuable resources for the scientific communities. Several online tools have made great steps to use these data by querying related datasets with users' own gene signatures or expression profiles. However, dataset annotation and result exhibition still need to be improved. RESULTS: ExpTreeDB is a database that allows for queries on human and mouse microarray experiments from Gene Expression Omnibus with gene signatures or profiles. Compared with similar applications, ExpTreeDB pays more attention to dataset annotations and result visualization. We introduced a multiple-level annotation system to depict and organize original experiments. For example, a tamoxifen-treated cell line experiment is hierarchically annotated as 'agent→drug→estrogen receptor antagonist→tamoxifen'. Consequently, retrieved results are exhibited by an interactive tree-structured graphics, which provide an overview for related experiments and might enlighten users on key items of interest. AVAILABILITY AND IMPLEMENTATION: The database is freely available at http://biotech.bmi.ac.cn/ExpTreeDB. Web site is implemented in Perl, PHP, R, MySQL and Apache.
Fuqiang Ye, Juanjuan Zhu, Bite Yang, Yongge Wu, Fei Li 0004, Shengqi Wang, Xiaochen Bo
Bioinform.12
2013 Exploring the role of human miRNAs in virus-host interactions using systematic overlap analysis
abstract
MOTIVATION: Human miRNAs have recently been found to have important roles in viral replication. Understanding the patterns and details of human miRNA interactions during virus-host interactions may help uncover novel antiviral therapies. Based on the abundance of knowledge available regarding protein-protein interactions (PPI), virus-host protein interactions, experimentally validated human miRNA-target pairs and transcriptional regulation of human miRNAs, it is possible to explore the complex regulatory network that exists between viral proteins and human miRNAs at the system level. RESULTS: By integrating current data regarding the virus-human interactome and human miRNA-target pairs, the overlap between targets of viral proteins and human miRNAs was identified and found to represent topologically important proteins (e.g. hubs or bottlenecks) at the global center of the human PPI network. Viral proteins and human miRNAs were also found to significantly target human PPI pairs. Furthermore, an overlap analysis of virus targets and transcription factors (TFs) of human miRNAs revealed that viral proteins preferentially target human miRNA TFs, representing a new pattern of virus-host interactions. Potential feedback loops formed by viruses, human miRNAs and miRNA TFs were also identified, and these may be exploited by viruses resulting in greater virulence and more effective replication strategies.
Xiuliang Cui, Fei Li 0004, Shengqi Wang, Xiaochen Bo
Bioinform.7
2010 PerturbationAnalyzer: a tool for investigating the effects of concentration perturbation on protein interaction networks
abstract
UNLABELLED: The propagation of perturbations in protein concentration through a protein interaction network (PIN) can shed light on network dynamics and function. In order to facilitate this type of study, PerturbationAnalyzer, which is an open source plugin for Cytoscape, has been developed. PerturbationAnalyzer can be used in manual mode for simulating user-defined perturbations, as well as in batch mode for evaluating network robustness and identifying significant proteins that cause large propagation effects in the PINs when their concentrations are perturbed. Results from PerturbationAnalyzer can be represented in an intuitive and customizable way and can also be exported for further exploration. PerturbationAnalyzer has great potential in mining the design principles of protein networks, and may be a useful tool for identifying drug targets. AVAILABILITY: PerturbationAnalyzer can be accessed from the Cytoscape web site http://www.cytoscape.org/plugins/index.php or http://biotech.bmi.ac.cn/PerturbationAnalyzer. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fei Li 0004, Wenjian Xu, Yuxing Peng 0001, Xiaochen Bo, Shengqi Wang
Bioinform.5
2010 GOSemSim: an R package for measuring semantic similarity among GO terms and gene products
abstract
SUMMARY: The semantic comparisons of Gene Ontology (GO) annotations provide quantitative ways to compute similarities between genes and gene groups, and have became important basis for many bioinformatics analysis approaches. GOSemSim is an R package for semantic similarity computation among GO terms, sets of GO terms, gene products and gene clusters. Four information content (IC)- and a graph-based methods are implemented in the GOSemSim package, multiple species including human, rat, mouse, fly and yeast are also supported. The functions provided by the GOSemSim offer flexibility for applications, and can be easily integrated into high-throughput analysis pipelines. AVAILABILITY: GOSemSim is released under the GNU General Public License within Bioconductor project, and freely available at http://bioconductor.org/packages/2.6/bioc/html/GOSemSim.html.
Guangchuang Yu, Fei Li 0004, Yide Qin, Xiaochen Bo, Shengqi Wang
Bioinform.4
2009 EvoRSR: an integrated system for exploring evolution of RNA structural robustness
abstract
BACKGROUND: Robustness, maintaining a constant phenotype despite perturbations, is a fundamental property of biological systems that is incorporated at various levels of biological complexity. Although robustness has been frequently observed in nature, its evolutionary origin remains unknown. Current hypotheses suggest that robustness originated as a direct consequence of natural selection, as an intrinsic property of adaptations, or as a congruent correlate of environment robustness. To elucidate the evolutionary origins of robustness, a convenient computational package is strongly needed. RESULTS: In this study, we developed the open-source integrated system EvoRSR (Evolution of RNA Structural Robustness) to explore the evolution of robustness based on biologically important landscapes induced by RNA folding. EvoRSR is object-oriented, modular, and freely available at http://biotech.bmi.ac.cn/EvoRSR under the GNU/GPL license. We present an overview of EvoRSR package and illustrate its features with the miRNA gene cel-mir-357. CONCLUSION: EvoRSR is a novel and flexible package for exploring the evolution of robustness. Accordingly, EvoRSR can be used for future studies to investigate the evolution and origin of robustness and to address other common questions about robustness. While the current EvoRSR environment is a versatile analysis framework, future versions can include features to enhance evolutionary studies of robustness.
Wenjie Shu 0001, Xiaochen Bo, Zhiqiang Zheng 0005, Shengqi Wang
BMC Bioinform.3
2008 A novel representation of RNA secondary structure based on element-contact graphs
abstract
BACKGROUND: Depending on their specific structures, noncoding RNAs (ncRNAs) play important roles in many biological processes. Interest in developing new topological indices based on RNA graphs has been revived in recent years, as such indices can be used to compare, identify and classify RNAs. Although the topological indices presented before characterize the main topological features of RNA secondary structures, information on RNA structural details is ignored to some degree. Therefore, it is necessity to identify topological features with low degeneracy based on complete and fine-grained RNA graphical representations. RESULTS: In this study, we present a complete and fine scheme for RNA graph representation as a new basis for constructing RNA topological indices. We propose a combination of three vertex-weighted element-contact graphs (ECGs) to describe the RNA element details and their adjacent patterns in RNA secondary structure. Both the stem and loop topologies are encoded completely in the ECGs. The relationship among the three typical topological index families defined by their ECGs and RNA secondary structures was investigated from a dataset of 6,305 ncRNAs. The applicability of topological indices is illustrated by three application case studies. Based on the applied small dataset, we find that the topological indices can distinguish true pre-miRNAs from pseudo pre-miRNAs with about 96% accuracy, and can cluster known types of ncRNAs with about 98% accuracy, respectively. CONCLUSION: The results indicate that the topological indices can characterize the details of RNA structures and may have a potential role in identifying and classifying ncRNAs. Moreover, these indices may lead to a new approach for discovering novel ncRNAs. However, further research is needed to fully resolve the challenging problem of predicting and classifying noncoding RNAs.
Wenjie Shu 0001, Xiaochen Bo, Zhiqiang Zheng 0005, Shengqi Wang
BMC Bioinform.2
2006 Selection of antisense oligonucleotides based on multiple predicted target mRNA structures
abstract
BACKGROUND: Local structures of target mRNAs play a significant role in determining the efficacies of antisense oligonucleotides (ODNs), but some structure-based target site selection methods are limited by uncertainties in RNA secondary structure prediction. If all the predicted structures of a given mRNA within a certain energy limit could be used simultaneously, target site selection would obviously be improved in both reliability and efficiency. In this study, some key problems in ODN target selection on the basis of multiple predicted target mRNA structures are systematically discussed. RESULTS: Two methods were considered for merging topologically different RNA structures into integrated representations. Several parameters were derived to characterize local target site structures. Statistical analysis on a dataset with 448 ODNs against 28 different mRNAs revealed 9 features quantitatively associated with efficacy. Features of structural consistency seemed to be more highly correlated with efficacy than indices of the proportion of bases in single-stranded or double-stranded regions. The local structures of the target site 5' and 3' termini were also shown to be important in target selection. Neural network efficacy predictors using these features, defined on integrated structures as inputs, performed well in "minus-one-gene" cross-validation experiments. CONCLUSION: Topologically different target mRNA structures can be merged into integrated representations and then used in computer-aided ODN design. The results of this paper imply that some features characterizing multiple predicted target site structures can be used to predict ODN efficacy.
Xiaochen Bo, Shaoke Lou, Daochun Sun, Wenjie Shu 0001, Shengqi Wang
BMC Bioinform.1
2006 RDMAS: a web server for RNA deleterious mutation analysis
abstract
BACKGROUND: The diverse functions of ncRNAs critically depend on their structures. Mutations in ncRNAs disrupting the structures of functional sites are expected to be deleterious. RNA deleterious mutations have attracted wide attentions because some of them in cells result in serious disease, and some others in microbes influence their fitness. RESULTS: The RDMAS web server we describe here is an online tool for evaluating structural deleteriousness of single nucleotide mutation in RNA genes. Several structure comparison methods have been integrated; sub-optimal structures predicted can be optionally involved to mitigate the uncertainty of secondary structure prediction. With a user-friendly interface, the web application is easy to use. Intuitive illustrations are provided along with the original computational results to facilitate quick analysis. CONCLUSION: RDMAS can be used to explore the structure alterations which cause mutations pathogenic, and to predict deleterious mutations which may help to determine the functionally critical regions. RDMAS is freely accessed via http://biosrv1.bmi.ac.cn/rdmas.
Wenjie Shu 0001, Xiaochen Bo, Rujia Liu, Zhiqiang Zheng 0005, Shengqi Wang
BMC Bioinform.2
2005 TargetFinder: a software for antisense oligonucleotide target site selection based on MAST and secondary structures of target mRNA
abstract
UNLABELLED: TargetFinder is a PC/Windows program for interactive effective antisense oligonucleotide (AO) selection based on mRNA accessible site tagging (MAST) and secondary structures of target mRNA. To make MAST result intuitive, both the alignment result and tag frequency profile is illustrated. As theoretical reference, secondary structure and single strand probability profile of target mRNA is also represented. All of these sequences and profiles are displayed in aligned mode, which facilitates identification of the accessible sites in target mRNA. Graphical, user-friendly interface makes TargetFinder a useful tool in AO target site selection. AVAILABILITY: The software is freely available at http://www.bioit.org.cn/ao/targetfinder.htm CONTACT: [email protected].
Xiaochen Bo, Shengqi Wang
Bioinform.1
2001 Evaluation of the Image Degradation for a Typical Watermarking Algorithm in the Block-DCT Domain
Xiaochen Bo, Lincheng Shen, Wensen Chang
ICICS1