Hong-Bin Shen

dblp:34/4415 · also Hongbin Shen · DBLP profile ↗
← Back
93ranked-venue papers
3as first author
38since 2021 · last 2026
0000-0002-4029-3325ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 68 · 1 first-author · 33 since 2021Artificial intelligence and machine learning · 25 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Identifying Nuclear Export Signals from Large-Scale Candidate Pools with a Deep Reranking Method
Zhixuan Piao, Xiaoyong Pan, Hong-Bin Shen
ISBRA (1)3
2026 Reconstructing cell-cell interaction network in single-cell spatial transcriptomics via directed heterogeneous graph autoencoder
abstract
MOTIVATION: Spatial transcriptome data have both gene expression information and cell spatial location information, offering exceptional prospects for analyzing cell-cell interaction (CCI) network. Most existing statistical and optimal transport-based methods rely only on known ligand-receptor pairs to infer CCI network. Furthermore, most current deep learning frameworks rely on symmetric decoders or undirected graph architectures. RESULTS: Taking advantage of spatial transcriptomic data and graph autoencoders, we present a directed heterogeneous graph autoencoder-based approach DualCellChat to reconstruct a complete and accurate CCI network from incomplete single cell spatial transcriptomics. Benchmarked on five single-cell spatial datasets from four different technologies, we demonstrate that DualCellChat outperforms existing deep learning-based methods and can inherently model the direction of cellular interactions. Furthermore, we introduce downstream analysis to infer signature genes involved in cellular interactions from the reconstructed CCI network and infer significant ligand-receptor pairs for specific cell types. AVAILABILITY AND IMPLEMENTATION: The dataset and code are available in GitHub (https://github.com/JinxianHu/DualCellChat) and Zenodo (DOI: 10.5281/zenodo.18512678).
Jin-Xian Hu, Xiaoyong Pan, Hong-Bin Shen
Bioinform.4
2026 DrugDL: dual-modal deep learning framework for multi-property drug prediction and targeted therapy discovery
abstract
MOTIVATION: The accurate and robust representation of drug molecule features, the prediction of drug-target biomacromolecule interactions, and the determination of physicochemical properties are crucial in drug development. However, these tasks remain challenging due to issues such as the limited generalizability of single-modal representations, the absence of multitask prediction frameworks, and weak adaptability in cold-start scenarios. RESULTS: In this study, we present DrugDL, a framework for comprehensive drug molecule representation and the prediction of multiple downstream tasks, including drug-target interactions, binding affinities, binding sites, physicochemical properties, toxicity, and drug-drug interactions. DrugDL jointly learns representations of the drug chemical space and the target protein biological space, while capturing multiscale interaction mechanisms between drug molecules and target proteins through the integration of cross-modal contrastive learning and single-modal feature enhancement algorithms. Specifically, DrugDL employs a multitask prediction framework to predict multiple properties of drug molecules. In practical applications, it consistently outperforms state-of-the-art methods, particularly in cold-start tasks. The framework has been successfully applied to high-throughput screening, the identification of inhibitors of SARS-CoV-2 and metabolic enzymes, and the prediction of cancer-targeted drugs. Experimental validations on EGFR and ALK targets further demonstrate its effectiveness as a precise drug discovery tool. By enabling accurate molecular representation and multi-property prediction, DrugDL provides end-to-end technical support for drug development, thereby significantly accelerating the drug discovery process. AVAILABILITY AND IMPLEMENTATION: The datasets and code are available at https://github.com/ZhangQi9910/DrugDL. The version of record is archived in Zenodo with the DOI: 10.5281/zenodo.20579718.
Yuxiao Wei, Yunpeng Xia, Long-Chen Shen, Hong-Bin Shen, Dongjun Yu
Bioinform.7
2026 SMENET: A Multi-View Semantic Model for Multi-Level Enzyme Function Prediction
abstract
Comprehending biological reproduction and cellular metabolism is facilitated by the Enzyme Commission, which matches protein sequences to the biochemical reactions they catalyse through EC numbers. In recent years, several methods have been proposed for predicting enzyme function. However, these methods still encounter challenges. Firstly, traditional methods for manually designing enzyme features are complex and cumbersome, lacking an effective generalized method for embedding enzyme sequences. Secondly, the distribution gap between different enzymes is significant, which resulting in existing methods struggling to predict multilevel enzyme functions. Thirdly, traditional enzyme function prediction models only extract single view feature of enzyme, so there is still room for further improving the ability of these models to extract enzyme data. To address these challenges, a new multilevel enzyme function prediction model (SMENET) based on multi-view semantics is proposed. This method uses protein large language model to extract semantic information. Subsequently, this semantic information is fed into multiple information extraction network modules, followed by using Biologic Sematic Attention to integrate these views' information. Finally, a multi-view adaptive fusion network is designed to extract the best common representation between multiple semantic views. Extensive experiments were conducted on multiple datasets to validate the effectiveness of SMENET.
Hanwen Zhou, Wei Zhang 0221, Zhaohong Deng, Guanjin Wang, Zhisheng Wei, Xiaoyong Pan, Hong-Bin Shen, Dongjun Yu, Jing Wu 0030
IEEE Trans. Comput. Biol. Bioinform.8
2025 Mutual exclusive gene expression reveals a stress-induced compensatory role of taurine uptake in dilated cardiomyopathy
abstract
Abstract Mutually exclusive gene expression, where gene pairs are expressed in strict alternation within individual cells, reflects fundamental inter-gene regulatory mechanisms and can reveal shifts in transcriptional programs during development or disease. Detecting such patterns is critical for resolving rare cellular subpopulations, temporally discrete states along pseudotime, and spatially segregated neighborhoods in single-cell and spatial multi-omics data. However, the sparsity and dropout inherent to single-cell data make mutually exclusive expression difficult to detect, leading conventional feature selection methods to overlook subtle yet functionally important genes. We present MULE, an unbiased framework that systematically organizes collective mutual exclusivity into a hierarchical taxonomy. Applying MULE to cardiac datasets, we uncovered robust upregulation of SPOCK1 and SLC6A6 in dilated cardiomyopathy, previously obscured by the inability to resolve pathological cardiomyocytes. In vivo and in vitro experiments demonstrated that stress-induced SLC6A6 upregulation serves a cardiomyocyte self-protective mechanism. Taurine supplementation reduced oxidative stress, restored calcium homeostasis, prevented cell death, and improved cardiac function post-injury. These findings elucidate a novel cardiomyocyte stress response and highlight the therapeutic promise of taurine supplementation for the treatment of dilated cardiomyopathy.
Jinpu Cai, Luqi Yang, Luting Zhou, Ziqi Rong, Linkang He, Xinzhu Jiang, Yu Zhao 0009, Jianhua Yao 0001, Hong-Bin Shen, Shyam Prabhakar, Qiuyu Lian, Hongyi Xin
Briefings Bioinform.16
2025 DRAG: design RNAs as hierarchical graphs with reinforcement learning
abstract
The rapid development of RNA vaccines and therapeutics puts forward intensive requirements on the sequence design of RNAs. RNA sequence design, or RNA inverse folding, aims to generate RNA sequences that can fold into specific target structures. To date, efficient and high-accuracy prediction models for secondary structures of RNAs have been developed. They provide a basis for computational RNA sequence design methods. Especially, reinforcement learning (RL) has emerged as a promising approach for RNA design due to its ability to learn from trial and error in generation tasks and work without ground truth data. However, existing RL methods are limited in considering complex hierarchical structures in RNA design environments. To address the above limitation, we propose DRAG, an RL method that builds design environments for target secondary structures with hierarchical division based on graph neural networks. Through extensive experiments on benchmark datasets, DRAG exhibits remarkable performance compared with current machine-learning approaches for RNA sequence design. This advantage is particularly evident in long and intricate tasks involving structures with significant depth.
Yichong Li, Xiaoyong Pan, Hong-Bin Shen, Yang Yang 0030
Briefings Bioinform.3
2025 Predicting transcriptional changes induced by molecules with MiTCP
abstract
Studying the changes in cellular transcriptional profiles induced by small molecules can significantly advance our understanding of cellular state alterations and response mechanisms under chemical perturbations, which plays a crucial role in drug discovery and screening processes. Considering that experimental measurements need substantial time and cost, we developed a deep learning-based method called Molecule-induced Transcriptional Change Predictor (MiTCP) to predict changes in transcriptional profiles (CTPs) of 978 landmark genes induced by molecules. MiTCP utilizes graph neural network-based approaches to simultaneously model molecular structure representation and gene co-expression relationships, and integrates them for CTP prediction. After training on the L1000 dataset, MiTCP achieves an average Pearson correlation coefficient (PCC) of 0.482 on the test set and an average PCC of 0.801 for predicting the top 50 differentially expressed genes, which outperforms other existing methods. Furthermore, we used MiTCP to predict CTPs of three cancer drugs, palbociclib, irinotecan and goserelin, and performed gene enrichment analysis on the top differentially expressed genes and found that the enriched pathways and Gene Ontology terms are highly relevant to the corresponding diseases, which reveals the potential of MiTCP in drug development.
Kaiyuan Yang 0009, Jiabei Cheng, Shenghao Cao, Xiaoyong Pan, Hong-Bin Shen
Briefings Bioinform.5
2025 Integration of multi-source gene interaction networks and omics data with graph attention networks to identify novel disease genes
abstract
MOTIVATION: The pathogenesis of diseases is closely associated with genes, and the discovery of disease genes holds significant importance for understanding disease mechanisms and designing targeted therapeutics. However, biological validation of all genes for diseases is expensive and challenging. RESULTS: In this study, we propose DGP-AMIO, a computational method based on graph attention networks, to rank all unknown genes and identify potential novel disease genes by integrating multi-omics and gene interaction networks from multiple data sources. DGP-AMIO outperforms other methods significantly on 20 disease datasets, with an average AUROC and AUPR exceeding 0.9. The superior performance of DGP-AMIO is attributed to the integration of multiomics and gene interaction networks from multiple databases, as well as triGAT, a proposed GAT-based method that enables precise identification of disease genes in directed gene networks. Enrichment analysis conducted on the top 100 genes predicted by DGP-AMIO and literature research revealed that a majority of enriched GO terms, KEGG pathways and top genes were associated with diseases supported by relevant studies. We believe that our method can serve as an effective tool for identifying disease genes and guiding subsequent experimental validation efforts. AVAILABILITY AND IMPLEMENTATION: DGP-AMIO is publicly available at https://github.com/yangkaiyuan1027/DGP-AMIO.
Kaiyuan Yang 0009, Jiabei Cheng, Shenghao Cao, Xiaoyong Pan, Hong-Bin Shen
Bioinform.5
2025 m2ST: dual multi-scale graph clustering for spatially resolved transcriptomics
abstract
MOTIVATION: Spatial clustering is a key analytical technique for exploring spatial transcriptomics data. Recent graph neural network-based methods have shown promise in spatial clustering but face notable challenges. One significant issue is that analyzing the functions and complex mechanisms of organisms from a single scale is difficult and most methods focus exclusively on the single-scale representation of transcriptomic data, potentially limiting the discriminative power of extracted features for spatial domain clustering. Furthermore, classical clustering algorithms are often applied directly to latent representation, making it a worthwhile endeavor to explore a tailored clustering method to further improve the accuracy of spatial domain annotation. RESULTS: To address these limitations, we propose m2ST, a novel dual multi-scale graph clustering method. m2ST first uses a multi-scale masked graph autoencoder to extract representations across different scales from spatial transcriptomic data. To effectively compress and distill meaningful knowledge embedded in the data, m2ST introduces a random masking mechanism for node features and uses a scaled cosine error as the loss function. Additionally, we introduce a tailored multi-scale clustering framework that integrates scale-common and scale-specific information exploration into the clustering process, achieving more robust annotation performance. Shannon entropy is finally utilized to dynamically adjust the importance of different scales. Extensive experiments on multiple spatial transcriptomic datasets demonstrate the superior performance of m2ST compared to existing methods. AVAILABILITY AND IMPLEMENTATION: https://github.com/BBKing49/m2ST.
Wei Zhang 0221, Hailong Yang 0001, Te Zhang, Zhaohong Deng, Xiaoyong Pan, Hong-Bin Shen, Dongjun Yu, Shitong Wang 0001
Bioinform.9
2025 CATransUnetLBP: Accurate Prediction of Protein-Ligand Binding Pockets Using a Hybrid Network
abstract
The development of intelligent methods capable of predicting protein-ligand binding sites has become a popular research field. Recently, deep learning based methods have been proposed as a promising solution for this task. However, some limitations still exist. For example, the network structure is not optimized for predicting protein binding pockets, which limits the model's capabilities. To address the aforementioned challenges, a novel method called CATransUnetLPB is proposed, in which a new network structure named CATransUnet is designed. The proposed CATransUnet combines CNN and Transformer models to accurately segment binding pocket regions from protein 3D structures. It outperforms existing representative methods on three test sets, demonstrating the effectiveness of optimizing the deep network model for detecting protein ligand binding pockets. Furthermore, we conduct thorough analysis on applying data augmentation to protein data structure and confirm that such technique can enhance the model's generalization ability, thereby ensuring good performance on new protein structures. Moreover, experiments show that the predicted binding pockets from our model can complement the results obtained from other methods. This suggests that integrating our method with existing approaches could further improve the prediction of protein-ligand binding pockets.
Cheng Cai, Zhaohong Deng, Andong Li, Yun Zuo 0001, Haoran Chen 0003, Zhisheng Wei, Xiaoyong Pan, Hong-Bin Shen, Dongjun Yu
IEEE Trans. Comput. Biol. Bioinform.9
2025 DMMAFS: Protein Function Prediction Based on Multi-Modal Multi-Attention Fusion Features
abstract
Intelligent prediction of protein function is more efficient and less resource-consuming and has achieved significant progress in recent years. However, most of the current methods are performed solely based on the sequence information of proteins. These methods overlook information of other modalities that the proteins themselves possess, which makes it difficult to achieve the desired predicted results. Furthermore, a few existing methods based on multiple modal information fuse them in a simple splicing manner and fail to fully exploit the complementary relation between different modalities. To address the above-mentioned challenges, we propose Multi-modal Multi-attention fusion Features (DMMAFS), a method based on deep learning, to predict protein function. On the one hand, DMMAFS gains the semantic information embedded in the sequence itself through the self-attention learning of the sequence. On the other hand, DMMAFS employs the 3D structural information of proteins to compensate for the sequence information. Particularly, a S-C cross-modal cross-attention fusion network module is proposed that not only optimizes the weights of the semantic information but also efficiently fuses the sequence features with the structural information, thus avoiding the simple splicing of different modal features. Our experimental results demonstrate that the proposed DMMAFS outperforms the state-of-the-art methods in protein function prediction.
Liangwen He, Zhaohong Deng, Fuping Hu, Yun Zuo 0001, Haoran Chen 0003, Xiaoyong Pan, Zhisheng Wei, Hong-Bin Shen, Dongjun Yu, Jing Wu 0030
IEEE Trans. Comput. Biol. Bioinform.11
2024 GenoHoption: Bridging Gene Network Graphs and Single-Cell Foundation Models
abstract
The remarkable success of foundation models has sparked growing interest in their application to single-cell biology. Models like Geneformer and scGPT promise to serve as versatile tools in this specialized field. However, representing a cell as a sequence of genes remains an open question since the order of genes is interchangeable. Injecting the gene network graph offers gene relative positions and compact data representation but poses a dilemma: limited receptive fields without in-layer message passing or parameter explosion with message passing in each layer. To pave the way forward, we propose GenoHoption, a new computational framework for single-cell sequencing data that effortlessly combines the strengths of these foundation models with explicit relationships in gene networks. We also introduce a constraint that lightens the model by focusing on learning the predefined graph structure while ensuring further hops are deducted to expand the receptive field. Empirical studies show that our model improves by an average of 1.27% on cell-type annotation and 3.86% on perturbation prediction. Furthermore, our method significantly decreases computational overhead. Overall, GenoHoption can function as an efficient and expressive bridge, connecting existing single-cell foundation models to gene network graphs.1
Jiabei Cheng, Kaiyuan Yang 0009, Hong-Bin Shen
BIBM4
2024 Multimodal variational contrastive learning for few-shot classification
Meihong Pan, Hong-Bin Shen
Appl. Intell.2
2024 MINDG: a drug-target interaction prediction method based on an integrated learning algorithm
abstract
MOTIVATION: Drug-target interaction (DTI) prediction refers to the prediction of whether a given drug molecule will bind to a specific target and thus exert a targeted therapeutic effect. Although intelligent computational approaches for drug target prediction have received much attention and made many advances, they are still a challenging task that requires further research. The main challenges are manifested as follows: (i) most graph neural network-based methods only consider the information of the first-order neighboring nodes (drug and target) in the graph, without learning deeper and richer structural features from the higher-order neighboring nodes. (ii) Existing methods do not consider both the sequence and structural features of drugs and targets, and each method is independent of each other, and cannot combine the advantages of sequence and structural features to improve the interactive learning effect. RESULTS: To address the above challenges, a Multi-view Integrated learning Network that integrates Deep learning and Graph Learning (MINDG) is proposed in this study, which consists of the following parts: (i) a mixed deep network is used to extract sequence features of drugs and targets, (ii) a higher-order graph attention convolutional network is proposed to better extract and capture structural features, and (iii) a multi-view adaptive integrated decision module is used to improve and complement the initial prediction results of the above two networks to enhance the prediction performance. We evaluate MINDG on two dataset and show it improved DTI prediction performance compared to state-of-the-art baselines. AVAILABILITY AND IMPLEMENTATION: https://github.com/jnuaipr/MINDG.
Hailong Yang 0001, Yun Zuo 0001, Zhaohong Deng, Xiaoyong Pan, Hong-Bin Shen, Kup-Sze Choi, Dongjun Yu
Bioinform.6
2024 Semantic-Based Implicit Feature Transform for Few-Shot Classification
Meihong Pan, Hongyi Xin, Hong-Bin Shen
Int. J. Comput. Vis.3
2024 Cross-modal de-deviation for enhancing few-shot classification
Meihong Pan, Hong-Bin Shen
Pattern Recognit.2
2023 Isoform Function Prediction Based on Heterogeneous Graph Attention Networks
abstract
Isoforms refer to different mRNA molecules transcribed from the same gene, which can be translated into proteins with varying structures and functions. Predicting the functions of isoforms is an essential topic in bioinformatics as it can provide valuable insights into the intricate mechanisms of gene regulation and biological processes. Conventionally, gene function labels are standardized in Gene Ontology (GO) terms. However, traditional methods for predicting isoform function are largely limited by the absence of isoform-specific labels, sparse annotations, and the vast number of GO terms. To address these issues, we propose HANIso, a deep learning-based method for isoform function prediction. HANIso leverages a pretrained protein language model to extract features from protein sequences. It also integrates heterogeneous information, such as isoform sequence features, GO annotations, and isoform interaction data, using a Heterogeneous Graph Attention Network (HAN). This allows the model to learn the importance of different sources of information and their semantic relationships through the attention mechanism. Our method can predict function labels at both the gene level and isoform level. We conduct experiments on two species datasets, and the results demonstrate that our method outperforms existing methods on both AUROC and AUPRC. HANIso has the potential to overcome the limitations of traditional methods and provide a more accurate and comprehensive understanding of isoform function.
Kuo Guo, Hong-Bin Shen, Yang Yang 0030
BIBM4
2023 Leveraging scaffold information to predict protein-ligand binding affinity with an empirical graph neural network
abstract
Protein-ligand binding affinity prediction is an important task in structural bioinformatics for drug discovery and design. Although various scoring functions (SFs) have been proposed, it remains challenging to accurately evaluate the binding affinity of a protein-ligand complex with the known bound structure because of the potential preference of scoring system. In recent years, deep learning (DL) techniques have been applied to SFs without sophisticated feature engineering. Nevertheless, existing methods cannot model the differential contribution of atoms in various regions of proteins, and the relationship between atom properties and intermolecular distance is also not fully explored. We propose a novel empirical graph neural network for accurate protein-ligand binding affinity prediction (EGNA). Graphs of protein, ligand and their interactions are constructed based on different regions of each bound complex. Proteins and ligands are effectively represented by graph convolutional layers, enabling the EGNA to capture interaction patterns precisely by simulating empirical SFs. The contributions of different factors on binding affinity can thus be transparently investigated. EGNA is compared with the state-of-the-art machine learning-based SFs on two widely used benchmark data sets. The results demonstrate the superiority of EGNA and its good generalization capability.
Chun-Qiu Xia, Shi-Hao Feng, Xiaoyong Pan, Hong-Bin Shen
Briefings Bioinform.5
2023 High-accuracy protein model quality assessment using attention graph neural networks
abstract
Great improvement has been brought to protein tertiary structure prediction through deep learning. It is important but very challenging to accurately rank and score decoy structures predicted by different models. CASP14 results show that existing quality assessment (QA) approaches lag behind the development of protein structure prediction methods, where almost all existing QA models degrade in accuracy when the target is a decoy of high quality. How to give an accurate assessment to high-accuracy decoys is particularly useful with the available of accurate structure prediction methods. Here we propose a fast and effective single-model QA method, QATEN, which can evaluate decoys only by their topological characteristics and atomic types. Our model uses graph neural networks and attention mechanisms to evaluate global and amino acid level scores, and uses specific loss functions to constrain the network to focus more on high-precision decoys and protein domains. On the CASP14 evaluation decoys, QATEN performs better than other QA models under all correlation coefficients when targeting average LDDT. QATEN shows promising performance when considering only high-accuracy decoys. Compared to the embedded evaluation modules of predicted ${C}_{\alpha^{-}} RMSD$ (pRMSD) in RosettaFold and predicted LDDT (pLDDT) in AlphaFold2, QATEN is complementary and capable of achieving better evaluation on some decoy structures generated by AlphaFold2 and RosettaFold. These results suggest that the new QATEN approach can be used as a reliable independent assessment algorithm for high-accuracy protein structure decoys.
Chun-Qiu Xia, Hong-Bin Shen
Briefings Bioinform.3
2023 De novodrug design by iterative multiobjective deep reinforcement learning with graph-based molecular quality assessment
abstract
MOTIVATION: Generating molecules of high quality and drug-likeness in the vast chemical space is a big challenge in the drug discovery. Most existing molecule generative methods focus on diversity and novelty of molecules, but ignoring drug potentials of the generated molecules during the generation process. RESULTS: In this study, we present a novel de novo multiobjective quality assessment-based drug design approach (QADD), which integrates an iterative refinement framework with a novel graph-based molecular quality assessment model on drug potentials. QADD designs a multiobjective deep reinforcement learning pipeline to generate molecules with multiple desired properties iteratively, where a graph neural network-based model for accurate molecular quality assessment on drug potentials is introduced to guide molecule generation. Experimental results show that QADD can jointly optimize multiple molecular properties with a promising performance and the quality assessment module is capable of guiding the generated molecules with high drug potentials. Furthermore, applying QADD to generate novel molecules binding to a biological target protein DRD2 also demonstrates the algorithm's efficacy. AVAILABILITY AND IMPLEMENTATION: QADD is freely available online for academic use at https://github.com/yifang000/QADD or http://www.csbio.sjtu.edu.cn/bioinf/QADD.
Xiaoyong Pan, Hong-Bin Shen
Bioinform.3
2023 MLNGCF: circRNA-disease associations prediction with multilayer attention neural graph-based collaborative filtering
abstract
MOTIVATION: CircRNAs play a critical regulatory role in physiological processes, and the abnormal expression of circRNAs can mediate the processes of diseases. Therefore, exploring circRNAs-disease associations is gradually becoming an important area of research. Due to the high cost of validating circRNA-disease associations using traditional wet-lab experiments, novel computational methods based on machine learning are gaining more and more attention in this field. However, current computational methods suffer to insufficient consideration of latent features in circRNA-disease interactions. RESULTS: In this study, a multilayer attention neural graph-based collaborative filtering (MLNGCF) is proposed. MLNGCF first enhances multiple biological information with autoencoder as the initial features of circRNAs and diseases. Then, by constructing a central network of different diseases and circRNAs, a multilayer cooperative attention-based message propagation is performed on the central network to obtain the high-order features of circRNAs and diseases. A neural network-based collaborative filtering is constructed to predict the unknown circRNA-disease associations and update the model parameters. Experiments on the benchmark datasets demonstrate that MLNGCF outperforms state-of-the-art methods, and the prediction results are supported by the literature in the case studies. AVAILABILITY AND IMPLEMENTATION: The source codes and benchmark datasets of MLNGCF are available at https://github.com/ABard0/MLNGCF.
Qunzhuo Wu, Zhaohong Deng, Wei Zhang 0221, Xiaoyong Pan, Kup-Sze Choi, Yun Zuo 0001, Hong-Bin Shen, Dongjun Yu
Bioinform.7
2023 Few-shot classification with task-adaptive semantic feature learning
abstract
Few-shot classification aims to learn a classifier that categorizes objects of unseen classes with limited samples. One general approach is to mine as much information as possible from limited samples. This can be achieved by incorporating data from multiple modalities. However, existing multi-modality methods only use additional modality in support samples while adhering to a single modal in query samples. Such approach could lead to information imbalance between support and query samples, which confounds model generalization from support to query samples. Towards this problem, we propose a task-adaptive semantic feature learning mechanism to incorporate semantic features for both support and query samples. The semantic feature learner is trained episodic-wisely by regressing from the feature vectors of the support samples. It is utilized to predict semantic features for the query samples. Such method maintains a consistent training scheme between support and query samples and enables direct model transfer from support to query data, which significantly improves model generalization. We conduct extensive experiments on four benchmarks in both inductive and transductive settings. Results show that the proposed TasNet outperforms state-of-the-art methods with an improvement of 1% to 5% in classification accuracy , demonstrating the superiority of our method. The exhaustive ablation studies further validate the effectiveness of our framework. The code is available at: https://github.com/pmhDL/TasNet
Meihong Pan, Hongyi Xin, Chun-Qiu Xia, Hong-Bin Shen
Pattern Recognit.4
2022 circRNA-binding protein site prediction based on multi-view deep learning, subspace learning and multi-view classifier
abstract
Circular RNAs (circRNAs) generally bind to RNA-binding proteins (RBPs) to play an important role in the regulation of autoimmune diseases. Thus, it is crucial to study the binding sites of RBPs on circRNAs. Although many methods, including traditional machine learning and deep learning, have been developed to predict the interactions between RNAs and RBPs, and most of them are focused on linear RNAs. At present, few studies have been done on the binding relationships between circRNAs and RBPs. Thus, in-depth research is urgently needed. In the existing circRNA-RBP binding site prediction methods, circRNA sequences are the main research subjects, but the relevant characteristics of circRNAs have not been fully exploited, such as the structure and composition information of circRNA sequences. Some methods have extracted different views to construct recognition models, but how to efficiently use the multi-view data to construct recognition models is still not well studied. Considering the above problems, this paper proposes a multi-view classification method called DMSK based on multi-view deep learning, subspace learning and multi-view classifier for the identification of circRNA-RBP interaction sites. In the DMSK method, first, we converted circRNA sequences into pseudo-amino acid sequences and pseudo-dipeptide components for extracting high-dimensional sequence features and component features of circRNAs, respectively. Then, the structure prediction method RNAfold was used to predict the secondary structure of the RNA sequences, and the sequence embedding model was used to extract the context-dependent features. Next, we fed the above four views' raw features to a hybrid network, which is composed of a convolutional neural network and a long short-term memory network, to obtain the deep features of circRNAs. Furthermore, we used view-weighted generalized canonical correlation analysis to extract four views' common features by subspace learning. Finally, the learned subspace common features and multi-view deep features were fed to train the downstream multi-view TSK fuzzy system to construct a fuzzy rule and fuzzy inference-based multi-view classifier. The trained classifier was used to predict the specific positions of the RBP binding sites on the circRNAs. The experiments show that the prediction performance of the proposed method DMSK has been improved compared with the existing methods. The code and dataset of this study are available at https://github.com/Rebecca3150/DMSK.
Zhaohong Deng, Xiaoyong Pan, Zhisheng Wei, Hong-Bin Shen, Kup-Sze Choi, Shitong Wang 0001, Jing Wu 0030
Briefings Bioinform.6
2022 SIFLoc: a self-supervised pre-training method for enhancing the recognition of protein subcellular localization in immunofluorescence microscopic images
abstract
With the rapid growth of high-resolution microscopy imaging data, revealing the subcellular map of human proteins has become a central task in the spatial proteome. The cell atlas of the Human Protein Atlas (HPA) provides precious resources for recognizing subcellular localization patterns at the cell level, and the large-scale annotated data enable learning via advanced deep neural networks. However, the existing predictors still suffer from the imbalanced class distribution and the lack of labeled data for minor classes. Thus, it is necessary to develop new methods for coping with these issues. We leverage the self-supervised learning protocol to address these problems. Especially, we propose a pre-training scheme to enhance the conventional supervised learning framework called SIFLoc. The pre-training is featured by a hybrid data augmentation method and a modified contrastive loss function, aiming to learn good feature representations from microscopic images. The experiments are performed on a large-scale immunofluorescence microscopic image dataset collected from the HPA database. Using the same deep neural networks as the classifier, the model pre-trained via SIFLoc not only outperforms the model without pre-training by a large margin but also shows advantages over the state-of-the-art self-supervised learning methods. Especially, SIFLoc improves the prediction accuracy for minor organelles significantly.
Yanlun Tu, Houchao Lei, Hong-Bin Shen, Yang Yang 0030
Briefings Bioinform.3
2022 Learning protein subcellular localization multi-view patterns from heterogeneous data of imaging, sequence and networks
abstract
Location proteomics seeks to provide automated high-resolution descriptions of protein location patterns within cells. Many efforts have been undertaken in location proteomics over the past decades, thereby producing plenty of automated predictors for protein subcellular localization. However, most of these predictors are trained solely from high-throughput microscopic images or protein amino acid sequences alone. Unifying heterogeneous protein data sources has yet to be exploited. In this paper, we present a pipeline called sequence, image, network-based protein subcellular locator (SIN-Locator) that constructs a multi-view description of proteins by integrating multiple data types including images of protein expression in cells or tissues, amino acid sequences and protein-protein interaction networks, to classify the patterns of protein subcellular locations. Proteins were encoded by both handcrafted features and deep learning features, and multiple combining methods were implemented. Our experimental results indicated that optimal integrations can considerately enhance the classification accuracy, and the utility of SIN-Locator has been demonstrated through applying to new released proteins in the human protein atlas. Furthermore, we also investigate the contribution of different data sources and influence of partial absence of data. This work is anticipated to provide clues for reconciliation and combination of multi-source data for protein location analysis.
Min-Qi Xue, Hong-Bin Shen, Ying-Ying Xu
Briefings Bioinform.3
2022 MDGF-MCEC: a multi-view dual attention embedding model with cooperative ensemble learning for CircRNA-disease association prediction
abstract
Circular RNA (circRNA) is closely involved in physiological and pathological processes of many diseases. Discovering the associations between circRNAs and diseases is of great significance. Due to the high-cost to verify the circRNA-disease associations by wet-lab experiments, computational approaches for predicting the associations become a promising research direction. In this paper, we propose a method, MDGF-MCEC, based on multi-view dual attention graph convolution network (GCN) with cooperative ensemble learning to predict circRNA-disease associations. First, MDGF-MCEC constructs two disease relation graphs and two circRNA relation graphs based on different similarities. Then, the relation graphs are fed into a multi-view GCN for representation learning. In order to learn high discriminative features, a dual-attention mechanism is introduced to adjust the contribution weights, at both channel level and spatial level, of different features. Based on the learned embedding features of diseases and circRNAs, nine different feature combinations between diseases and circRNAs are treated as new multi-view data. Finally, we construct a multi-view cooperative ensemble classifier to predict the associations between circRNAs and diseases. Experiments conducted on the CircR2Disease database demonstrate that the proposed MDGF-MCEC model achieves a high area under curve of 0.9744 and outperforms the state-of-the-art methods. Promising results are also obtained from experiments on the circ2Disease and circRNADisease databases. Furthermore, the predicted associated circRNAs for hepatocellular carcinoma and gastric cancer are supported by the literature. The code and dataset of this study are available at https://github.com/ABard0/MDGF-MCEC.
Qunzhuo Wu, Zhaohong Deng, Xiaoyong Pan, Hong-Bin Shen, Kup-Sze Choi, Shitong Wang 0001, Jing Wu 0030, Dongjun Yu
Briefings Bioinform.4
2022 Accurate flexible refinement for atomic-level protein structure using cryo-EM density maps and deep learning
abstract
With the rapid progress of deep learning in cryo-electron microscopy and protein structure prediction, improving the accuracy of the protein structure model by using a density map and predicted contact/distance map through deep learning has become an urgent need for robust methods. Thus, designing an effective protein structure optimization strategy based on the density map and predicted contact/distance map is critical to improving the accuracy of structure refinement. In this article, a protein structure optimization method based on the density map and predicted contact/distance map by deep-learning technology was proposed in accordance with the result of matching between the density map and the initial model. Physics- and knowledge-based energy functions, integrated with Cryo-EM density map data and deep-learning data, were used to optimize the protein structure in the simulation. The dynamic confidence score was introduced to the iterative process for choosing whether it is a density map or a contact/distance map to dominate the movement in the simulation to improve the accuracy of refinement. The protocol was tested on a large set of 224 non-homologous membrane proteins and generated 214 structural models with correct folds, where 4.5% of structural models were generated from structural models with incorrect folds. Compared with other state-of-the-art methods, the major advantage of the proposed methods lies in the skills for using density map and contact/distance map in the simulation, as well as the new energy function in the re-assembly simulations. Overall, the results demonstrated that this strategy is a valuable approach and ready to use for atomic-level structure refinement using cryo-EM density map and predicted contact/distance map.
Yang Zhang 0040, Hong-Bin Shen, Guijun Zhang
Briefings Bioinform.4
2022 CoCoPRED: coiled-coil protein structural feature prediction from amino acid sequence using deep neural networks
abstract
MOTIVATION: Coiled-coil is composed of two or more helices that are wound around each other. It widely exists in proteins and has been discovered to play a variety of critical roles in biology processes. Generally, there are three types of structural features in coiled-coil: coiled-coil domain (CCD), oligomeric state and register. However, most of the existing computational tools only focus on one of them. RESULTS: Here, we describe a new deep learning model, CoCoPRED, which is based on convolutional layers, bidirectional long short-term memory, and attention mechanism. It has three networks, i.e. CCD network, oligomeric state network, and register network, corresponding to the three types of structural features in coiled-coil. This means CoCoPRED has the ability of fulfilling comprehensive prediction for coiled-coil proteins. Through the 5-fold cross-validation experiment, we demonstrate that CoCoPRED can achieve better performance than the state-of-the-art models on both CCD prediction and oligomeric state prediction. Further analysis suggests the CCD prediction may be a performance indicator of the oligomeric state prediction in CoCoPRED. The attention heads in CoCoPRED indicate that registers a, b and e are more crucial for the oligomeric state prediction. AVAILABILITY AND IMPLEMENTATION: CoCoPRED is available at http://www.csbio.sjtu.edu.cn/bioinf/CoCoPRED. The datasets used in this research can also be downloaded from the website. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Shi-Hao Feng, Chun-Qiu Xia, Hong-Bin Shen
Bioinform.3
2022 GraphLoc: a graph neural network model for predicting protein subcellular localization from immunohistochemistry images
abstract
MOTIVATION: Recognition of protein subcellular distribution patterns and identification of location biomarker proteins in cancer tissues are important for understanding protein functions and related diseases. Immunohistochemical (IHC) images enable visualizing the distribution of proteins at the tissue level, providing an important resource for the protein localization studies. In the past decades, several image-based protein subcellular location prediction methods have been developed, but the prediction accuracies still have much space to improve due to the complexity of protein patterns resulting from multi-label proteins and the variation of location patterns across cell types or states. RESULTS: Here, we propose a multi-label multi-instance model based on deep graph convolutional neural networks, GraphLoc, to recognize protein subcellular location patterns. GraphLoc builds a graph of multiple IHC images for one protein, learns protein-level representations by graph convolutions and predicts multi-label information by a dynamic threshold method. Our results show that GraphLoc is a promising model for image-based protein subcellular location prediction with model interpretability. Furthermore, we apply GraphLoc to the identification of candidate location biomarkers and potential members for protein networks. A large portion of the predicted results have supporting evidence from the existing literatures and the new candidates also provide guidance for further experimental screening. AVAILABILITY AND IMPLEMENTATION: The dataset and code are available at: www.csbio.sjtu.edu.cn/bioinf/GraphLoc. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jin-Xian Hu, Yang Yang 0030, Ying-Ying Xu, Hong-Bin Shen
Bioinform.4
2022 Accurate inference of gene regulatory interactions from spatial gene expression with deep contrastive learning
abstract
MOTIVATION: Reverse engineering of gene regulatory networks (GRNs) has long been an attractive research topic in system biology. Computational prediction of gene regulatory interactions has remained a challenging problem due to the complexity of gene expression and scarce information resources. The high-throughput spatial gene expression data, like in situ hybridization images that exhibit temporal and spatial expression patterns, has provided abundant and reliable information for the inference of GRNs. However, computational tools for analyzing the spatial gene expression data are highly underdeveloped. RESULTS: In this study, we develop a new method for identifying gene regulatory interactions from gene expression images, called ConGRI. The method is featured by a contrastive learning scheme and deep Siamese convolutional neural network architecture, which automatically learns high-level feature embeddings for the expression images and then feeds the embeddings to an artificial neural network to determine whether or not the interaction exists. We apply the method to a Drosophila embryogenesis dataset and identify GRNs of eye development and mesoderm development. Experimental results show that ConGRI outperforms previous traditional and deep learning methods by a large margin, which achieves accuracies of 76.7% and 68.7% for the GRNs of early eye development and mesoderm development, respectively. It also reveals some master regulators for Drosophila eye development. AVAILABILITYAND IMPLEMENTATION: https://github.com/lugimzheng/ConGRI. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lujing Zheng, Zhenhuan Liu, Yang Yang 0030, Hong-Bin Shen
Bioinform.4
2022 Fast protein structure comparison through effective representation learning with contrastive graph neural networks
abstract
Protein structure alignment algorithms are often time-consuming, resulting in challenges for large-scale protein structure similarity-based retrieval. There is an urgent need for more efficient structure comparison approaches as the number of protein structures increases rapidly. In this paper, we propose an effective graph-based protein structure representation learning method, GraSR, for fast and accurate structure comparison. In GraSR, a graph is constructed based on the intra-residue distance derived from the tertiary structure. Then, deep graph neural networks (GNNs) with a short-cut connection learn graph representations of the tertiary structures under a contrastive learning framework. To further improve GraSR, a novel dynamic training data partition strategy and length-scaling cosine distance are introduced. We objectively evaluate our method GraSR on SCOPe v2.07 and a new released independent test set from PDB database with a designed comprehensive performance metric. Compared with other state-of-the-art methods, GraSR achieves about 7%-10% improvement on two benchmark datasets. GraSR is also much faster than alignment-based methods. We dig into the model and observe that the superiority of GraSR is mainly brought by the learned discriminative residue-level and global descriptors. The web-server and source code of GraSR are freely available at www.csbio.sjtu.edu.cn/bioinf/GraSR/ for academic use.
Chun-Qiu Xia, Shi-Hao Feng, Xiaoyong Pan, Hong-Bin Shen
PLoS Comput. Biol.5
2022 Ab-Initio Membrane Protein Amphipathic Helix Structure Prediction Using Deep Neural Networks
abstract
Amphipathic helix (AH)features the segregation of polar and nonpolar residues and plays important roles in many membrane-associated biological processes through interacting with both the lipid and the soluble phases. Although the AH structure has been discovered for a long time, few ab initio machine learning-based prediction models have been reported, due to the limited amount of training data. In this study, we report a new deep learning-based prediction model, which is composed of a residual neural network and the uneven-thresholds decision algorithm. It is constructed on 121 membrane proteins, in total 51640 residue samples, which are curated from an up-to-date membrane protein structure database. Through a rigid 10-fold nested cross-validation experiment, we demonstrate that our model can achieve promising predictions and exceed current state-of-the-art approaches in this field. This presents a new avenue for accurately predicting AHs. Analysis on the contribution of the input residues and some cases further reveals the high interpretability and the generalization of our model.
Shi-Hao Feng, Chun-Qiu Xia, Hong-Bin Shen
IEEE ACM Trans. Comput. Biol. Bioinform.4
2022 Transductive Multiview Modeling With Interpretable Rules, Matrix Factorization, and Cooperative Learning
abstract
Multiview fuzzy systems aim to deal with fuzzy modeling in multiview scenarios effectively and to obtain the interpretable model through multiview learning. However, current studies of multiview fuzzy systems still face several challenges, one of which is how to achieve efficient collaboration between multiple views when there are few labeled data. To address this challenge, this article explores a novel transductive multiview fuzzy modeling method. The dependency on labeled data is reduced by integrating transductive learning into the fuzzy model to simultaneously learn both the model and the labels using a novel learning criterion. Matrix factorization is incorporated to further improve the performance of the fuzzy model. In addition, collaborative learning between multiple views is used to enhance the robustness of the model. The experimental results indicate that the proposed method is highly competitive with other multiview learning methods.
Wei Zhang 0221, Zhaohong Deng, Jun Wang 0024, Kup-Sze Choi, Te Zhang, Xiaoqing Luo, Hong-Bin Shen, Wenhao Ying, Shitong Wang 0001
IEEE Trans. Cybern.7
2021 Recognizing binding sites of poorly characterized RNA-binding proteins on circular RNAs using attention Siamese network
abstract
Circular RNAs (circRNAs) interact with RNA-binding proteins (RBPs) to play crucial roles in gene regulation and disease development. Computational approaches have attracted much attention to quickly predict highly potential RBP binding sites on circRNAs using the sequence or structure statistical binding knowledge. Deep learning is one of the popular learning models in this area but usually requires a lot of labeled training data. It would perform unsatisfactorily for the less characterized RBPs with a limited number of known target circRNAs. How to improve the prediction performance for such small-size labeled characterized RBPs is a challenging task for deep learning-based models. In this study, we propose an RBP-specific method iDeepC for predicting RBP binding sites on circRNAs from sequences. It adopts a Siamese neural network consisting of a lightweight attention module and a metric module. We have found that Siamese neural network effectively enhances the network capability of capturing mutual information between circRNAs with pairwise metric learning. To further deal with the small-sample size problem, we have performed the pretraining using available labeled data from other RBPs and also demonstrate the efficacy of this transfer-learning pipeline. We comprehensively evaluated iDeepC on the benchmark datasets of RBP-binding circRNAs, and the results suggest iDeepC achieving promising results on the poorly characterized RBPs. The source code is available at https://github.com/hehew321/iDeepC.
Hehe Wu, Xiaoyong Pan, Yang Yang 0030, Hong-Bin Shen
Briefings Bioinform.4
2021 RNA-binding protein recognition based on multi-view deep feature and multi-label learning
abstract
RNA-binding protein (RBP) is a class of proteins that bind to and accompany RNAs in regulating biological processes. An RBP may have multiple target RNAs, and its aberrant expression can cause multiple diseases. Methods have been designed to predict whether a specific RBP can bind to an RNA and the position of the binding site using binary classification model. However, most of the existing methods do not take into account the binding similarity and correlation between different RBPs. While methods employing multiple labels and Long Short Term Memory Network (LSTM) are proposed to consider binding similarity between different RBPs, the accuracy remains low due to insufficient feature learning and multi-label learning on RNA sequences. In response to this challenge, the concept of RNA-RBP Binding Network (RRBN) is proposed in this paper to provide theoretical support for multi-label learning to identify RBPs that can bind to RNAs. It is experimentally shown that the RRBN information can significantly improve the prediction of unknown RNA-RBP interactions. To further improve the prediction accuracy, we present the novel computational method iDeepMV which integrates multi-view deep learning technology under the multi-label learning framework. iDeepMV first extracts data from the views of amino acid sequence and dipeptide component based on the RNA sequences as the original view. Deep neural network models are then designed for the respective views to perform deep feature learning. The extracted deep features are fed into multi-label classifiers which are trained with the RNA-RBP interaction information for the three views. Finally, a voting mechanism is designed to make comprehensive decision on the results of the multi-label classifiers. Our experimental results show that the prediction performance of iDeepMV, which combines multi-view deep feature learning models with RNA-RBP interaction information, is significantly better than that of the state-of-the-art methods. iDeepMV is freely available at http://www.csbio.sjtu.edu.cn/bioinf/iDeepMV for academic use. The code is freely available at http://github.com/uchihayht/iDeepMV.
Zhaohong Deng, Xiaoyong Pan, Hong-Bin Shen, Kup-Sze Choi, Shitong Wang 0001, Jing Wu 0030
Briefings Bioinform.4
2021 lncLocator 2.0: a cell-line-specific subcellular localization predictor for long non-coding RNAs with interpretable deep learning
abstract
MOTIVATION: Long non-coding RNAs (lncRNAs) are generally expressed in a tissue-specific way, and subcellular localizations of lncRNAs depend on the tissues or cell lines that they are expressed. Previous computational methods for predicting subcellular localizations of lncRNAs do not take this characteristic into account, they train a unified machine learning model for pooled lncRNAs from all available cell lines. It is of importance to develop a cell-line-specific computational method to predict lncRNA locations in different cell lines. RESULTS: In this study, we present an updated cell-line-specific predictor lncLocator 2.0, which trains an end-to-end deep model per cell line, for predicting lncRNA subcellular localization from sequences. We first construct benchmark datasets of lncRNA subcellular localizations for 15 cell lines. Then we learn word embeddings using natural language models, and these learned embeddings are fed into convolutional neural network, long short-term memory and multilayer perceptron to classify subcellular localizations. lncLocator 2.0 achieves varying effectiveness for different cell lines and demonstrates the necessity of training cell-line-specific models. Furthermore, we adopt Integrated Gradients to explain the proposed model in lncLocator 2.0, and find some potential patterns that determine the subcellular localizations of lncRNAs, suggesting that the subcellular localization of lncRNAs is linked to some specific nucleotides. AVAILABILITYAND IMPLEMENTATION: The lncLocator 2.0 is available at www.csbio.sjtu.edu.cn/bioinf/lncLocator2 and the source code can be found at https://github.com/Yang-J-LIN/lncLocator2.
Xiaoyong Pan, Hong-Bin Shen
Bioinform.3
2021 ToxDL: deep learning using primary structure and domain embeddings for assessing protein toxicity
abstract
MOTIVATION: Genetically engineering food crops involves introducing proteins from other species into crop plant species or modifying already existing proteins with gene editing techniques. In addition, newly synthesized proteins can be used as therapeutic protein drugs against diseases. For both research and safety regulation purposes, being able to assess the potential toxicity of newly introduced/synthesized proteins is of high importance. RESULTS: In this study, we present ToxDL, a deep learning-based approach for in silico prediction of protein toxicity from sequence alone. ToxDL consists of (i) a module encompassing a convolutional neural network that has been designed to handle variable-length input sequences, (ii) a domain2vec module for generating protein domain embeddings and (iii) an output module that classifies proteins as toxic or non-toxic, using the outputs of the two aforementioned modules. Independent test results obtained for animal proteins and cross-species transferability results obtained for bacteria proteins indicate that ToxDL outperforms traditional homology-based approaches and state-of-the-art machine-learning techniques. Furthermore, through visualizations based on saliency maps, we are able to verify that the proposed network learns known toxic motifs. Moreover, the saliency maps allow for directed in silico modification of a sequence, thus making it possible to alter its predicted protein toxicity. AVAILABILITY AND IMPLEMENTATION: ToxDL is freely available at http://www.csbio.sjtu.edu.cn/bioinf/ToxDL/. The source code can be found at https://github.com/xypan1232/ToxDL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiaoyong Pan, Jasper Zuallaert, Hong-Bin Shen, Elda Posada Campos, Denys O. Marushchak, Wesley De Neve
Bioinform.4
2021 FlyIT: Drosophila Embryogenesis Image Annotation based on Image Tiling and Convolutional Neural Networks
abstract
With the rise of image-based transcriptomics, spatial gene expression data has become increasingly important for understanding gene regulations from the tissue level down to the cell level. Especially, the gene expression images of Drosophila embryos provide a new data source in the study of Drosophila embryogenesis. It is imperative to develop automatic annotation tools since manual annotation is labor-intensive and requires professional knowledge. Although a lot of image annotation methods have been proposed in the computer vision field, they may not work well for gene expression images, due to the great difference between these two annotation tasks. Besides the apparent difference on images, the annotation is performed at the gene level rather than the image level, where the expression patterns of a gene are recorded in multiple images. Moreover, the annotation terms often correspond to local expression patterns of images, yet they are assigned collectively to groups of images and the relations between the terms and single images are unknown. In order to learn the spatial expression patterns comprehensively for genes, we propose a new method, called FlyIT (image annotation based on Image Tiling and convolutional neural networks for fruit Fly). We implement two versions of FlyIT, learning at image-level and gene-level, respectively. The gene-level version employs an image tiling strategy to get a combined image feature representation for each gene. FlyIT uses a pre-trained ResNet model to obtain feature representation and a new loss function to deal with the class imbalance problem. As the annotation of Drosophila images is a multi-label classification problem, the new loss function considers the difficulty levels for recognizing different labels of the same sample and adjusts the sample weights accordingly. The experimental results on the FlyExpress database show that both the image tiling strategy and the deep architecture lead to the great enhancement of the annotation performance. FlyIT outperforms the existing annotators by a large margin (over 9 percent on AUC and 12 percent on macro F1 for predicting the top 10 terms). It also shows advantages over other deep learning models, including both single-instance and multi-instance learning frameworks.
Tiange Li, Yang Yang 0030, Hong-Bin Shen
IEEE ACM Trans. Comput. Biol. Bioinform.4
2020 NiuEM: A Nested-iterative Unsupervised Learning Model for Single-particle Cryo-EM Image Processing
abstract
Cryo-electron microscopy (cryo-EM) has become a mainstream technology for solving spatial structures of biomacromolecules, while the processing of cryo-EM images is a very challenging task. One of the great challenges is the high noise in the images. A common method is to cluster the images with close projecting angles to get mean images, which are used for 3D reconstruction. However, due to the extremely low signal-to-noise-ratio, common clustering methods often fail to obtain high-quality mean images, leading to poorly reconstructed structures. In this study, we present a new unsupervised learning framework, called NiuEM, to discriminate images captured from different angles and yield cluster-mean images. NiuEM first generates pseudo-labels and then exploits both contrastive loss and cross-entropy loss for training convolutional layers to learn feature representations. Moreover, the pseudo-labels are updated iteratively to enhance the reliability of labels. We assess the performance of NiuEM on four data sets via both visualized and quantitative experiments. Especially, two kinds of metrics are adopted to measure the performance, regarding the clustering quality and the resolution of reconstructed 3D models, respectively. The experimental results show that NiuEM achieves very competitive clustering accuracy in the comparison with the state-of-the-art image clustering methods. Moreover, the cluster mean images yielded by NiuEM lead to better initial 3D models compared with the mainstream reconstruction tools.
Jia-Ming Cai, Wangjie Zheng, Yang Yang 0030, Hong-Bin Shen
BIBM5
2020 ImPLoc: a multi-instance deep learning model for the prediction of protein subcellular localization based on immunohistochemistry images
abstract
MOTIVATION: The tissue atlas of the human protein atlas (HPA) houses immunohistochemistry (IHC) images visualizing the protein distribution from the tissue level down to the cell level, which provide an important resource to study human spatial proteome. Especially, the protein subcellular localization patterns revealed by these images are helpful for understanding protein functions, and the differential localization analysis across normal and cancer tissues lead to new cancer biomarkers. However, computational tools for processing images in this database are highly underdeveloped. The recognition of the localization patterns suffers from the variation in image quality and the difficulty in detecting microscopic targets. RESULTS: We propose a deep multi-instance multi-label model, ImPLoc, to predict the subcellular locations from IHC images. In this model, we employ a deep convolutional neural network-based feature extractor to represent image features, and design a multi-head self-attention encoder to aggregate multiple feature vectors for subsequent prediction. We construct a benchmark dataset of 1186 proteins including 7855 images from HPA and 6 subcellular locations. The experimental results show that ImPLoc achieves significant enhancement on the prediction accuracy compared with the current computational methods. We further apply ImPLoc to a test set of 889 proteins with images from both normal and cancer tissues, and obtain 8 differentially localized proteins with a significance level of 0.05. AVAILABILITY AND IMPLEMENTATION: https://github.com/yl2019lw/ImPloc. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yang Yang 0030, Hong-Bin Shen
Bioinform.3
2020 Artificial intelligence-based multi-objective optimization protocol for protein structure refinement
abstract
MOTIVATION: Protein structure refinement is an important step of protein structure prediction. Existing approaches have generally used a single scoring function combined with Monte Carlo method or Molecular Dynamics algorithm. The one-dimension optimization of a single energy function may take the structure too far away without a constraint. The basic motivation of our study is to reduce the bias problem caused by minimizing only a single energy function due to the very diversity of different protein structures. RESULTS: We report a new Artificial Intelligence-based protein structure Refinement method called AIR. Its fundamental idea is to use multiple energy functions as multi-objectives in an effort to correct the potential inaccuracy from a single function. A multi-objective particle swarm optimization algorithm-based structure refinement is designed, where each structure is considered as a particle in the protocol. With the refinement iterations, the particles move around. The quality of particles in each iteration is evaluated by three energy functions, and the non-dominated particles are put into a set called Pareto set. After enough iteration times, particles from the Pareto set are screened and part of the top solutions are outputted as the final refined structures. The multi-objective energy function optimization strategy designed in the AIR protocol provides a different constraint view of the structure, by extending the one-dimension optimization to a new three-dimension space optimization driven by the multi-objective particle swarm optimization engine. Experimental results on CASP11, CASP12 refinement targets and blind tests in CASP 13 turn to be promising. AVAILABILITY AND IMPLEMENTATION: The AIR is available online at: www.csbio.sjtu.edu.cn/bioinf/AIR/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ling Geng, Yu-Jun Zhao, Yang Yang 0030, Yang Zhang 0040, Hong-Bin Shen
Bioinform.7
2020 Protein-ligand binding residue prediction enhancement through hybrid deep heterogeneous learning of sequence and structure data
abstract
MOTIVATION: Knowledge of protein-ligand binding residues is important for understanding the functions of proteins and their interaction mechanisms. From experimentally solved protein structures, how to accurately identify its potential binding sites of a specific ligand on the protein is still a challenging problem. Compared with structure-alignment-based methods, machine learning algorithms provide an alternative flexible solution which is less dependent on annotated homogeneous protein structures. Several factors are important for an efficient protein-ligand prediction model, e.g. discriminative feature representation and effective learning architecture to deal with both the large-scale and severely imbalanced data. RESULTS: In this study, we propose a novel deep-learning-based method called DELIA for protein-ligand binding residue prediction. In DELIA, a hybrid deep neural network is designed to integrate 1D sequence-based features with 2D structure-based amino acid distance matrices. To overcome the problem of severe data imbalance between the binding and nonbinding residues, strategies of oversampling in mini-batch, random undersampling and stacking ensemble are designed to enhance the model. Experimental results on five benchmark datasets demonstrate the effectiveness of proposed DELIA pipeline. AVAILABILITY AND IMPLEMENTATION: The web server of DELIA is available at www.csbio.sjtu.edu.cn/bioinf/delia/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Chun-Qiu Xia, Xiaoyong Pan, Hong-Bin Shen
Bioinform.3
2020 Learning complex subcellular distribution patterns of proteins via analysis of immunohistochemistry images
abstract
MOTIVATION: Systematic and comprehensive analysis of protein subcellular location as a critical part of proteomics ('location proteomics') has been studied for many years, but annotating protein subcellular locations and understanding variation of the location patterns across various cell types and states is still challenging. RESULTS: In this work, we used immunohistochemistry images from the Human Protein Atlas as the source of subcellular location information, and built classification models for the complex protein spatial distribution in normal and cancerous tissues. The models can automatically estimate the fractions of protein in different subcellular locations, and can help to quantify the changes of protein distribution from normal to cancer tissues. In addition, we examined the extent to which different annotated protein pathways and complexes showed similarity in the locations of their member proteins, and then predicted new potential proteins for these networks. AVAILABILITY AND IMPLEMENTATION: The dataset and code are available at: www.csbio.sjtu.edu.cn/bioinf/complexsubcellularpatterns. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ying-Ying Xu, Hong-Bin Shen, Robert F. Murphy
Bioinform.2
2020 Scoring disease-microRNA associations by integrating disease hierarchy into graph convolutional networks
Xiaoyong Pan, Hong-Bin Shen
Pattern Recognit.2
2019 AnnoFly: annotating Drosophila embryonic images based on an attention-enhanced RNN model
abstract
MOTIVATION: In the post-genomic era, image-based transcriptomics have received huge attention, because the visualization of gene expression distribution is able to reveal spatial and temporal expression pattern, which is significantly important for understanding biological mechanisms. The Berkeley Drosophila Genome Project has collected a large-scale spatial gene expression database for studying Drosophila embryogenesis. Given the expression images, how to annotate them for the study of Drosophila embryonic development is the next urgent task. In order to speed up the labor-intensive labeling work, automatic tools are highly desired. However, conventional image annotation tools are not applicable here, because the labeling is at the gene-level rather than the image-level, where each gene is represented by a bag of multiple related images, showing a multi-instance phenomenon, and the image quality varies by image orientations and experiment batches. Moreover, different local regions of an image correspond to different CV annotation terms, i.e. an image has multiple labels. Designing an accurate annotation tool in such a multi-instance multi-label scenario is a very challenging task. RESULTS: To address these challenges, we develop a new annotator for the fruit fly embryonic images, called AnnoFly. Driven by an attention-enhanced RNN model, it can weight images of different qualities, so as to focus on the most informative image patterns. We assess the new model on three standard datasets. The experimental results reveal that the attention-based model provides a transparent approach for identifying the important images for labeling, and it substantially enhances the accuracy compared with the existing annotation methods, including both single-instance and multi-instance learning methods. AVAILABILITY AND IMPLEMENTATION: http://www.csbio.sjtu.edu.cn/bioinf/annofly/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yang Yang 0030, Qingwei Fang, Hong-Bin Shen
Bioinform.4
2019 Identifying RNA-binding proteins using multi-label deep learning
Xiaoyong Pan, Yong-Xian Fan, Jue Jia, Hong-Bin Shen
Sci. China Inf. Sci.4
2019 Predicting gene regulatory interactions based on spatial gene expression data and deep learning
abstract
Reverse engineering of gene regulatory networks (GRNs) is a central task in systems biology. Most of the existing methods for GRN inference rely on gene co-expression analysis or TF-target binding information, where the determination of co-expression is often unreliable merely based on gene expression levels, and the TF-target binding data from high-throughput experiments may be noisy, leading to a high ratio of false links and missed links, especially for large-scale networks. In recent years, the microscopy images recording spatial gene expression have become a new resource in GRN reconstruction, as the spatial and temporal expression patterns contain much abundant gene interaction information. Till now, the spatial expression resources have been largely underexploited, and only a few traditional image processing methods have been employed in the image-based GRN reconstruction. Moreover, co-expression analysis using conventional measurements based on image similarity may be inaccurate, because it is the local-pattern consistency rather than global-image-similarity that determines gene-gene interactions. Here we present GripDL (Gene regulatory interaction prediction via Deep Learning), which incorporates high-confidence TF-gene regulation knowledge from previous studies, and constructs GRNs for Drosophila eye development based on Drosophila embryonic gene expression images. Benefitting from the powerful representation ability of deep neural networks and the supervision information of known interactions, the new method outperforms traditional methods with a large margin and reveals new intriguing knowledge about Drosophila eye development.
Yang Yang 0030, Qingwei Fang, Hong-Bin Shen
PLoS Comput. Biol.3
2018 IterVM: An Iterative Model for Single-Particle Cryo-EM Image Clustering Based on Variational Autoencoder and Multi-Reference Alignment
Guowei Ji, Yang Yang 0030, Hong-Bin Shen
BIBM3
2018 HMIML: Hierarchical Multi-Instance Multi-Label Learning of Drosophila Embryogenesis Images Using Convolutional Neural Networks
Tiange Li, Yang Yang 0030, Hong-Bin Shen
BIBM3
2018 Prediction of MicroRNA Subcellular Localization by Using a Sequence-to-Sequence Model
abstract
The subcellular localization of microRNAs (miR-NAs) is closely related with their biological functions. Some recent studies have discovered that microRNAs can target to various cellular compartments, and have abundant localization patterns in cells. However, to the best of our knowledge, there has been no computational tool for predicting miRNA subcellular locations to date. The major reason is that the lack of useful information source largely limits the prediction performance using traditional statistical learning approaches. In this study, we regard this prediction task as a Sequence-to-Sequence learning process and propose an attention-based encoder-decoder model, miRLocator, to identify subcellular locations of human miRNAs. The designed miRLocator uses a bidirectional long short-term memory (BiLSTM) module to encode the input sequences, and an LSTM module to decode these context vectors as location sets. Especially, a new encoding method for RNAs, RNA2Vec, and an entropy-based method are incorporated in the model to determine the input and output representations, respectively. The experimental results show that miRLocator achieves promising prediction accuracy with the limited input information, and outperforms the models using hand-designed features and conventional RNN models.
Yiqun Xiao, Jiaxun Cai, Yang Yang 0030, Hai Zhao 0001, Hong-Bin Shen
ICDM5
2018 The lncLocator: a subcellular localization predictor for long non-coding RNAs based on a stacked ensemble classifier
abstract
Motivation: The long non-coding RNA (lncRNA) studies have been hot topics in the field of RNA biology. Recent studies have shown that their subcellular localizations carry important information for understanding their complex biological functions. Considering the costly and time-consuming experiments for identifying subcellular localization of lncRNAs, computational methods are urgently desired. However, to the best of our knowledge, there are no computational tools for predicting the lncRNA subcellular locations to date. Results: In this study, we report an ensemble classifier-based predictor, lncLocator, for predicting the lncRNA subcellular localizations. To fully exploit lncRNA sequence information, we adopt both k-mer features and high-level abstraction features generated by unsupervised deep models, and construct four classifiers by feeding these two types of features to support vector machine (SVM) and random forest (RF), respectively. Then we use a stacked ensemble strategy to combine the four classifiers and get the final prediction results. The current lncLocator can predict five subcellular localizations of lncRNAs, including cytoplasm, nucleus, cytosol, ribosome and exosome, and yield an overall accuracy of 0.59 on the constructed benchmark dataset. Availability and implementation: The lncLocator is available at www.csbio.sjtu.edu.cn/bioinf/lncLocator. Supplementary information: Supplementary data are available at Bioinformatics online.
Xiaoyong Pan, Yang Yang 0030, Hong-Bin Shen
Bioinform.5
2018 Predicting RNA-protein binding sites and motifs through combining local and global deep convolutional neural networks
abstract
Motivation: RNA-binding proteins (RBPs) take over 5-10% of the eukaryotic proteome and play key roles in many biological processes, e.g. gene regulation. Experimental detection of RBP binding sites is still time-intensive and high-costly. Instead, computational prediction of the RBP binding sites using patterns learned from existing annotation knowledge is a fast approach. From the biological point of view, the local structure context derived from local sequences will be recognized by specific RBPs. However, in computational modeling using deep learning, to our best knowledge, only global representations of entire RNA sequences are employed. So far, the local sequence information is ignored in the deep model construction process. Results: In this study, we present a computational method iDeepE to predict RNA-protein binding sites from RNA sequences by combining global and local convolutional neural networks (CNNs). For the global CNN, we pad the RNA sequences into the same length. For the local CNN, we split a RNA sequence into multiple overlapping fixed-length subsequences, where each subsequence is a signal channel of the whole sequence. Next, we train deep CNNs for multiple subsequences and the padded sequences to learn high-level features, respectively. Finally, the outputs from local and global CNNs are combined to improve the prediction. iDeepE demonstrates a better performance over state-of-the-art methods on two large-scale datasets derived from CLIP-seq. We also find that the local CNN runs 1.8 times faster than the global CNN with comparable performance when using GPUs. Our results show that iDeepE has captured experimentally verified binding motifs. Availability and implementation: https://github.com/xypan1232/iDeepE. Supplementary information: Supplementary data are available at Bioinformatics online.
Xiaoyong Pan, Hong-Bin Shen
Bioinform.2
2018 MiRGOFS: a GO-based functional similarity measurement for miRNAs, with applications to the prediction of miRNA subcellular localization and miRNA-disease association
abstract
Motivation: Benefiting from high-throughput experimental technologies, whole-genome analysis of microRNAs (miRNAs) has been more and more common to uncover important regulatory roles of miRNAs and identify miRNA biomarkers for disease diagnosis. As a complementary information to the high-throughput experimental data, domain knowledge like the Gene Ontology and KEGG pathway is usually used to guide gene function analysis. However, functional annotation for miRNAs is scarce in the public databases. Till now, only a few methods have been proposed for measuring the functional similarity between miRNAs based on public annotation data, and these methods cover a very limited number of miRNAs, which are not applicable to large-scale miRNA analysis. Results: In this paper, we propose a new method to measure the functional similarity for miRNAs, called miRGOFS, which has two notable features: (i) it adopts a new GO semantic similarity metric which considers both common ancestors and descendants of GO terms; (i) it computes similarity between GO sets in an asymmetric manner, and weights each GO term by its statistical significance. The miRGOFS-based predictor achieves an F1 of 61.2% on a benchmark dataset of miRNA localization, and AUC values of 87.7 and 81.1% on two benchmark sets of miRNA-disease association, respectively. Compared with the existing functional similarity measurements of miRNAs, miRGOFS has the advantages of higher accuracy and larger coverage of human miRNAs (over 1000 miRNAs). Availability and implementation: http://www.csbio.sjtu.edu.cn/bioinf/MiRGOFS/. Supplementary information: Supplementary data are available at Bioinformatics online.
Yang Yang 0030, Xiaofeng Fu, Wenhao Qu, Yiqun Xiao, Hong-Bin Shen
Bioinform.5
2018 MemBrain-contact 2.0: a new two-stage machine learning model for the prediction enhancement of transmembrane protein residue contacts in the full chain
abstract
MOTIVATION: Inter-residue contacts in proteins have been widely acknowledged to be valuable for protein 3 D structure prediction. Accurate prediction of long-range transmembrane inter-helix residue contacts can significantly improve the quality of simulated membrane protein models. RESULTS: In this paper, we present an updated MemBrain predictor, which aims to predict transmembrane protein residue contacts. Our new model benefits from an efficient learning algorithm that can mine latent structural features, which exist in original feature space. The new MemBrain is a two-stage inter-helix contact predictor. The first stage takes sequence-based features as inputs and outputs coarse contact probabilities for each residue pair, which will be further fed into convolutional neural network together with predictions from three direct-coupling analysis approaches in the second stage. Experimental results on the training dataset show that our method achieves an average accuracy of 81.6% for the top L/5 predictions using a strict sequence-based jackknife cross-validation. Evaluated on the test dataset, MemBrain can achieve 79.4% prediction accuracy. Moreover, for the top L/5 predicted long-range loop contacts, the prediction performance can reach an accuracy of 56.4%. These results demonstrate that the new MemBrain is promising for transmembrane protein's contact map prediction. AVAILABILITY AND IMPLEMENTATION: http://www.csbio.sjtu.edu.cn/bioinf/MemBrain/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hong-Bin Shen
Bioinform.2
2018 Bioimage-based protein subcellular location prediction: a comprehensive review
Ying-Ying Xu, Li-Xiu Yao, Hong-Bin Shen
Frontiers Comput. Sci.3
2018 Learning distributed representations of RNA sequences and its application for predicting RNA-protein binding sites with a convolutional neural network
Xiaoyong Pan, Hong-Bin Shen
Neurocomputing2
2018 High density cell tracking with accurate centroid detections and active area-based tracklet clustering
Xu-Hao Zhi, Shu Meng, Hong-Bin Shen
Neurocomputing3
2018 Saliency driven region-edge-based top down level set evolution reveals the asynchronous focus in image segmentation
Xu-Hao Zhi, Hong-Bin Shen
Pattern Recognit.2
2018 An Organelle Correlation-Guided Feature Selection Approach for Classifying Multi-Label Subcellular Bio-Images
abstract
Nowadays, with the advances in microscopic imaging, accurate classification of bioimage-based protein subcellular location pattern has attracted as much attention as ever. One of the basic challenging problems is how to select the useful feature components among thousands of potential features to describe the images. This is not an easy task especially considering there is a high ratio of multi-location proteins. Existing feature selection methods seldom take the correlation among different cellular compartments into consideration, and thus may miss some features that will be co-important for several subcellular locations. To deal with this problem, we make use of the important structural correlation among different cellular compartments and propose an organelle structural correlation regularized feature selection method CSF (Common-Sets of Features) in this paper. We formulate the multi-label classification problem by adopting a group-sparsity regularizer to select common subsets of relevant features from different cellular compartments. In addition, we also add a cell structural correlation regularized Laplacian term, which utilizes the prior biological structural information to capture the intrinsic dependency among different cellular compartments. The CSF provides a new feature selection strategy for multi-label bio-image subcellular pattern classifications, and the experimental results also show its superiority when comparing with several existing algorithms.
Wei Shao 0005, Mingxia Liu 0001, Ying-Ying Xu, Hong-Bin Shen, Daoqiang Zhang
IEEE ACM Trans. Comput. Biol. Bioinform.4
2017 NeBcon: protein contact map prediction using neural network training coupled with naïve Bayes classifiers
abstract
MOTIVATION: Recent CASP experiments have witnessed exciting progress on folding large-size non-humongous proteins with the assistance of co-evolution based contact predictions. The success is however anecdotal due to the requirement of the contact prediction methods for the high volume of sequence homologs that are not available to most of the non-humongous protein targets. Development of efficient methods that can generate balanced and reliable contact maps for different type of protein targets is essential to enhance the success rate of the ab initio protein structure prediction. RESULTS: We developed a new pipeline, NeBcon, which uses the naïve Bayes classifier (NBC) theorem to combine eight state of the art contact methods that are built from co-evolution and machine learning approaches. The posterior probabilities of the NBC model are then trained with intrinsic structural features through neural network learning for the final contact map prediction. NeBcon was tested on 98 non-redundant proteins, which improves the accuracy of the best co-evolution based meta-server predictor by 22%; the magnitude of the improvement increases to 45% for the hard targets that lack sequence and structural homologs in the databases. Detailed data analysis showed that the major contribution to the improvement is due to the optimized NBC combination of the complementary information from both co-evolution and machine learning predictions. The neural network training also helps to improve the coupling of the NBC posterior probability and the intrinsic structural features, which were found particularly important for the proteins that do not have sufficient number of homologous sequences to derive reliable co-evolution profiles. AVAILIABLITY AND IMPLEMENTATION: On-line server and standalone package of the program are available at http://zhanglab.ccmb.med.umich.edu/NeBcon/ . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bao-Ji He, S. M. Mortuza, Hong-Bin Shen, Yang Zhang 0040
Bioinform.4
2017 Hum-mPLoc 3.0: prediction enhancement of human protein subcellular localization through modeling the hidden correlations of gene ontology and functional domain features
abstract
Motivation: Protein subcellular localization prediction has been an important research topic in computational biology over the last decade. Various automatic methods have been proposed to predict locations for large scale protein datasets, where statistical machine learning algorithms are widely used for model construction. A key step in these predictors is encoding the amino acid sequences into feature vectors. Many studies have shown that features extracted from biological domains, such as gene ontology and functional domains, can be very useful for improving the prediction accuracy. However, domain knowledge usually results in redundant features and high-dimensional feature spaces, which may degenerate the performance of machine learning models. Results: In this paper, we propose a new amino acid sequence-based human protein subcellular location prediction approach Hum-mPLoc 3.0, which covers 12 human subcellular localizations. The sequences are represented by multi-view complementary features, i.e. context vocabulary annotation-based gene ontology (GO) terms, peptide-based functional domains, and residue-based statistical features. To systematically reflect the structural hierarchy of the domain knowledge bases, we propose a novel feature representation protocol denoted as HCM (Hidden Correlation Modeling), which will create more compact and discriminative feature vectors by modeling the hidden correlations between annotation terms. Experimental results on four benchmark datasets show that HCM improves prediction accuracy by 5-11% and F 1 by 8-19% compared with conventional GO-based methods. A large-scale application of Hum-mPLoc 3.0 on the whole human proteome reveals proteins co-localization preferences in the cell. Availability and Implementation: www.csbio.sjtu.edu.cn/bioinf/Hum-mPLoc3/. Contacts: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Yang Yang 0030, Hong-Bin Shen
Bioinform.3
2017 RNA-protein binding motifs mining with a new hybrid deep learning based cross-domain knowledge integration approach
abstract
BACKGROUND: RNAs play key roles in cells through the interactions with proteins known as the RNA-binding proteins (RBP) and their binding motifs enable crucial understanding of the post-transcriptional regulation of RNAs. How the RBPs correctly recognize the target RNAs and why they bind specific positions is still far from clear. Machine learning-based algorithms are widely acknowledged to be capable of speeding up this process. Although many automatic tools have been developed to predict the RNA-protein binding sites from the rapidly growing multi-resource data, e.g. sequence, structure, their domain specific features and formats have posed significant computational challenges. One of current difficulties is that the cross-source shared common knowledge is at a higher abstraction level beyond the observed data, resulting in a low efficiency of direct integration of observed data across domains. The other difficulty is how to interpret the prediction results. Existing approaches tend to terminate after outputting the potential discrete binding sites on the sequences, but how to assemble them into the meaningful binding motifs is a topic worth of further investigation. RESULTS: In viewing of these challenges, we propose a deep learning-based framework (iDeep) by using a novel hybrid convolutional neural network and deep belief network to predict the RBP interaction sites and motifs on RNAs. This new protocol is featured by transforming the original observed data into a high-level abstraction feature space using multiple layers of learning blocks, where the shared representations across different domains are integrated. To validate our iDeep method, we performed experiments on 31 large-scale CLIP-seq datasets, and our results show that by integrating multiple sources of data, the average AUC can be improved by 8% compared to the best single-source-based predictor; and through cross-domain knowledge integration at an abstraction level, it outperforms the state-of-the-art predictors by 6%. Besides the overall enhanced prediction performance, the convolutional neural network module embedded in iDeep is also able to automatically capture the interpretable binding motifs for RBPs. Large-scale experiments demonstrate that these mined binding motifs agree well with the experimentally verified results, suggesting iDeep is a promising approach in the real-world applications. CONCLUSION: The iDeep framework not only can achieve promising performance than the state-of-the-art predictors, but also easily capture interpretable binding motifs. iDeep is available at http://www.csbio.sjtu.edu.cn/bioinf/iDeep.
Xiaoyong Pan, Hong-Bin Shen
BMC Bioinform.2
2017 Deep model-based feature extraction for predicting protein subcellular localizations from bio-images
Wei Shao 0005, Hong-Bin Shen, Daoqiang Zhang
Frontiers Comput. Sci.3
2017 Predicting Protein-DNA Binding Residues by Weightedly Combining Sequence-Based Features and Boosting Multiple SVMs
abstract
Protein-DNA interactions are ubiquitous in a wide variety of biological processes. Correctly locating DNA-binding residues solely from protein sequences is an important but challenging task for protein function annotations and drug discovery, especially in the post-genomic era where large volumes of protein sequences have quickly accumulated. In this study, we report a new predictor, named TargetDNA, for targeting protein-DNA binding residues from primary sequences. TargetDNA uses a protein's evolutionary information and its predicted solvent accessibility as two base features and employs a centered linear kernel alignment algorithm to learn the weights for weightedly combining the two features. Based on the weightedly combined feature, multiple initial predictors with SVM as classifiers are trained by applying a random under-sampling technique to the original dataset, the purpose of which is to cope with the severe imbalance phenomenon that exists between the number of DNA-binding and non-binding residues. The final ensembled predictor is obtained by boosting the multiple initially trained predictors. Experimental simulation results demonstrate that the proposed TargetDNA achieves a high prediction performance and outperforms many existing sequence-based protein-DNA binding residue predictors. The TargetDNA web server and datasets are freely available at http://csbio.njust.edu.cn/bioinf/TargetDNA/ for academic use.
Jun Hu 0011, Yang Li 0107, Ming Zhang 0033, Xibei Yang, Hong-Bin Shen, Dongjun Yu
IEEE ACM Trans. Comput. Biol. Bioinform.5
2017 A Nonhomogeneous Cuckoo Search Algorithm Based on Quantum Mechanism for Real Parameter Optimization
abstract
Cuckoo search (CS) algorithm is a nature-inspired search algorithm, in which all the individuals have identical search behaviors. However, this simple homogeneous search behavior is not always optimal to find the potential solution to a special problem, and it may trap the individuals into local regions leading to premature convergence. To overcome the drawback, this paper presents a new variant of CS algorithm with nonhomogeneous search strategies based on quantum mechanism to enhance search ability of the classical CS algorithm. Featured contributions in this paper include: 1) quantum-based strategy is developed for nonhomogeneous update laws and 2) we, for the first time, present a set of theoretical analyses on CS algorithm as well as the proposed algorithm, respectively, and conclude a set of parameter boundaries guaranteeing the convergence of the CS algorithm and the proposed algorithm. On 24 benchmark functions, we compare our method with five existing CS-based methods and other ten state-of-the-art algorithms. The numerical results demonstrate that the proposed algorithm is significantly better than the original CS algorithm and the rest of compared methods according to two nonparametric tests.
Ngaam J. Cheung, Xueming Ding, Hong-Bin Shen
IEEE Trans. Cybern.3
2016 Incorporating organelle correlations into semi-supervised learning for protein subcellular localization prediction
abstract
MOTIVATION: Bioimages of subcellular protein distribution as a new data source have attracted much attention in the field of automated prediction of proteins subcellular localization. Performance of existing systems is significantly limited by the small number of high-quality images with explicit annotations, resulting in the small sample size learning problem. This limitation is more serious for the multi-location proteins that co-exist at two or more organelles, because it is difficult to accurately annotate those proteins by biological experiments or automated systems. RESULTS: In this study, we designed a new protein subcellular localization prediction pipeline aiming to deal with the small sample size learning and multi-location proteins annotation problems. Five semi-supervised algorithms that can make use of lower-quality data were integrated, and a new multi-label classification approach by incorporating the correlations among different organelles in cells was proposed. The organelle correlations were modeled by the Bayesian network, and the topology of the correlation graph was used to guide the order of binary classifiers training in the multi-label classification to reflect the label dependence relationship. The proposed protocol was applied on both immunohistochemistry and immunofluorescence images, and our experimental results demonstrated its efficiency. AVAILABILITY AND IMPLEMENTATION: The datasets and code are available at: www.csbio.sjtu.edu.cn/bioinf/CorrASemiB CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ying-Ying Xu, Hong-Bin Shen
Bioinform.3
2016 R2C: improving ab initio residue contact map prediction using dynamic fusion strategy and Gaussian noise filter
abstract
MOTIVATION: Inter-residue contacts in proteins dictate the topology of protein structures. They are crucial for protein folding and structural stability. Accurate prediction of residue contacts especially for long-range contacts is important to the quality of ab inito structure modeling since they can enforce strong restraints to structure assembly. RESULTS: In this paper, we present a new Residue-Residue Contact predictor called R2C that combines machine learning-based and correlated mutation analysis-based methods, together with a two-dimensional Gaussian noise filter to enhance the long-range residue contact prediction. Our results show that the outputs from the machine learning-based method are concentrated with better performance on short-range contacts; while for correlated mutation analysis-based approach, the predictions are widespread with higher accuracy on long-range contacts. An effective query-driven dynamic fusion strategy proposed here takes full advantages of the two different methods, resulting in an impressive overall accuracy improvement. We also show that the contact map directly from the prediction model contains the interesting Gaussian noise, which has not been discovered before. Different from recent studies that tried to further enhance the quality of contact map by removing its transitive noise, we designed a new two-dimensional Gaussian noise filter, which was especially helpful for reinforcing the long-range residue contact prediction. Tested on recent CASP10/11 datasets, the overall top L/5 accuracy of our final R2C predictor is 17.6%/15.5% higher than the pure machine learning-based method and 7.8%/8.3% higher than the correlated mutation analysis-based approach for the long-range residue contact prediction. AVAILABILITY AND IMPLEMENTATION: http://www.csbio.sjtu.edu.cn/bioinf/R2C/Contact:[email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Qi-Yu Jin, Hong-Bin Shen
Bioinform.4
2016 KNN-based dynamic query-driven sample rescaling strategy for class imbalance learning
Jun Hu 0011, Yang Li 0107, Wuxia Yan, Jing-Yu Yang 0001, Hong-Bin Shen, Dongjun Yu
Neurocomputing5
2016 Protein-protein interaction sites prediction by ensembling SVM and sample-weighted random forests
Zhisen Wei, Jing-Yu Yang 0001, Hong-Bin Shen, Dongjun Yu
Neurocomputing4
2016 Enhancing the Prediction of Transmembrane β-Barrel Segments with Chain Learning and Feature Sparse Representation
abstract
Transmembrane β-barrels (TMBs) are one important class of membrane proteins that play crucial functions in the cell. Membrane proteins are difficult wet-lab targets of structural biology, which call for accurate computational prediction approaches. Here, we developed a novel method named MemBrain-TMB to predict the spanning segments of transmembrane β-barrel from amino acid sequence. MemBrain-TMB is a statistical machine learning-based model, which is constructed using a new chain learning algorithm with input features encoded by the image sparse representation approach. We considered the relative status information between neighboring residues for enhancing the performance, and the matrix of features was translated into feature image by sparse coding algorithm for noise and dimension reduction. To deal with the diverse loop length problem, we applied a dynamic threshold method, which is particularly useful for enhancing the recognition of short loops and tight turns. Our experiments demonstrate that the new protocol designed in MemBrain-TMB effectively helps improve prediction performance.
Xi Yin 0004, Ying-Ying Xu, Hong-Bin Shen
IEEE ACM Trans. Comput. Biol. Bioinform.3
2015 A supervised particle swarm algorithm for real-parameter optimization
Ngaam J. Cheung, Xueming Ding, Hong-Bin Shen
Appl. Intell.3
2015 MACOED: a multi-objective ant colony optimization algorithm for SNP epistasis detection in genome-wide association studies
abstract
MOTIVATION: The existing methods for genetic-interaction detection in genome-wide association studies are designed from different paradigms, and their performances vary considerably for different disease models. One important reason for this variability is that their construction is based on a single-correlation model between SNPs and disease. Due to potential model preference and disease complexity, a single-objective method will therefore not work well in general, resulting in low power and a high false-positive rate. METHOD: In this work, we present a multi-objective heuristic optimization methodology named MACOED for detecting genetic interactions. In MACOED, we combine both logistical regression and Bayesian network methods, which are from opposing schools of statistics. The combination of these two evaluation objectives proved to be complementary, resulting in higher power with a lower false-positive rate than observed for optimizing either objective independently. To solve the space and time complexity for high-dimension problems, a memory-based multi-objective ant colony optimization algorithm is designed in MACOED that is able to retain non-dominated solutions found in past iterations. RESULTS: We compared MACOED with other recent algorithms using both simulated and real datasets. The experimental results demonstrate that our method outperforms others in both detection power and computational feasibility for large datasets. AVAILABILITY AND IMPLEMENTATION: Codes and datasets are available at: www.csbio.sjtu.edu.cn/bioinf/MACOED/.
Peng-Jie Jing, Hong-Bin Shen
Bioinform.2
2015 Bioimaging-based detection of mislocalized proteins in human cancers by semi-supervised learning
abstract
MOTIVATION: There is a long-term interest in the challenging task of finding translocated and mislocated cancer biomarker proteins. Bioimages of subcellular protein distribution are new data sources which have attracted much attention in recent years because of their intuitive and detailed descriptions of protein distribution. However, automated methods in large-scale biomarker screening suffer significantly from the lack of subcellular location annotations for bioimages from cancer tissues. The transfer prediction idea of applying models trained on normal tissue proteins to predict the subcellular locations of cancerous ones is arbitrary because the protein distribution patterns may differ in normal and cancerous states. RESULTS: We developed a new semi-supervised protocol that can use unlabeled cancer protein data in model construction by an iterative and incremental training strategy. Our approach enables us to selectively use the low-quality images in normal states to expand the training sample space and provides a general way for dealing with the small size of annotated images used together with large unannotated ones. Experiments demonstrate that the new semi-supervised protocol can result in improved accuracy and sensitivity of subcellular location difference detection. AVAILABILITY AND IMPLEMENTATION: The data and code are available at: www.csbio.sjtu.edu.cn/bioinf/SemiBiomarker/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ying-Ying Xu, Yang Zhang 0040, Hong-Bin Shen
Bioinform.4
2015 Accurate disulfide-bonding network predictions improve ab initio structure prediction of cysteine-rich proteins
abstract
MOTIVATION: Cysteine-rich proteins cover many important families in nature but there are currently no methods specifically designed for modeling the structure of these proteins. The accuracy of disulfide connectivity pattern prediction, particularly for the proteins of higher-order connections, e.g., >3 bonds, is too low to effectively assist structure assembly simulations. RESULTS: We propose a new hierarchical order reduction protocol called Cyscon for disulfide-bonding prediction. The most confident disulfide bonds are first identified and bonding prediction is then focused on the remaining cysteine residues based on SVR training. Compared with purely machine learning-based approaches, Cyscon improved the average accuracy of connectivity pattern prediction by 21.9%. For proteins with more than 5 disulfide bonds, Cyscon improved the accuracy by 585% on the benchmark set of PDBCYS. When applied to 158 non-redundant cysteine-rich proteins, Cyscon predictions helped increase (or decrease) the TM-score (or RMSD) of the ab initio QUARK modeling by 12.1% (or 14.4%). This result demonstrates a new avenue to improve the ab initio structure modeling for cysteine-rich proteins. AVAILABILITY AND IMPLEMENTATION: http://www.csbio.sjtu.edu.cn/bioinf/Cyscon/ CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bao-Ji He, Richard Jang, Yang Zhang 0040, Hong-Bin Shen
Bioinform.5
2015 Disulfide Connectivity Prediction Based on Modelled Protein 3D Structural Information and Random Forest Regression
abstract
Disulfide connectivity is an important protein structural characteristic. Accurately predicting disulfide connectivity solely from protein sequence helps to improve the intrinsic understanding of protein structure and function, especially in the post-genome era where large volume of sequenced proteins without being functional annotated is quickly accumulated. In this study, a new feature extracted from the predicted protein 3D structural information is proposed and integrated with traditional features to form discriminative features. Based on the extracted features, a random forest regression model is performed to predict protein disulfide connectivity. We compare the proposed method with popular existing predictors by performing both cross-validation and independent validation tests on benchmark datasets. The experimental results demonstrate the superiority of the proposed method over existing predictors. We believe the superiority of the proposed method benefits from both the good discriminative capability of the newly developed features and the powerful modelling capability of the random forest. The web server implementation, called TargetDisulfide, and the benchmark datasets are freely available at: http://csbio.njust.edu.cn/bioinf/TargetDisulfide for academic use.
Dongjun Yu, Yang Li 0107, Jun Hu 0011, Xibei Yang, Jing-Yu Yang 0001, Hong-Bin Shen
IEEE ACM Trans. Comput. Biol. Bioinform.6
2014 Enhancing protein-vitamin binding residues prediction by multiple heterogeneous subspace SVMs ensemble
abstract
BACKGROUND: Vitamins are typical ligands that play critical roles in various metabolic processes. The accurate identification of the vitamin-binding residues solely based on a protein sequence is of significant importance for the functional annotation of proteins, especially in the post-genomic era, when large volumes of protein sequences are accumulating quickly without being functionally annotated. RESULTS: In this paper, a new predictor called TargetVita is designed and implemented for predicting protein-vitamin binding residues using protein sequences. In TargetVita, features derived from the position-specific scoring matrix (PSSM), predicted protein secondary structure, and vitamin binding propensity are combined to form the original feature space; then, several feature subspaces are selected by performing different feature selection methods. Finally, based on the selected feature subspaces, heterogeneous SVMs are trained and then ensembled for performing prediction. CONCLUSIONS: The experimental results obtained with four separate vitamin-binding benchmark datasets demonstrate that the proposed TargetVita is superior to the state-of-the-art vitamin-specific predictor, and an average improvement of 10% in terms of the Matthews correlation coefficient (MCC) was achieved over independent validation tests. The TargetVita web server and the datasets used are freely available for academic use at http://csbio.njust.edu.cn/bioinf/TargetVita or http://www.csbio.sjtu.edu.cn/bioinf/TargetVita.
Dongjun Yu, Jun Hu 0011, Xibei Yang, Jing-Yu Yang 0001, Hong-Bin Shen
BMC Bioinform.6
2014 Predicting pupylation sites in prokaryotic proteins using pseudo-amino acid composition and extreme learning machine
Yong-Xian Fan, Hong-Bin Shen
Neurocomputing2
2014 Image-based classification of protein subcellular location patterns in human reproductive tissue by ensemble learning global and local features
Ying-Ying Xu, Shitong Wang 0001, Hong-Bin Shen
Neurocomputing4
2014 OptiFel: A Convergent Heterogeneous Particle Swarm Optimization Algorithm for Takagi-Sugeno Fuzzy Modeling
abstract
Data-driven design of accurate and reliable Takagi-Sugeno (T-S) fuzzy systems has attracted a lot of attention, where the model structures and parameters are important and often solved in an optimization framework. The particle swarm optimization (PSO) algorithm is widely applied in the field. However, the classical PSO suffers from premature convergence, and it is trapped easily into local optima, which will significantly affect the model accuracy. To overcome these drawbacks, we have developed a new T-S fuzzy system parameters searching strategy called OptiFel with a heterogeneous multiswarm PSO (MsPSO) to enhance the searching performance. MsPSO groups the whole population into multiple cooperative subswarms, which perform different search behaviors for the potential solutions. We have found that the multiple subswarms strategy proposed in this paper is greatly helpful for finding the optimal parameters suitable for the subspaces of the T-S fuzzy model. Our theoretical proof has also demonstrated that the cooperation among the subswarms can maintain a balance between exploration and exploitation to ensure the particles converge to stable points. Experimental results show that MsPSO performs significantly better than traditional PSO algorithms on six benchmark functions. With the improved MsPSO, OptiFel can generate a good fuzzy system model with high accuracy and strong generalization ability.
Ngaam J. Cheung, Xueming Ding, Hong-Bin Shen
IEEE Trans. Fuzzy Syst.3
2013 An image-based multi-label human protein subcellular localization predictor (iLocator) reveals protein mislocalizations in cancer tissues
abstract
MOTIVATION: Human cells are organized into compartments of different biochemical cellular processes. Having proteins appear at the right time to the correct locations in the cellular compartments is required to conduct their functions in normal cells, whereas mislocalization of proteins can result in pathological diseases, including cancer. RESULTS: To reveal the cancer-related protein mislocalizations, we developed an image-based multi-label subcellular location predictor, iLocator, which covers seven cellular localizations. The iLocator incorporates both global and local image descriptors and generates predictions by using an ensemble multi-label classifier. The algorithm has the ability to treat both single- and multiple-location proteins. We first trained and tested iLocator on 3240 normal human tissue images that have known subcellular location information from the human protein atlas. The iLocator was then used to generate protein localization predictions for 3696 protein images from seven cancer tissues that have no location annotations in the human protein atlas. By comparing the output data from normal and cancer tissues, we detected eight potential cancer biomarker proteins that have significant localization differences with P-value < 0.01. AVAILABILITY: http://www.csbio.sjtu.edu.cn/bioinf/iLocator/
Ying-Ying Xu, Yang Zhang 0040, Hong-Bin Shen
Bioinform.4
2013 High-accuracy prediction of transmembrane inter-helix contacts and application to GPCR 3D structure modeling
abstract
MOTIVATION: Residue-residue contacts across the transmembrane helices dictate the three-dimensional topology of alpha-helical membrane proteins. However, contact determination through experiments is difficult because most transmembrane proteins are hard to crystallize. RESULTS: We present a novel method (MemBrain) to derive transmembrane inter-helix contacts from amino acid sequences by combining correlated mutations and multiple machine learning classifiers. Tested on 60 non-redundant polytopic proteins using a strict leave-one-out cross-validation protocol, MemBrain achieves an average accuracy of 62%, which is 12.5% higher than the current best method from the literature. When applied to 13 recently solved G protein-coupled receptors, the MemBrain contact predictions helped increase the TM-score of the I-TASSER models by 37% in the transmembrane region. The number of foldable cases (TM-score >0.5) increased by 100%, where all G protein-coupled receptor templates and homologous templates with sequence identity >30% were excluded. These results demonstrate significant progress in contact prediction and a potential for contact-driven structure modeling of transmembrane proteins. AVAILABILITY: www.csbio.sjtu.edu.cn/bioinf/MemBrain/
Richard Jang, Yang Zhang 0040, Hong-Bin Shen
Bioinform.4
2013 GFO: A data driven approach for optimizing the Gaussian function based similarity metric in computational biology
Jian-Bo Lei, Jiang-Bo Yin, Hong-Bin Shen
Neurocomputing3
2013 Improving protein-ATP binding residues prediction by boosting SVMs with random under-sampling
Dongjun Yu, Jun Hu 0011, Zhenmin Tang, Hong-Bin Shen, Jian Yang 0003, Jing-Yu Yang 0001
Neurocomputing4
2013 Designing Template-Free Predictor for Targeting Protein-Ligand Binding Sites with Classifier Ensemble and Spatial Clustering
abstract
Accurately identifying the protein-ligand binding sites or pockets is of significant importance for both protein function analysis and drug design. Although much progress has been made, challenges remain, especially when the 3D structures of target proteins are not available or no homology templates can be found in the library, where the template-based methods are hard to be applied. In this paper, we report a new ligand-specific template-free predictor called TargetS for targeting protein-ligand binding sites from primary sequences. TargetS first predicts the binding residues along the sequence with ligand-specific strategy and then further identifies the binding sites from the predicted binding residues through a recursive spatial clustering algorithm. Protein evolutionary information, predicted protein secondary structure, and ligand-specific binding propensities of residues are combined to construct discriminative features; an improved AdaBoost classifier ensemble scheme based on random undersampling is proposed to deal with the serious imbalance problem between positive (binding) and negative (nonbinding) samples. Experimental results demonstrate that TargetS achieves high performances and outperforms many existing predictors. TargetS web server and data sets are freely available at: http://www.csbio.sjtu.edu.cn/bioinf/TargetS/ for academic use.
Dongjun Yu, Jun Hu 0011, Hong-Bin Shen, Jinhui Tang 0001, Jing-Yu Yang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2012 Predicting protein-ATP binding sites from primary sequence through fusing bi-profile sampling of multi-view features
abstract
BACKGROUND: Adenosine-5'-triphosphate (ATP) is one of multifunctional nucleotides and plays an important role in cell biology as a coenzyme interacting with proteins. Revealing the binding sites between protein and ATP is significantly important to understand the functionality of the proteins and the mechanisms of protein-ATP complex. RESULTS: In this paper, we propose a novel framework for predicting the proteins' functional residues, through which they can bind with ATP molecules. The new prediction protocol is achieved by combination of sequence evolutional information and bi-profile sampling of multi-view sequential features and the sequence derived structural features. The hypothesis for this strategy is single-view feature can only represent partial target's knowledge and multiple sources of descriptors can be complementary. CONCLUSIONS: Prediction performances evaluated by both 5-fold and leave-one-out jackknife cross-validation tests on two benchmark datasets consisting of 168 and 227 non-homologous ATP binding proteins respectively demonstrate the efficacy of the proposed protocol. Our experimental results also reveal that the residue structural characteristics of real protein-ATP binding sites are significant different from those normal ones, for example the binding residues do not show high solvent accessibility propensities, and the bindings prefer to occur at the conjoint points between different secondary structure segments. Furthermore, results also show that performance is affected by the imbalanced training datasets by testing multiple ratios between positive and negative samples in the experiments. Increasing the dataset scale is also demonstrated useful for improving the prediction performances.
Dongjun Yu, Shu-Sen Li, Yong-Xian Fan, Hong-Bin Shen
BMC Bioinform.6
2011 Gaussian kernel optimization: Complex problem and a simple solution
Jiang-Bo Yin, Hong-Bin Shen
Neurocomputing3
2010 Cascleave: towards more accurate prediction of caspase substrate cleavage sites
abstract
MOTIVATION: The caspase family of cysteine proteases play essential roles in key biological processes such as programmed cell death, differentiation, proliferation, necrosis and inflammation. The complete repertoire of caspase substrates remains to be fully characterized. Accordingly, systematic computational screening studies of caspase substrate cleavage sites may provide insight into the substrate specificity of caspases and further facilitating the discovery of putative novel substrates. RESULTS: In this article we develop an approach (termed Cascleave) to predict both classical (i.e. following a P(1) Asp) and non-typical caspase cleavage sites. When using local sequence-derived profiles, Cascleave successfully predicted 82.2% of the known substrate cleavage sites, with a Matthews correlation coefficient (MCC) of 0.667. We found that prediction performance could be further improved by incorporating information such as predicted solvent accessibility and whether a cleavage sequence lies in a region that is most likely natively unstructured. Novel bi-profile Bayesian signatures were found to significantly improve the prediction performance and yielded the best performance with an overall accuracy of 87.6% and a MCC of 0.747, which is higher accuracy than published methods that essentially rely on amino acid sequence alone. It is anticipated that Cascleave will be a powerful tool for predicting novel substrate cleavage sites of caspases and shedding new insights on the unknown caspase-substrate interactivity relationship. AVAILABILITY: http://sunflower.kuicr.kyoto-u.ac.jp/ approximately sjn/Cascleave/ CONTACT: [email protected]; [email protected]; james; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jiangning Song, Hong-Bin Shen, Khalid Mahmood 0001, Sarah E. Boyd, Geoffrey I. Webb, Tatsuya Akutsu, James C. Whisstock
Bioinform.3
2006 Ensemble classifier for protein fold pattern recognition
abstract
MOTIVATION: Prediction of protein folding patterns is one level deeper than that of protein structural classes, and hence is much more complicated and difficult. To deal with such a challenging problem, the ensemble classifier was introduced. It was formed by a set of basic classifiers, with each trained in different parameter systems, such as predicted secondary structure, hydrophobicity, van der Waals volume, polarity, polarizability, as well as different dimensions of pseudo-amino acid composition, which were extracted from a training dataset. The operation engine for the constituent individual classifiers was OET-KNN (optimized evidence-theoretic k-nearest neighbors) rule. Their outcomes were combined through a weighted voting to give a final determination for classifying a query protein. The recognition was to find the true fold among the 27 possible patterns. RESULTS: The overall success rate thus obtained was 62% for a testing dataset where most of the proteins have <25% sequence identity with the proteins used in training the classifier. Such a rate is 6-21% higher than the corresponding rates obtained by various existing NN (neural networks) and SVM (support vector machines) approaches, implying that the ensemble classifier is very promising and might become a useful vehicle in protein science, as well as proteomics and bioinformatics. AVAILABILITY: The ensemble classifier, called PFP-Pred, is available as a web-server at http://202.120.37.186/bioinf/fold/PFP-Pred.htm for public usage.
Hong-Bin Shen, Kuo-Chen Chou
Bioinform.1
2006 Attribute weighted mercer kernel based fuzzy clustering algorithm for general non-spherical datasets
Hong-Bin Shen, Jie Yang 0002, Shitong Wang 0001
Soft Comput.1
2005 Performing clustering analysis on collaborative models
Hong-Bin Shen, Jie Yang 0002, Ningjiang Chen, Shitong Wang 0001
Intell. Data Anal.1
2005 Fuzzy taxonomy, quantitative database and mining generalized association rules
Shitong Wang 0001, Korris Fu-Lai Chung, Hong-Bin Shen
Intell. Data Anal.3
2004 Note on the relationship between probabilistic and fuzzy clustering
Korris Fu-Lai Chung, Shitong Wang 0001, Hong-Bin Shen, Ruiqiang Zhu
Soft Comput.3
2004 Note on the relationship between probabilistic and fuzzy clustering
Shitong Wang 0001, Korris Fu-Lai Chung, Hong-Bin Shen, Ruiqiang Zhu
Soft Comput.3