EDBT 2026 Demo / reviewers in the wild / expert
Bin Liu 0014
dblp:35/837-14
· DBLP profile ↗
86ranked-venue papers
20as first author
49since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 73 · 19 first-author · 46 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PepLM-GNN: A graph neural network framework leveraging pre-trained language models for peptide-protein binding predictionabstractMOTIVATION: The precise prediction of peptide-protein interaction (PepPI) is a core support for promoting breakthroughs in peptide drug research, as well as understanding the regulatory mechanisms of biomolecules. Researchers have developed several computational methods to predict PepPI. However, existing computational methods also have significant limitations. At the level of data feature characterisation, the problem of PepPI does not conform to the Euclidean axioms, making it difficult for conventional prediction methods to effectively measure the underlying correlations between peptides and proteins. At the level of model generalisation performance, existing approaches are often hampered by insufficient generalisation ability, as manifested by their markedly degraded performance in cold start scenarios involving novel peptides, novel proteins, and novel binding pairs. RESULTS: In this study, we propose a computing framework, PepLM-GNN, that integrates a pre-trained language ProtT5 model with a hybrid graph network for accurate identification of PepPI. This model constructs a graph by using ProtT5-extracted semantic context features of peptides and proteins to form heterogeneous nodes, with edges connecting interacting peptide-protein pairs. The hybrid graph network Graph Convolutional Networks (GCN) provides the comprehensive information of the peptide and protein sequences, while employing the Graph Isomorphism Network (GIN) to capture the global interactions between them. Specifically, the GCN aggregates both the semantic context information of node sequences and local neighbourhood information, effectively representing non-Euclidean data. To capture the global associations, we adopt a GIN strategy to optimize the cross-node feature interaction and transfer process, thereby enhancing the generalisation performance of addressing the cold start scenario. Compared with the existing advanced methods, PepLM-GNN demonstrated highly accurate performance and robustness in predicting the PepPI. We further demonstrated the capabilities of PepLM-GNN in virtual peptide drug screening, which is expected to facilitate the discovery of peptide drugs and the elucidation of protein functions. Ke Yan 0003, Meijing Li, Shutao Chen, Bin Liu 0014 |
PLoS Comput. Biol. | 6 |
| 2026 | MMLmiRLocNet: miRNA Subcellular Localization Prediction Based on Multi-View Multi-Label Learning for Drug DesignabstractIdentifying subcellular localization of microRNAs (miRNAs) is essential for comprehensive understanding of cellular function and has significant implications for drug design. In the past, several computational methods for miRNA subcellular localization is being used for uncovering multiple facets of RNA function to facilitate the biological applications. Unfortunately, most existing classification methods rely on a single sequence-based view, making the effective fusion of data from multiple heterogeneous networks a primary challenge. Inspired by multi-view multi-label learning strategy, we propose a computational method, named MMLmiRLocNet, for predicting the subcellular localizations of miRNAs. The MMLmiRLocNet predictor extracts multi-perspective sequence representations by analyzing lexical, syntactic, and semantic aspects of biological sequences. Specifically, it integrates lexical attributes derived from k-mer physicochemical profiles, syntactic characteristics obtained via word2vec embeddings, and semantic representations generated by pre-trained feature embeddings. Finally, module for extracting multi-view consensus-level features and specific-level features was constructed to capture consensus and specific features from various perspectives. The full connection networks are utilized as the output module to predict the miRNA subcellular localization. Experimental results suggest that MMLmiRLocNet outperforms existing methods in terms of F1, subACC, and Accuracy, and achieves best performance with the help of multi-view consensus features and specific features extract network. Junxi Xie, Bin Liu 0014 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | MT-IDR: Disordered Flexible Linkers and Molecular Recognition Features Prediction Based on Multi-Task LearningabstractWith the continuous development of research on intrinsically disordered proteins, experimental functional annotation methods provide researchers with some functional annotations of intrinsically disordered regions(IDRs), shifting research on the functions of IDRs toward data-driven approaches. Disordered Flexible Linkers(DFLs) and Molecular Recognition Features(MoRFs) are two of the most annotated functions. Moreover, the amount of labeled data is too small, and it is difficult to build a sequence-level model. Therefore, for DFLs and MoRFs prediction, we proposed MT-IDR (intrinsically disordered region function prediction based on multi-task learning), a solution based on sequence hierarchy with multi-task learning. First, based on the amino acid sequence of the protein, evolutionary information, physicochemical properties and structural properties were calculated, then the enhanced protein representation module was used to improve the features. Second, the enhanced protein representation is fed into a multi-task learning model for training. In the independent test sets of two functional prediction tasks, MT-IDR has excellent performance. In the DFLs prediction task, the AUC values of both the TE82 and TE64 reached 0.790 and 0.812. In the MoRFs prediction task, the EXP53 achieves the best performance of single models. The higher performance of MT-IDR provides the opportunity for large-scale screening of DFLs and MoRFs. Liang Yu 0002, Haozheng Li, JingTao Liu, Yuchuan Peng, Bin Liu 0014 |
BIBM | 6 |
| 2025 | FusionEncoder: identification of intrinsically disordered regions based on multi-feature fusionabstractMOTIVATION: Intrinsic disorder regions (IDRs) play a significant role in diverse biological processes and are widely distributed in proteins. Thus, accurately predicting these regions is essential for analyzing protein structure and function. Amino acid feature extraction servers as a foundational process in the development of computational predictive models. Existing methods typically rely on traditional biological features (e.g. PSSM) or use pre-trained protein language models (PPLMs) to capture sequence semantic information, often resorting to straightforward feature concatenation. However, these approaches fail to capture the multi-semantic interactions between traditional biological features and PPLMs-based features. RESULTS: In this study, we propose a method named FusionEncoder designed for the integration of traditional biological and PPLMs-based features of the protein. FusionEncoder is a fusion network built on a variant of long short-term memory (LSTM). We consider traditional biological features and PPLMs-based features to be two types of semantic inputs within a "multi-semantic" space. Traditional features are input into the cell state of the LSTM, while PPLMs-based features are fed into the input part. A fusion cell is then utilized to fuse these two types of features. This strategy leverages the capability of LSTM to encode long sequences, enhancing context-aware semantic learning of amino acid sequences. Finally, a transformer-based encoder layer is employed to predict the IDRs. Evaluation on four independent test datasets indicate that FusionEncoder obviously improves the accuracy of amino acid feature representation and achieves superior performance compared to the other existing methods. AVAILABILITY AND IMPLEMENTATION: To facilitate accessibility for experimental researchers, a user-friendly and publicly available webserver for the FusionEncoder predictor has been deployed at http://bliulab.net/FusionEncoder/. FusionEncoder is expected to serve as a valuable tool for the accurate identification of IDRs. Sicen Liu, Shutao Chen, Bin Liu 0014 |
Bioinform. | 4 |
| 2025 | Accurate prediction of toxicity peptide and its function using multi-view tensor learning and latent semantic learning frameworkabstractMOTIVATION: Therapeutic peptide is an important ingredient in the treatment of various diseases and drug discovery. The toxicity of peptides is one of the major challenges in peptide drug therapy. With the abundance of therapeutic peptides generated in the post-genomics era, it is a challenge to promptly identify toxicity peptides using computational methods. Although several efforts have been made, few algorithms are designed to identify whether a query peptide exhibits toxicity. Considering the varied levels of biological activities, the toxicity peptides should be further classified into multi-functional peptides. RESULTS: This study introduces a two-level predictor, ToxPre-2L, developed using the multi-view tensor learning and latent semantic learning framework. The proposed method utilized multi-label learning with feature induced labels to avoid the redundancy of information from each view. Then the multi-view tensor learning was employed to establish the latent semantic information among different views, while low-rank constraint learning was leveraged to exploit the correlation information among multi-labels. Finally, we constructed an updated toxicity peptide benchmark dataset to assess the effectiveness of the proposed method. Experimental results demonstrated that ToxPre-2L achieves a better performance than alternative computational methods in the prediction of toxicity peptides and their multi-functional types. AVAILABILITY AND IMPLEMENTATION: The source code and data of ToxPre-2L can be accessed at http://bliulab.net/ToxPre-2L. Ke Yan 0003, Shutao Chen, Bin Liu 0014, Hao Wu 0066 |
Bioinform. | 3 |
| 2025 | A Systemic Pipeline of Identifying lncRNA-Disease Associations to the Prognosis and Treatment of Hepatocellular CarcinomaabstractExploring disease mechanisms at the lncRNA level provides valuable guidance for disease prognosis and treatment. Recently, there has been a surge of interest in exploring disease mechanisms via computational methods to overcome the challenge of tremendous manpower and material resources in biological experiments. However, current computational methods suffer from two main limitations: simple data structures that do not consider the close association between multiple types of data, and the lack of a systematic pathogenesis analysis that identified disease-associated lncRNAs are not applied to the downstream disease prognosis and therapeutic analysis from the perspective of data analysis. In this end, we present a systemic pipeline including disease-associated lncRNAs identification and downstream pathogenesis analysis on how the predicted lncRNAs are involved in the disease prognosis and therapy. Due to the importance of identifying disease-associated lncRNAs and the weak interpretability of existing computational identification methods, we propose a novel approach named iLncDA-PT to identify disease-associated lncRNAs considering the interactions between various bio-entities outperforming the other state-of-the-art methods, and then we conduct a systematically subsequent analysis on prognosis and therapy for a specific disease, hepatocellular carcinoma (HCC), as an example. Finally, we reveal a significant association between immune checkpoint expression, tumor microenvironment, and drug treatment. Wenxiang Zhang, Ye Yuan 0001, Hang Wei 0005, Bin Liu 0014 |
IEEE Trans. Big Data | 5 |
| 2025 | MMFmiRLocEL: A Multi-Model Fusion and Ensemble Learning Approach for Identifying miRNA Subcellular Localization Using RNA Structure Language ModelabstractMiRNA subcellular localizations (MSLs) are essential for uncovering and understanding miRNA functions in various biological processes. Several computational methods have been proposed for measuring MSL. However, existing methods only rely on manually crafted features based on sequence without considering RNA 3D structure information, and most methods often rely on single-model approaches, which fail to capture the full complexity of biological systems, further hindering predictive accuracy and performance. In this study, we introduce a deep learning-based approach, MMFmiRLocEL, which integrates multi-model fusion and ensemble learning for MSL identification. To the best of our knowledge, MMFmiRLocEL is the first method to combine sequence, structure, and function three information for MSL prediction. Specifically, it employs RNA 3D structure generated by the predicted structural model to construct a structure-based approach for MSL prediction. It also develops a sequence-based prediction method using sequence features and convolutional neural networks, while constructing a function-based prediction method using miRNA-disease association networks and deep residual neural networks. Furthermore, a multi-model fusion approach, employing weighted ensemble strategies, integrates sequence, structure, and function models to enhance the robustness and accuracy of MSL identification. Experimental results demonstrate that MMFmiRLocEL outperforms existing state-of-the-art methods, and then ablation analysis confirmed the significant contribution of the multi-model fusion mechanism to improve the prediction performance. Junxi Xie, Bin Liu 0014 |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Protein Language Pragmatic Analysis and Progressive Transfer Learning for Profiling Peptide-Protein InteractionsabstractProtein complex structural data are growing at an unprecedented pace, but its complexity and diversity pose significant challenges for protein function research. Although deep learning models have been widely used to capture the syntactic structure, word semantics, or semantic meanings of polypeptide and protein sequences, these models often overlook the complex contextual information of sequences. Here, we propose interpretable interaction deep learning (IIDL)-peptide-protein interaction (PepPI), a deep learning model designed to tackle these challenges using data-driven and interpretable pragmatic analysis to profile PepPIs. IIDL-PepPI constructs bidirectional attention modules to represent the contextual information of peptides and proteins, enabling pragmatic analysis. It then adopts a progressive transfer learning framework to simultaneously predict PepPIs and identify binding residues for specific interactions, providing a solution for multilevel in-depth profiling. We validate the performance and robustness of IIDL-PepPI in accurately predicting peptide-protein binary interactions and identifying binding residues compared with the state-of-the-art methods. We further demonstrate the capability of IIDL-PepPI in peptide virtual drug screening and binding affinity assessment, which is expected to advance artificial intelligence-based peptide drug discovery and protein function elucidation. Shutao Chen, Ke Yan 0003, Xuelong Li 0001, Bin Liu 0014 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | ProFun-SOM: Protein Function Prediction for Specific Ontology Based on Multiple Sequence Alignment ReconstructionabstractProtein function prediction is crucial for understanding species evolution, including viral mutations. Gene ontology (GO) is a standardized representation framework for describing protein functions with annotated terms. Each ontology is a specific functional category containing multiple child ontologies, and the relationships of parent and child ontologies create a directed acyclic graph. Protein functions are categorized using GO, which divides them into three main groups: cellular component ontology, molecular function ontology, and biological process ontology. Therefore, the GO annotation of protein is a hierarchical multilabel classification problem. This hierarchical relationship introduces complexities such as mixed ontology problem, leading to performance bottlenecks in existing computational methods due to label dependency and data sparsity. To overcome bottleneck issues brought by mixed ontology problem, we propose ProFun-SOM, an innovative multilabel classifier that utilizes multiple sequence alignments (MSAs) to accurately annotate gene ontologies. ProFun-SOM enhances the initial MSAs through a reconstruction process and integrates them into a deep learning architecture. It then predicts annotations within the cellular component, molecular function, biological process, and mixed ontologies. Our evaluation results on three datasets (CAFA3, SwissProt, and NetGO2) demonstrate that ProFun-SOM surpasses state-of-the-art methods. This study confirmed that utilizing MSAs of proteins can effectively overcome the two main bottlenecks issues, label dependency and data sparsity, thereby alleviating the root problem, mixed ontology. A freely accessible web server is available at http://bliulab.net/ ProFun-SOM/. Jiangyi Shao, Junjie Chen 0004, Bin Liu 0014 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | MVLncStack: A novel method for lncRNA subcellular localization prediction based on multi-view featuresabstractLncRNA performs different life activities in different subcellulars. For the subcellular multi-classification localization task of lncRNA, this paper proposed an ensemble learning prediction model named MVLncStack based on random forest and deep learning. Random forest in MVLncStack was employed to further analyze the sequence-level features of lncRNA, and selected the optimal feature subset with incremental feature selection. The deep learning model in MVLncStack was employed to learn the residue-level features of lncRNA with position encoding. RNA multi-view features provided the sequence information and structure information of lncRNA for MVLncStack, which can enhance the prediction performance of the model. Compared with the existing methods on the independent test set, the accuracy of multi classification prediction was improved by 5.2%, and other indicators were improved by about 1% - 2%. Through the performance evaluation and related experimental analysis, the importance of multi-view features of RNA was illustrated, and the prediction performance of the models can be effectively improved by selecting appropriate feature coding for RNA. Dongdong Jiang, Bin Liu 0014, Liang Yu 0002 |
BIBM | 4 |
| 2024 | DiSMVC: a multi-view graph collaborative learning framework for measuring disease similarityabstractMOTIVATION: Exploring potential associations between diseases can help in understanding pathological mechanisms of diseases and facilitating the discovery of candidate biomarkers and drug targets, thereby promoting disease diagnosis and treatment. Some computational methods have been proposed for measuring disease similarity. However, these methods describe diseases without considering their latent multi-molecule regulation and valuable supervision signal, resulting in limited biological interpretability and efficiency to capture association patterns. RESULTS: In this study, we propose a new computational method named DiSMVC. Different from existing predictors, DiSMVC designs a supervised graph collaborative framework to measure disease similarity. Multiple bio-entity associations related to genes and miRNAs are integrated via cross-view graph contrastive learning to extract informative disease representation, and then association pattern joint learning is implemented to compute disease similarity by incorporating phenotype-annotated disease associations. The experimental results show that DiSMVC can draw discriminative characteristics for disease pairs, and outperform other state-of-the-art methods. As a result, DiSMVC is a promising method for predicting disease associations with molecular interpretability. AVAILABILITY AND IMPLEMENTATION: Datasets and source codes are available at https://github.com/Biohang/DiSMVC. Hang Wei 0005, Lin Gao 0006, Shuai Wu 0001, Yina Jiang, Bin Liu 0014 |
Bioinform. | 5 |
| 2024 | TPpred-SC: multi-functional therapeutic peptide prediction based on multi-label supervised contrastive learning
Ke Yan 0003, Hongwu Lv, Jiangyi Shao, Shutao Chen, Bin Liu 0014 |
Sci. China Inf. Sci. | 5 |
| 2024 | Multiple types of disease-associated RNAs identification for disease prognosis and therapy using heterogeneous graph learning
Wenxiang Zhang, Hang Wei 0005, Hao Wu 0066, Bin Liu 0014 |
Sci. China Inf. Sci. | 5 |
| 2023 | DAmiRLocGNet: miRNA subcellular localization prediction by combining miRNA-disease associations and graph convolutional networksabstractMicroRNAs (miRNAs) are human post-transcriptional regulators in humans, which are involved in regulating various physiological processes by regulating the gene expression. The subcellular localization of miRNAs plays a crucial role in the discovery of their biological functions. Although several computational methods based on miRNA functional similarity networks have been presented to identify the subcellular localization of miRNAs, it remains difficult for these approaches to effectively extract well-referenced miRNA functional representations due to insufficient miRNA-disease association representation and disease semantic representation. Currently, there has been a significant amount of research on miRNA-disease associations, making it possible to address the issue of insufficient miRNA functional representation. In this work, a novel model is established, named DAmiRLocGNet, based on graph convolutional network (GCN) and autoencoder (AE) for identifying the subcellular localizations of miRNA. The DAmiRLocGNet constructs the features based on miRNA sequence information, miRNA-disease association information and disease semantic information. GCN is utilized to gather the information of neighboring nodes and capture the implicit information of network structures from miRNA-disease association information and disease semantic information. AE is employed to capture sequence semantics from sequence similarity networks. The evaluation demonstrates that the performance of DAmiRLocGNet is superior to other competing computational approaches, benefiting from implicit features captured by using GCNs. The DAmiRLocGNet has the potential to be applied to the identification of subcellular localization of other non-coding RNAs. Moreover, it can facilitate further investigation into the functional mechanisms underlying miRNA localization. The source code and datasets are accessed at http://bliulab.net/DAmiRLocGNet. Ke Yan 0003, Bin Liu 0014 |
Briefings Bioinform. | 3 |
| 2023 | LncRNA-disease association identification using graph auto-encoder and learning to rankabstractDiscovering the relationships between long non-coding RNAs (lncRNAs) and diseases is significant in the treatment, diagnosis and prevention of diseases. However, current identified lncRNA-disease associations are not enough because of the expensive and heavy workload of wet laboratory experiments. Therefore, it is greatly important to develop an efficient computational method for predicting potential lncRNA-disease associations. Previous methods showed that combining the prediction results of the lncRNA-disease associations predicted by different classification methods via Learning to Rank (LTR) algorithm can be effective for predicting potential lncRNA-disease associations. However, when the classification results are incorrect, the ranking results will inevitably be affected. We propose the GraLTR-LDA predictor based on biological knowledge graphs and ranking framework for predicting potential lncRNA-disease associations. Firstly, homogeneous graph and heterogeneous graph are constructed by integrating multi-source biological information. Then, GraLTR-LDA integrates graph auto-encoder and attention mechanism to extract embedded features from the constructed graphs. Finally, GraLTR-LDA incorporates the embedded features into the LTR via feature crossing statistical strategies to predict priority order of diseases associated with query lncRNAs. Experimental results demonstrate that GraLTR-LDA outperforms the other state-of-the-art predictors and can effectively detect potential lncRNA-disease associations. Availability and implementation: Datasets and source codes are available at http://bliulab.net/GraLTR-LDA. Wenxiang Zhang, Hao Wu 0066, Bin Liu 0014 |
Briefings Bioinform. | 4 |
| 2023 | PreHom-PCLM: protein remote homology detection by combing motifs and protein cubic language modelabstractProtein remote homology detection is essential for structure prediction, function prediction, disease mechanism understanding, etc. The remote homology relationship depends on multiple protein properties, such as structural information and local sequence patterns. Previous studies have shown the challenges for predicting remote homology relationship by protein features at sequence level (e.g. position-specific score matrix). Protein motifs have been used in structure and function analysis due to their unique sequence patterns and implied structural information. Therefore, designing a usable architecture to fuse multiple protein properties based on motifs is urgently needed to improve protein remote homology detection performance. To make full use of the characteristics of motifs, we employed the language model called the protein cubic language model (PCLM). It combines multiple properties by constructing a motif-based neural network. Based on the PCLM, we proposed a predictor called PreHom-PCLM by extracting and fusing multiple motif features for protein remote homology detection. PreHom-PCLM outperforms the other state-of-the-art methods on the test set and independent test set. Experimental results further prove the effectiveness of multiple features fused by PreHom-PCLM for remote homology detection. Furthermore, the protein features derived from the PreHom-PCLM show strong discriminative power for proteins from different structural classes in the high-dimensional space. Availability and Implementation: http://bliulab.net/PreHom-PCLM. Jiangyi Shao, Ke Yan 0003, Bin Liu 0014 |
Briefings Bioinform. | 4 |
| 2023 | CFAGO: cross-fusion of network and attributes based on attention mechanism for protein function predictionabstractMOTIVATION: Protein function annotation is fundamental to understanding biological mechanisms. The abundant genome-scale protein-protein interaction (PPI) networks, together with other protein biological attributes, provide rich information for annotating protein functions. As PPI networks and biological attributes describe protein functions from different perspectives, it is highly challenging to cross-fuse them for protein function prediction. Recently, several methods combine the PPI networks and protein attributes via the graph neural networks (GNNs). However, GNNs may inherit or even magnify the bias caused by noisy edges in PPI networks. Besides, GNNs with stacking of many layers may cause the over-smoothing problem of node representations. RESULTS: We develop a novel protein function prediction method, CFAGO, to integrate single-species PPI networks and protein biological attributes via a multi-head attention mechanism. CFAGO is first pre-trained with an encoder-decoder architecture to capture the universal protein representation of the two sources. It is then fine-tuned to learn more effective protein representations for protein function prediction. Benchmark experiments on human and mouse datasets show CFAGO outperforms state-of-the-art single-species network-based methods by at least 7.59%, 6.90%, 11.68% in terms of m-AUPR, M-AUPR, and Fmax, respectively, demonstrating cross-fusion by multi-head attention mechanism can greatly improve the protein function prediction. We further evaluate the quality of captured protein representations in terms of Davies Bouldin Score, whose results show that cross-fused protein representations by multi-head attention mechanism are at least 2.7% better than that of original and concatenated representations. We believe CFAGO is an effective tool for protein function prediction. AVAILABILITY AND IMPLEMENTATION: The source code of CFAGO and experiments data are available at: http://bliulab.net/CFAGO/. Zhourun Wu, Mingyue Guo 0001, Xiaopeng Jin, Junjie Chen 0004, Bin Liu 0014 |
Bioinform. | 5 |
| 2023 | PreTP-2L: identification of therapeutic peptides and their types using two-layer ensemble learning frameworkabstractMOTIVATION: Therapeutic peptides play an important role in immune regulation. Recently various therapeutic peptides have been used in the field of medical research, and have great potential in the design of therapeutic schedules. Therefore, it is essential to utilize the computational methods to predict the therapeutic peptides. However, the therapeutic peptides cannot be accurately predicted by the existing predictors. Furthermore, chaotic datasets are also an important obstacle of the development of this important field. Therefore, it is still challenging to develop a multi-classification model for identification of therapeutic peptides and their types. RESULTS: In this work, we constructed a general therapeutic peptide dataset. An ensemble-learning method named PreTP-2L was developed for predicting various therapeutic peptide types. PreTP-2L consists of two layers. The first layer predicts whether a peptide sequence belongs to therapeutic peptide, and the second layer predicts if a therapeutic peptide belongs to a particular species. AVAILABILITY AND IMPLEMENTATION: A user-friendly webserver PreTP-2L can be accessed at http://bliulab.net/PreTP-2L. Ke Yan 0003, Bin Liu 0014 |
Bioinform. | 3 |
| 2023 | sAMPpred-GAT: prediction of antimicrobial peptide by graph attention network and predicted peptide structureabstractMOTIVATION: Antimicrobial peptides (AMPs) are essential components of therapeutic peptides for innate immunity. Researchers have developed several computational methods to predict the potential AMPs from many candidate peptides. With the development of artificial intelligent techniques, the protein structures can be accurately predicted, which are useful for protein sequence and function analysis. Unfortunately, the predicted peptide structure information has not been applied to the field of AMP prediction so as to improve the predictive performance. RESULTS: In this study, we proposed a computational predictor called sAMPpred-GAT for AMP identification. To the best of our knowledge, sAMPpred-GAT is the first approach based on the predicted peptide structures for AMP prediction. The sAMPpred-GAT predictor constructs the graphs based on the predicted peptide structures, sequence information and evolutionary information. The Graph Attention Network (GAT) is then performed on the graphs to learn the discriminative features. Finally, the full connection networks are utilized as the output module to predict whether the peptides are AMP or not. Experimental results show that sAMPpred-GAT outperforms the other state-of-the-art methods in terms of AUC, and achieves better or highly comparable performance in terms of the other metrics on the eight independent test datasets, demonstrating that the predicted peptide structure information is important for AMP prediction. AVAILABILITY AND IMPLEMENTATION: A user-friendly webserver of sAMPpred-GAT can be accessed at http://bliulab.net/sAMPpred-GAT and the source code is available at https://github.com/HongWuL/sAMPpred-GAT/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ke Yan 0003, Hongwu Lv, Bin Liu 0014 |
Bioinform. | 5 |
| 2023 | iPiDA-SWGCN: Identification of piRNA-disease associations based on Supplementarily Weighted Graph Convolutional NetworkabstractAccurately identifying potential piRNA-disease associations is of great importance in uncovering the pathogenesis of diseases. Recently, several machine-learning-based methods have been proposed for piRNA-disease association detection. However, they are suffering from the high sparsity of piRNA-disease association network and the Boolean representation of piRNA-disease associations ignoring the confidence coefficients. In this study, we propose a supplementarily weighted strategy to solve these disadvantages. Combined with Graph Convolutional Networks (GCNs), a novel predictor called iPiDA-SWGCN is proposed for piRNA-disease association prediction. There are three main contributions of iPiDA-SWGCN: (i) Potential piRNA-disease associations are preliminarily supplemented in the sparse piRNA-disease network by integrating various basic predictors to enrich network structure information. (ii) The original Boolean piRNA-disease associations are assigned with different relevance confidence to learn node representations from neighbour nodes in varying degrees. (iii) The experimental results show that iPiDA-SWGCN achieves the best performance compared with the other state-of-the-art methods, and can predict new piRNA-disease associations. Jialu Hou, Hang Wei 0005, Bin Liu 0014 |
PLoS Comput. Biol. | 3 |
| 2023 | iDRBP-EL: Identifying DNA- and RNA- Binding Proteins Based on Hierarchical Ensemble LearningabstractIdentification of DNA-binding proteins (DBPs) and RNA-binding proteins (RBPs) from the primary sequences is essential for further exploring protein-nucleic acid interactions. Previous studies have shown that machine-learning-based methods can efficiently identify DBPs or RBPs. However, the information used in these methods is slightly unitary, and most of them only can predict DBPs or RBPs. In this study, we proposed a computational predictor iDRBP-EL to identify DNA- and RNA- binding proteins, and introduced hierarchical ensemble learning to integrate three level information. The method can integrate the information of different features, machine learning algorithms and data into one multi-label model. The ablation experiment showed that the fusion of different information can improve the prediction performance and overcome the cross-prediction problem. Experimental results on the independent datasets showed that iDRBP-EL outperformed all the other competing methods. Moreover, we established a user-friendly webserver iDRBP-EL (http://bliulab.net/iDRBP-EL), which can predict both DBPs and RBPs only based on protein sequences. Ning Wang 0054, Jun Zhang 0078, Bin Liu 0014 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | PreTP-Stack: Prediction of Therapeutic Peptides Based on the Stacked Ensemble LearingabstractTherapeutic peptide prediction is critical for drug development and therapeutic therapy. Researchers have developed several computational methods to identify different therapeutic peptide types. However, most computational methods focus on identifying the specific type of therapeutic peptides and fail to accurately predict all types of therapeutic peptides. Moreover, it is still challenging to utilize different properties features to predict the therapeutic peptides. In this study, a novel stacking framework PreTP-Stack is proposed for predicting different types of therapeutic peptide. PreTP-Stack is constructed based on ten different features and four predictors (Random Forest, Linear Discriminant Analysis, XGBoost and Support Vector Machine). Then the proposed method constructs an auto-weighted multi-view learning model as a final meta-classifier to enhance the performance of the basic models. Experimental results showed that the proposed method achieved better or highly comparable performance with the state-of-the-art methods for predicting eight types of therapeutic peptides A user-friendly web-server predictor is available at http://bliulab.net/PreTP-Stack. Ke Yan 0003, Hongwu Lv, Jie Wen 0001, Yong Xu 0001, Bin Liu 0014 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2023 | iSnoDi-MDRF: Identifying snoRNA-Disease Associations Based on Multiple Biological Data by Ranking FrameworkabstractAccumulating evidence indicates that the dysregulation of small nucleolar RNAs (snoRNAs) is relevant with diseases. Identifying snoRNA-disease associations by computational methods is desired for biologists, which can save considerable costs and time compared biological experiments. However, it still faces some challenges as followings: (i) Many snoRNAs are detected in recent years, but only a few snoRNAs have been proved to be associated with diseases; (ii) Computational predictors trained with only a few known snoRNA-disease associations fail to accurately identify the snoRNA-disease associations. In this study, we propose a ranking framework, called iSnoDi-MDRF, to identify potential snoRNA-disease associations based on multiple biological data, which has the following highlights: (i) iSnoDi-MDRF integrates ranking framework, which is not only able to identify potential associations between known snoRNAs and diseases, but also can identify diseases associated with new snoRNAs. (ii) Known gene-disease associations are employed to help train a mature model for predicting snoRNA-disease association. Experimental results illustrate that iSnoDi-MDRF is very suitable for identifying potential snoRNA-disease associations. The web server of iSnoDi-MDRF predictor is freely available at http://bliulab.net/iSnoDi-MDRF/. Wenxiang Zhang, Bin Liu 0014 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | Correction to: PHR-search: a search framework for protein remote homology detection based on the predicted protein hierarchical relationships
Xiaopeng Jin, Xiaoling Luo 0001, Bin Liu 0014 |
Briefings Bioinform. | 3 |
| 2022 | PHR-search: a search framework for protein remote homology detection based on the predicted protein hierarchical relationshipsabstractProtein remote homology detection is one of the most fundamental research tool for protein structure and function prediction. Most search methods for protein remote homology detection are evaluated based on the Structural Classification of Proteins-extended (SCOPe) benchmark, but the diverse hierarchical structure relationships between the query protein and candidate proteins are ignored by these methods. In order to further improve the predictive performance for protein remote homology detection, a search framework based on the predicted protein hierarchical relationships (PHR-search) is proposed. In the PHR-search framework, the superfamily level prediction information is obtained by extracting the local and global features of the Hidden Markov Model (HMM) profile through a convolution neural network and it is converted to the fold level and class level prediction information according to the hierarchical relationships of SCOPe. Based on these predicted protein hierarchical relationships, filtering strategy and re-ranking strategy are used to construct the two-level search of PHR-search. Experimental results show that the PHR-search framework achieves the state-of-the-art performance by employing five basic search methods, including HHblits, JackHMMER, PSI-BLAST, DELTA-BLAST and PSI-BLASTexB. Furthermore, the web server of PHR-search is established, which can be accessed at http://bliulab.net/PHR-search. Xiaopeng Jin, Xiaoling Luo 0001, Bin Liu 0014 |
Briefings Bioinform. | 3 |
| 2022 | iDRNA-ITF: identifying DNA- and RNA-binding residues in proteins based on induction and transfer frameworkabstractProtein-DNA and protein-RNA interactions are involved in many biological activities. In the post-genome era, accurate identification of DNA- and RNA-binding residues in protein sequences is of great significance for studying protein functions and promoting new drug design and development. Therefore, some sequence-based computational methods have been proposed for identifying DNA- and RNA-binding residues. However, they failed to fully utilize the functional properties of residues, leading to limited prediction performance. In this paper, a sequence-based method iDRNA-ITF was proposed to incorporate the functional properties in residue representation by using an induction and transfer framework. The properties of nucleic acid-binding residues were induced by the nucleic acid-binding residue feature extraction network, and then transferred into the feature integration modules of the DNA-binding residue prediction network and the RNA-binding residue prediction network for the final prediction. Experimental results on four test sets demonstrate that iDRNA-ITF achieves the state-of-the-art performance, outperforming the other existing sequence-based methods. The webserver of iDRNA-ITF is freely available at http://bliulab.net/iDRNA-ITF. Ning Wang 0054, Ke Yan 0003, Jun Zhang 0078, Bin Liu 0014 |
Briefings Bioinform. | 4 |
| 2022 | idenMD-NRF: a ranking framework for miRNA-disease association identificationabstractIdentifying miRNA-disease associations is an important task for revealing pathogenic mechanism of complicated diseases. Different computational methods have been proposed. Although these methods obtained encouraging performance for detecting missing associations between known miRNAs and diseases, how to accurately predict associated diseases for new miRNAs is still a difficult task. In this regard, a ranking framework named idenMD-NRF is proposed for miRNA-disease association identification. idenMD-NRF treats the miRNA-disease association identification as an information retrieval task. Given a novel query miRNA, idenMD-NRF employs Learning to Rank algorithm to rank associated diseases based on high-level association features and various predictors. The experimental results on two independent test datasets indicate that idenMD-NRF is superior to other compared predictors. A user-friendly web server of idenMD-NRF predictor is freely available at http://bliulab.net/idenMD-NRF/. Wenxiang Zhang, Hang Wei 0005, Bin Liu 0014 |
Briefings Bioinform. | 3 |
| 2022 | DeepIDP-2L: protein intrinsically disordered region prediction by combining convolutional attention network and hierarchical attention networkabstractMOTIVATION: Intrinsically disordered regions (IDRs) are widely distributed in proteins. Accurate prediction of IDRs is critical for the protein structure and function analysis. The IDRs are divided into long disordered regions (LDRs) and short disordered regions (SDRs) according to their lengths. Previous studies have shown that LDRs and SDRs have different proprieties. However, the existing computational methods fail to extract different features for LDRs and SDRs separately. As a result, they achieve unstable performance on datasets with different ratios of LDRs and SDRs. RESULTS: In this study, a two-layer predictor was proposed called DeepIDP-2L. In the first layer, two kinds of attention-based models are used to extract different features for LDRs and SDRs, respectively. The hierarchical attention network is used to capture the distribution pattern features of LDRs, and convolutional attention network is used to capture the local correlation features of SDRs. The second layer of DeepIDP-2L maps the feature extracted in the first layer into a new feature space. Convolutional network and bidirectional long short term memory are used to capture the local and long-range information for predicting both SDRs and LDRs. Experimental results show that DeepIDP-2L can achieve more stable performance than other exiting predictors on independent test sets with different ratios of SDRs and LDRs. AVAILABILITY AND IMPLEMENTATION: For the convenience of most experimental scientists, a user-friendly and publicly accessible web-server for the new predictor has been established at http://bliulab.net/DeepIDP-2L/. It is anticipated that DeepIDP-2L will become a very useful tool for identification of intrinsically disordered regions. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yi-Jun Tang, Yihe Pang, Bin Liu 0014 |
Bioinform. | 3 |
| 2022 | TPpred-ATMV: therapeutic peptide prediction by adaptive multi-view tensor learning modelabstractMOTIVATION: Therapeutic peptide prediction is important for the discovery of efficient therapeutic peptides and drug development. Researchers have developed several computational methods to identify different therapeutic peptide types. However, these computational methods focus on identifying some specific types of therapeutic peptides, failing to predict the comprehensive types of therapeutic peptides. Moreover, it is still challenging to utilize different properties to predict the therapeutic peptides. RESULTS: In this study, an adaptive multi-view based on the tensor learning framework TPpred-ATMV is proposed for predicting different types of therapeutic peptides. TPpred-ATMV constructs the class and probability information based on various sequence features. We constructed the latent subspace among the multi-view features and constructed an auto-weighted multi-view tensor learning model to utilize the high correlation based on the multi-view features. Experimental results showed that the TPpred-ATMV is better than or highly comparable with the other state-of-the-art methods for predicting eight types of therapeutic peptides. AVAILABILITY AND IMPLEMENTATION: The code of TPpred-ATMV is accessed at: https://github.com/cokeyk/TPpred-ATMV. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ke Yan 0003, Hongwu Lv, Yongyong Chen, Hao Wu 0066, Bin Liu 0014 |
Bioinform. | 6 |
| 2022 | PreRBP-TL: prediction of species-specific RNA-binding proteins based on transfer learningabstractMOTIVATION: RNA-binding proteins (RBPs) play crucial roles in post-transcriptional regulation. Accurate identification of RBPs helps to understand gene expression, regulation, etc. In recent years, some computational methods were proposed to identify RBPs. However, these methods fail to accurately identify RBPs from some specific species with limited data, such as bacteria. RESULTS: In this study, we introduce a computational method called PreRBP-TL for identifying species-specific RBPs based on transfer learning. The weights of the prediction model were initialized by pretraining with the large general RBP dataset and then fine-tuned with the small species-specific RPB dataset by using transfer learning. The experimental results show that the PreRBP-TL achieves better performance for identifying the species-specific RBPs from Human, Arabidopsis, Escherichia coli and Salmonella, outperforming eight state-of-the-art computational methods. It is anticipated PreRBP-TL will become a useful method for identifying RBPs. AVAILABILITY AND IMPLEMENTATION: For the convenience of researchers to identify RBPs, the web server of PreRBP-TL was established, freely available at http://bliulab.net/PreRBP-TL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jun Zhang 0078, Ke Yan 0003, Qingcai Chen, Bin Liu 0014 |
Bioinform. | 4 |
| 2022 | iPiDA-GCN: Identification of piRNA-disease associations based on Graph Convolutional NetworkabstractMOTIVATION: Piwi-interacting RNAs (piRNAs) play a critical role in the progression of various diseases. Accurately identifying the associations between piRNAs and diseases is important for diagnosing and prognosticating diseases. Although some computational methods have been proposed to detect piRNA-disease associations, it is challenging for these methods to effectively capture nonlinear and complex relationships between piRNAs and diseases because of the limited training data and insufficient association representation. RESULTS: With the growth of piRNA-disease association data, it is possible to design a more complex machine learning method to solve this problem. In this study, we propose a computational method called iPiDA-GCN for piRNA-disease association identification based on graph convolutional networks (GCNs). The iPiDA-GCN predictor constructs the graphs based on piRNA sequence information, disease semantic information and known piRNA-disease associations. Two GCNs (Asso-GCN and Sim-GCN) are used to extract the features of both piRNAs and diseases by capturing the association patterns from piRNA-disease interaction network and two similarity networks. GCNs can capture complex network structure information from these networks, and learn discriminative features. Finally, the full connection networks and inner production are utilized as the output module to predict piRNA-disease association scores. Experimental results demonstrate that iPiDA-GCN achieves better performance than the other state-of-the-art methods, benefitted from the discriminative features extracted by Asso-GCN and Sim-GCN. The iPiDA-GCN predictor is able to detect new piRNA-disease associations to reveal the potential pathogenesis at the RNA level. The data and source code are available at http://bliulab.net/iPiDA-GCN/. Jialu Hou, Hang Wei 0005, Bin Liu 0014 |
PLoS Comput. Biol. | 3 |
| 2022 | SelfAT-Fold: Protein Fold Recognition Based on Residue-Based and Motif-Based Self-Attention NetworksabstractThe protein fold recognition is a fundamental and crucial step of tertiary structure determination. In this regard, several computational predictors have been proposed. Recently, the predictive performance has been obviously improved by the fold-specific features generated by deep learning techniques. However, these methods failed to measure the global associations among residues or motifs along the protein sequences. Furthermore, these deep learning techniques are often treated as black boxes without interpretability. Inspired by the similarities between protein sequences and natural language sentences, we applied the self-attention mechanism derived from natural language processing (NLP) field to protein fold recognition. The motif-based self-attention network (MSAN) and the residue-based self-attention network (RSAN) were constructed based on a training set to capture the global associations among the structure motifs and residues along the protein sequences, respectively. The fold-specific attention features trained and generated from the training set were then combined with Support Vector Machines (SVMs) to predict the samples in the widely used LE benchmark dataset, which is fully independent from the training set. Experimental results showed that the proposed two SelfAT-Fold predictors outperformed 34 existing state-of-the-art computational predictors. The two SelfAT-Fold predictors were further tested on an independent dataset SCOP_TEST, and they can achieve stable performance. Furthermore, the fold-specific attention features can be used to analyse the characteristics of protein folds. The trained models and data of SelfAT-Fold can be downloaded from http://bliulab.net/selfAT_fold/. Yihe Pang, Bin Liu 0014 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | ProtRe-CN: Protein Remote Homology Detection by Combining Classification Methods and Network Methods via Learning to RankabstractProtein remote homology detection is one of fundamental research tasks for downstream analysis (i.e., protein structure and function prediction). Many advanced methods are proposed from different views with complementary detection ability, such as the classification method, the network method, and the ranking method. A framework integrating these heterogeneous methods is urgently desired to reduce the false positive rate and predictive bias. We propose a novel ranking method called ProtRe-CN by fusing the classification methods and network methods via Learning to Rank. Experimental results on the benchmark dataset and the independent dataset show that ProtRe-CN outperforms other existing state-of-the-art predictors. ProtRe-CN improves the detective performance via correcting the false positives in the ranking list by combining the heterogeneous methods. The web server of ProtRe-CN can be accessed at http://bliulab.net/ProtRe-CN. Jiangyi Shao, Junjie Chen 0004, Bin Liu 0014 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | IDRBP-PPCT: Identifying Nucleic Acid-Binding Proteins Based on Position-Specific Score Matrix and Position-Specific Frequency Matrix Cross TransformationabstractDNA-binding proteins (DBPs) and RNA-binding proteins (RBPs) are two important nucleic acid-binding proteins (NABPs), which play important roles in biological processes such as replication, translation and transcription of genetic material. Some proteins (DRBPs) bind to both DNA and RNA, also play a key role in gene expression. Identification of DBPs, RBPs and DRBPs is important to study protein-nucleic acid interactions. Computational methods are increasingly being proposed to automatically identify DNA- or RNA-binding proteins based only on protein sequences. One challenge is to design an effective protein representation method to convert protein sequences into fixed-dimension feature vectors. In this study, we proposed a novel protein representation method called Position-Specific Scoring Matrix (PSSM) and Position-Specific Frequency Matrix (PSFM) Cross Transformation (PPCT) to represent protein sequences. This method contains the evolutionary information in PSSM and PSFM, and their correlations. A new computational predictor called IDRBP-PPCT was proposed by combining PPCT and the two-layer framework based on the random forest algorithm to identify DBPs, RBPs and DRBPs. The experimental results on the independent dataset and the tomato genome proved the effectiveness of the proposed method. A user-friendly web-server of IDRBP-PPCT was constructed, which is freely available at http://bliulab.net/IDRBP-PPCT. Ning Wang 0054, Jun Zhang 0078, Bin Liu 0014 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | PreTP-EL: prediction of therapeutic peptides based on ensemble learningabstractTherapeutic peptides are important for understanding the correlation between peptides and their therapeutic diagnostic potential. The therapeutic peptides can be further divided into different types based on therapeutic function sharing different characteristics. Although some computational approaches have been proposed to predict different types of therapeutic peptides, they failed to accurately predict all types of therapeutic peptides. In this study, a predictor called PreTP-EL has been proposed via employing the ensemble learning approach to fuse the different features and machine learning techniques in order to capture the different characteristics of various therapeutic peptides. Experimental results showed that PreTP-EL outperformed other competing methods. Availability and implementation: A user-friendly web-server of PreTP-EL predictor is available at http://bliulab.net/PreTP-EL. Ke Yan 0003, Hongwu Lv, Bin Liu 0014 |
Briefings Bioinform. | 4 |
| 2021 | PL-search: a profile-link-based search method for protein remote homology detectionabstractProtein remote homology detection is a fundamental and important task for protein structure and function analysis. Several search methods have been proposed to improve the detection performance of the remote homologues and the accuracy of ranking lists. The position-specific scoring matrix (PSSM) profile and hidden Markov model (HMM) profile can contribute to improving the performance of the state-of-the-art search methods. In this paper, we improved the profile-link (PL) information for constructing PSSM or HMM profiles, and proposed a PL-based search method (PL-search). In PL-search, more robust PLs are constructed through the double-link and iterative extending strategies, and an accurate similarity score of sequence pairs is calculated from the two-level Jaccard distance for remote homologues. We tested our method on two widely used benchmark datasets. Our results show that whether HHblits, JackHMMER or position-specific iterated-BLAST is used, PL-search obviously improves the search performance in terms of ranking quality as well as the number of detected remote homologues. For ease of use of PL-search, both its stand-alone tool and the web server are constructed, which can be accessed at http://bliulab.net/PL-search/. Xiaopeng Jin, Qing Liao 0001, Bin Liu 0014 |
Briefings Bioinform. | 3 |
| 2021 | FoldRec-C2C: protein fold recognition by combining cluster-to-cluster model and protein similarity networkabstractAs a key for studying the protein structures, protein fold recognition is playing an important role in predicting the protein structures associated with COVID-19 and other important structures. However, the existing computational predictors only focus on the protein pairwise similarity or the similarity between two groups of proteins from 2-folds. However, the homology relationship among proteins is in a hierarchical structure. The global protein similarity network will contribute to the performance improvement. In this study, we proposed a predictor called FoldRec-C2C to globally incorporate the interactions among proteins into the prediction. For the FoldRec-C2C predictor, protein fold recognition problem is treated as an information retrieval task in nature language processing. The initial ranking results were generated by a surprised ranking algorithm Learning to Rank, and then three re-ranking algorithms were performed on the ranking lists to adjust the results globally based on the protein similarity network, including seq-to-seq model, seq-to-cluster model and cluster-to-cluster model (C2C). When tested on a widely used and rigorous benchmark dataset LINDAHL dataset, FoldRec-C2C outperforms other 34 state-of-the-art methods in this field. The source code and data of FoldRec-C2C can be downloaded from http://bliulab.net/FoldRec-C2C/download. Jiangyi Shao, Ke Yan 0003, Bin Liu 0014 |
Briefings Bioinform. | 3 |
| 2021 | iPiDi-PUL: identifying Piwi-interacting RNA-disease associations based on positive unlabeled learningabstractAccumulated researches have revealed that Piwi-interacting RNAs (piRNAs) are regulating the development of germ and stem cells, and they are closely associated with the progression of many diseases. As the number of the detected piRNAs is increasing rapidly, it is important to computationally identify new piRNA-disease associations with low cost and provide candidate piRNA targets for disease treatment. However, it is a challenging problem to learn effective association patterns from the positive piRNA-disease associations and the large amount of unknown piRNA-disease pairs. In this study, we proposed a computational predictor called iPiDi-PUL to identify the piRNA-disease associations. iPiDi-PUL extracted the features of piRNA-disease associations from three biological data sources, including piRNA sequence information, disease semantic terms and the available piRNA-disease association network. Principal component analysis (PCA) was then performed on these features to extract the key features. The training datasets were constructed based on known positive associations and the negative associations selected from the unknown pairs. Various random forest classifiers trained with these different training sets were merged to give the predictive results via an ensemble learning approach. Finally, the web server of iPiDi-PUL was established at http://bliulab.net/iPiDi-PUL to help the researchers to explore the associated diseases for newly discovered piRNAs. Hang Wei 0005, Yong Xu 0001, Bin Liu 0014 |
Briefings Bioinform. | 3 |
| 2021 | idenPC-CAP: Identify protein complexes from weighted RNA-protein heterogeneous interaction networks using co-assemble partner relationabstractProtein complexes play important roles in most cellular processes. The available genome-wide protein-protein interaction (PPI) data make it possible for computational methods identifying protein complexes from PPI networks. However, PPI datasets usually contain a large ratio of false positive noise. Moreover, different types of biomolecules in a living cell cooperate to form a union interaction network. Because previous computational methods focus only on PPIs ignoring other types of biomolecule interactions, their predicted protein complexes often contain many false positive proteins. In this study, we develop a novel computational method idenPC-CAP to identify protein complexes from the RNA-protein heterogeneous interaction network consisting of RNA-RNA interactions, RNA-protein interactions and PPIs. By considering interactions among proteins and RNAs, the new method reduces the ratio of false positive proteins in predicted protein complexes. The experimental results demonstrate that idenPC-CAP outperforms the other state-of-the-art methods in this field. Zhourun Wu, Qing Liao 0001, Shixi Fan, Bin Liu 0014 |
Briefings Bioinform. | 4 |
| 2021 | idenPC-MIIP: identify protein complexes from weighted PPI networks using mutual important interacting partner relationabstractProtein complexes are key units for studying a cell system. During the past decades, the genome-scale protein-protein interaction (PPI) data have been determined by high-throughput approaches, which enables the identification of protein complexes from PPI networks. However, the high-throughput approaches often produce considerable fraction of false positive and negative samples. In this study, we propose the mutual important interacting partner relation to reflect the co-complex relationship of two proteins based on their interaction neighborhoods. In addition, a new algorithm called idenPC-MIIP is developed to identify protein complexes from weighted PPI networks. The experimental results on two widely used datasets show that idenPC-MIIP outperforms 17 state-of-the-art methods, especially for identification of small protein complexes with only two or three proteins. Zhourun Wu, Qing Liao 0001, Bin Liu 0014 |
Briefings Bioinform. | 3 |
| 2021 | NCBRPred: predicting nucleic acid binding residues in proteins based on multilabel learningabstractThe interactions between proteins and nucleic acid sequences play many important roles in gene expression and some cellular activities. Accurate prediction of the nucleic acid binding residues in proteins will facilitate the research of the protein functions, gene expression, drug design, etc. In this regard, several computational methods have been proposed to predict the nucleic acid binding residues in proteins. However, these methods cannot satisfactorily measure the global interactions among the residues along protein. Furthermore, these methods are suffering cross-prediction problem, new strategies should be explored to solve this problem. In this study, a new computational method called NCBRPred was proposed to predict the nucleic acid binding residues based on the multilabel sequence labeling model. NCBRPred used the bidirectional Gated Recurrent Units (BiGRUs) to capture the global interactions among the residues, and treats this task as a multilabel learning task. Experimental results on three widely used benchmark datasets and an independent dataset showed that NCBRPred achieved higher predictive results with lower cross-prediction, outperforming 10 existing state-of-the-art predictors. The web-server and a stand-alone package of NCBRPred are freely available at http://bliulab.net/NCBRPred. It is anticipated that NCBRPred will become a very useful tool for identifying nucleic acid binding residues. Jun Zhang 0078, Qingcai Chen, Bin Liu 0014 |
Briefings Bioinform. | 3 |
| 2021 | S2L-PSIBLAST: a supervised two-layer search framework based on PSI-BLAST for protein remote homology detectionabstractMOTIVATION: Protein remote homology detection is a challenging task for the studies of protein evolutionary relationships. PSI-BLAST is an important and fundamental search method for detecting homology proteins. Although many improved versions of PSI-BLAST have been proposed, their performance is limited by the search processes of PSI-BLAST. RESULTS: For further improving the performance of PSI-BLAST for protein remote homology detection, a supervised two-layer search framework based on PSI-BLAST (S2L-PSIBLAST) is proposed. S2L-PSIBLAST consists of a two-level search: the first-level search provides high-quality search results by using SMI-BLAST framework and double-link strategy to filter the non-homology protein sequences, the second-level search detects more homology proteins by profile-link similarity, and more accurate ranking lists for those detected protein sequences are obtained by learning to rank strategy. Experimental results on the updated version of Structural Classification of Proteins-extended benchmark dataset show that S2L-PSIBLAST not only obviously improves the performance of PSI-BLAST, but also achieves better performance on two improved versions of PSI-BLAST: DELTA-BLAST and PSI-BLASTexB. AVAILABILITY AND IMPLEMENTATION: http://bliulab.net/S2L-PSIBLAST. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaopeng Jin, Qing Liao 0001, Bin Liu 0014 |
Bioinform. | 3 |
| 2021 | SMI-BLAST: a novel supervised search framework based on PSI-BLAST for protein remote homology detectionabstractMOTIVATION: As one of the most important and widely used mainstream iterative search tool for protein sequence search, an accurate Position-Specific Scoring Matrix (PSSM) is the key of PSI-BLAST. However, PSSMs containing non-homologous information obviously reduce the performance of PSI-BLAST for protein remote homology. RESULTS: To further study this problem, we summarize three types of Incorrectly Selected Homology (ISH) errors in PSSMs. A new search tool Supervised-Manner-based Iterative BLAST (SMI-BLAST) is proposed based on PSI-BLAST for solving these errors. SMI-BLAST obviously outperforms PSI-BLAST on the Structural Classification of Proteins-extended (SCOPe) dataset. Compared with PSI-BLAST on the ISH error subsets of SCOPe dataset, SMI-BLAST detects 1.6-2.87 folds more remote homologous sequences, and outperforms PSI-BLAST by 35.66% in terms of ROC1 scores. Furthermore, this framework is applied to JackHMMER, DELTA-BLAST and PSI-BLASTexB, and their performance is further improved. AVAILABILITY AND IMPLEMENTATION: User-friendly webservers for SMI-BLAST, JackHMMER, DELTA-BLAST and PSI-BLASTexB are established at http://bliulab.net/SMI-BLAST/, by which the users can easily get the results without the need to go through the mathematical details. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaopeng Jin, Qing Liao 0001, Hang Wei 0005, Jun Zhang 0078, Bin Liu 0014 |
Bioinform. | 5 |
| 2021 | iCircDA-LTR: identification of circRNA-disease associations based on Learning to RankabstractMOTIVATION: Due to the inherent stability and close relationship with the progression of diseases, circRNAs are serving as important biomarkers and drug targets. Efficient predictors for identifying circRNA-disease associations are highly required. The existing predictors consider circRNA-disease association prediction as a classification task or a recommendation problem, failing to capture the ranking information among the associations and detect the diseases associated with new circRNAs. However, more and more circRNAs are discovered. Identification of the diseases associated with these new circRNAs remains a challenging task. RESULTS: In this study, we proposed a new predictor called iCricDA-LTR for circRNA-disease association prediction. Different from any existing predictor, iCricDA-LTR employed a ranking framework to model the global ranking associations among the query circRNAs and the diseases. The Learning to Rank (LTR) algorithm was employed to rank the associations based on various predictors and features in a supervised manner. The experimental results on two independent test datasets showed that iCircDA-LTR outperformed the other competing methods, especially for predicting the diseases associated with new circRNAs. As a result, iCircDA-LTR is more suitable for the real-world applications. AVAILABILITY AND IMPLEMENTATION: For the convenience of researchers to detect new circRNA-disease associations. The web server of iCircDA-LTR was established and freely available at http://bliulab.net/iCircDA-LTR/. Hang Wei 0005, Yong Xu 0001, Bin Liu 0014 |
Bioinform. | 3 |
| 2021 | MLDH-Fold: Protein fold recognition based on multi-view low-rank modeling
Ke Yan 0003, Jie Wen 0001, Yong Xu 0001, Bin Liu 0014 |
Neurocomputing | 4 |
| 2021 | iLncRNAdis-FB: Identify lncRNA-Disease Associations by Fusing Biological Feature Blocks Through Deep Neural NetworkabstractIdentification of lncRNA-disease associations is not only important for exploring the disease mechanism, but will also facilitate the molecular targeting drug discovery. Fusing multiple biological information is able to generate a more comprehensive view of lncRNA-disease association feature. However, the existing fusion strategies in this field fail to remove the noisy and irrelevant information from each data source. As a result, their predictive performance is still too low to be applied to real world applications. In this regard, a novel computational predictor called iLncRNAdis-FB is proposed based on the Convolution Neural Network (CNN) to integrate different data sources by using the feature blocks in a supervised manner. The lncRNA similarity matrix and disease similarity matrix are constructed, based on which the three-dimensional feature blocks are generated. These feature blocks are then fed into CNN to train the model so as to predict unknown lncRNA-disease associations. Experimental results show that iLncRNAdis-FB achieves better performance compared with other state-of-the-art predictors. Furthermore, a web server of iLncRNAdis-FB has been established at http://bliulab.net/iLncRNAdis-FB/, by which users can submit lncRNA sequences to detect their potential associated diseases. Hang Wei 0005, Qing Liao 0001, Bin Liu 0014 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | Protein Fold Recognition by Combining Support Vector Machines and Pairwise Sequence Similarity ScoresabstractProtein fold recognition is one of the most essential steps for protein structure prediction, aiming to classify proteins into known protein folds. There are two main computational approaches: one is the template-based method based on the alignment scores between query-template protein pairs and the other is the machine learning method based on the feature representation and classifier. These two approaches have their own advantages and disadvantages. Can we combine these methods to establish more accurate predictors for protein fold recognition? In this study, we made an initial attempt and proposed two novel algorithms: TSVM-fold and ESVM-fold. TSVM-fold was based on the Support Vector Machines (SVMs), which utilizes a set of pairwise sequence similarity scores generated by three complementary template-based methods, including HHblits, SPARKS-X, and DeepFR. These scores measured the global relationships between query sequences and templates. The comprehensive features of the attributes of the sequences were fed into the SVMs for the prediction. Then the TSVM-fold was further combined with the HHblits algorithm so as to improve its generalization ability. The combined method is called ESVM-fold. Experimental results in two rigorous benchmark datasets (LE and YK datasets) showed that the proposed methods outperform some state-of-the-art methods, indicating that the TSVM-fold and ESVM-fold are efficient predictors for protein fold recognition. Ke Yan 0003, Jie Wen 0001, Jin-Xing Liu 0001, Yong Xu 0001, Bin Liu 0014 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2021 | Protein Fold Recognition Based on Auto-Weighted Multi-View Graph Embedding Learning ModelabstractProtein fold recognition is critical for studies of the protein structure prediction and drug design. Several methods have been proposed to obtain discriminative features from the protein sequences for fold recognition. However, the ensemble methods that combine the various features to improve predictive performance remain the challenge problems. In this study, we proposed two novel algorithms: AWMG and EMfold. AWMG used a novel predictor based on the multi-view learning framework for fold recognition. Each view was treated as the intermediate representation of the corresponding data source of proteins, including the evolutionary information and the retrieval information. AWMG calculated the auto-weight for each view respectively and constructed the latent subspace which contains the common information shared by different views. The marginalized constraint was employed to enlarge the margins between different folds, improving the predictive performance of AWMG. Furthermore, we proposed a novel ensemble method called EMfold, which combines two complementary methods AWMG and DeepSS. The later method was a template-based algorithm using the SPARKS-X and DeepFR programs. EMfold integrated the advantages of template-based assignment and machine learning classifier. Experimental results on the two widely datasets (LE and YK) showed that the proposed methods outperformed some state-of-the-art methods, indicating that AWMG and EMfold are useful tools for protein fold recognition. Ke Yan 0003, Jie Wen 0001, Yong Xu 0001, Bin Liu 0014 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2021 | DeepDRBP-2L: A New Genome Annotation Predictor for Identifying DNA-Binding Proteins and RNA-Binding Proteins Using Convolutional Neural Network and Long Short-Term MemoryabstractDNA-binding proteins (DBPs) and RNA-binding proteins (RBPs) are two kinds of crucial proteins, which are associated with various cellule activities and some important diseases. Accurate identification of DBPs and RBPs facilitate both theoretical research and real world application. Existing sequence-based DBP predictors can accurately identify DBPs but incorrectly predict many RBPs as DBPs, and vice versa, resulting in low prediction precision. Moreover, some proteins (DRBPs) interacting with both DNA and RNA play important roles in gene expression and cannot be identified by existing computational methods. In this study, a two-level predictor named DeepDRBP-2L was proposed by combining Convolutional Neural Network (CNN) and the Long Short-Term Memory (LSTM). It is the first computational method that is able to identify DBPs, RBPs and DRBPs. Rigorous cross-validations and independent tests showed that DeepDRBP-2L is able to overcome the shortcoming of the existing methods and can go one further step to identify DRBPs. Application of DeepDRBP-2L to tomato genome further demonstrated its performance. The webserver of DeepDRBP-2L is freely available at http://bliulab.net/DeepDRBP-2L. Jun Zhang 0078, Qingcai Chen, Bin Liu 0014 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2020 | HITS-PR-HHblits: protein remote homology detection by combining PageRank and Hyperlink-Induced Topic SearchabstractAs one of the most important fundamental problems in protein sequence analysis, protein remote homology detection is critical for both theoretical research (protein structure and function studies) and real world applications (drug design). Although several computational predictors have been proposed, their detection performance is still limited. In this study, we treat protein remote homology detection as a document retrieval task, where the proteins are considered as documents and its aim is to find the highly related documents with the query documents in a database. A protein similarity network was constructed based on the true labels of proteins in the database, and the query proteins were then connected into the network based on the similarity scores calculated by three ranking methods, including PSI-BLAST, Hmmer and HHblits. The PageRank algorithm and Hyperlink-Induced Topic Search (HITS) algorithm were respectively performed on this network to move the homologous proteins of query proteins to the neighbors of the query proteins in the network. Finally, PageRank and HITS algorithms were combined, and a predictor called HITS-PR-HHblits was proposed to further improve the predictive performance. Tested on the SCOP and SCOPe benchmark datasets, the experimental results showed that the proposed protocols outperformed other state-of-the-art methods. For the convenience of the most experimental scientists, a web server for HITS-PR-HHblits was established at http://bioinformatics.hitsz.edu.cn/HITS-PR-HHblits, by which the users can easily get the results without the need to go through the mathematical details. The HITS-PR-HHblits predictor is a protocol for protein remote homology detection using different sets of programs, which will become a very useful computational tool for proteome analysis. Bin Liu 0014, Shuangyan Jiang, Quan Zou 0001 |
Briefings Bioinform. | 1 |
| 2020 | DeepSVM-fold: protein fold recognition by combining support vector machines and pairwise sequence similarity scores generated by deep learning networksabstractProtein fold recognition is critical for studying the structures and functions of proteins. The existing protein fold recognition approaches failed to efficiently calculate the pairwise sequence similarity scores of the proteins in the same fold sharing low sequence similarities. Furthermore, the existing feature vectorization strategies are not able to measure the global relationships among proteins from different protein folds. In this article, we proposed a new computational predictor called DeepSVM-fold for protein fold recognition by introducing a new feature vector based on the pairwise sequence similarity scores calculated from the fold-specific features extracted by deep learning networks. The feature vectors are then fed into a support vector machine to construct the predictor. Experimental results on the benchmark dataset (LE) show that DeepSVM-fold obviously outperforms all the other competing methods. Bin Liu 0014, Chen-Chen Li, Ke Yan 0003 |
Briefings Bioinform. | 1 |
| 2020 | Fold-LTR-TCP: protein fold recognition based on triadic closure principleabstractAs an important task in protein structure and function studies, protein fold recognition has attracted more and more attention. The existing computational predictors in this field treat this task as a multi-classification problem, ignoring the relationship among proteins in the dataset. However, previous studies showed that their relationship is critical for protein homology analysis. In this study, the protein fold recognition is treated as an information retrieval task. The Learning to Rank model (LTR) was employed to retrieve the query protein against the template proteins to find the template proteins in the same fold with the query protein in a supervised manner. The triadic closure principle (TCP) was performed on the ranking list generated by the LTR to improve its accuracy by considering the relationship among the query protein and the template proteins in the ranking list. Finally, a predictor called Fold-LTR-TCP was proposed. The rigorous test on the LE benchmark dataset showed that the Fold-LTR-TCP predictor achieved an accuracy of 73.2%, outperforming all the other competing methods. Bin Liu 0014, Ke Yan 0003 |
Briefings Bioinform. | 1 |
| 2020 | iCircDA-MF: identification of circRNA-disease associations based on matrix factorizationabstractCircular RNAs (circRNAs) are a group of novel discovered non-coding RNAs with closed-loop structure, which play critical roles in various biological processes. Identifying associations between circRNAs and diseases is critical for exploring the complex disease mechanism and facilitating disease-targeted therapy. Although several computational predictors have been proposed, their performance is still limited. In this study, a novel computational method called iCircDA-MF is proposed. Because the circRNA-disease associations with experimental validation are very limited, the potential circRNA-disease associations are calculated based on the circRNA similarity and disease similarity extracted from the disease semantic information and the known associations of circRNA-gene, gene-disease and circRNA-disease. The circRNA-disease interaction profiles are then updated by the neighbour interaction profiles so as to correct the false negative associations. Finally, the matrix factorization is performed on the updated circRNA-disease interaction profiles to predict the circRNA-disease associations. The experimental results on a widely used benchmark dataset showed that iCircDA-MF outperforms other state-of-the-art predictors and can identify new circRNA-disease associations effectively. Hang Wei 0005, Bin Liu 0014 |
Briefings Bioinform. | 2 |
| 2020 | A comprehensive review and evaluation of computational methods for identifying protein complexes from protein-protein interaction networksabstractProtein complexes are the fundamental units for many cellular processes. Identifying protein complexes accurately is critical for understanding the functions and organizations of cells. With the increment of genome-scale protein-protein interaction (PPI) data for different species, various computational methods focus on identifying protein complexes from PPI networks. In this article, we give a comprehensive and updated review on the state-of-the-art computational methods in the field of protein complex identification, especially focusing on the newly developed approaches. The computational methods are organized into three categories, including cluster-quality-based methods, node-affinity-based methods and ensemble clustering methods. Furthermore, the advantages and disadvantages of different methods are discussed, and then, the performance of 17 state-of-the-art methods is evaluated on two widely used benchmark data sets. Finally, the bottleneck problems and their potential solutions in this important field are discussed. Zhourun Wu, Qing Liao 0001, Bin Liu 0014 |
Briefings Bioinform. | 3 |
| 2019 | BioSeq-Analysis: a platform for DNA, RNA and protein sequence analysis based on machine learning approachesabstractWith the avalanche of biological sequences generated in the post-genomic age, one of the most challenging problems is how to computationally analyze their structures and functions. Machine learning techniques are playing key roles in this field. Typically, predictors based on machine learning techniques contain three main steps: feature extraction, predictor construction and performance evaluation. Although several Web servers and stand-alone tools have been developed to facilitate the biological sequence analysis, they only focus on individual step. In this regard, in this study a powerful Web server called BioSeq-Analysis (http://bioinformatics.hitsz.edu.cn/BioSeq-Analysis/) has been proposed to automatically complete the three main steps for constructing a predictor. The user only needs to upload the benchmark data set. BioSeq-Analysis can generate the optimized predictor based on the benchmark data set, and the performance measures can be reported as well. Furthermore, to maximize user's convenience, its stand-alone program was also released, which can be downloaded from http://bioinformatics.hitsz.edu.cn/BioSeq-Analysis/download/, and can be directly run on Windows, Linux and UNIX. Applied to three sequence analysis tasks, experimental results showed that the predictors generated by BioSeq-Analysis even outperformed some state-of-the-art methods. It is anticipated that BioSeq-Analysis will become a useful tool for biological sequence analysis. Bin Liu 0014 |
Briefings Bioinform. | 1 |
| 2019 | A comprehensive review and comparison of existing computational methods for intrinsically disordered protein and region predictionabstractIntrinsically disordered proteins and regions are widely distributed in proteins, which are associated with many biological processes and diseases. Accurate prediction of intrinsically disordered proteins and regions is critical for both basic research (such as protein structure and function prediction) and practical applications (such as drug development). During the past decades, many computational approaches have been proposed, which have greatly facilitated the development of this important field. Therefore, a comprehensive and updated review is highly required. In this regard, we give a review on the computational methods for intrinsically disordered protein and region prediction, especially focusing on the recent development in this field. These computational approaches are divided into four categories based on their methodologies, including physicochemical-based method, machine-learning-based method, template-based method and meta method. Furthermore, their advantages and disadvantages are also discussed. The performance of 40 state-of-the-art predictors is directly compared on the target proteins in the task of disordered region prediction in the 10th Critical Assessment of protein Structure Prediction. A more comprehensive performance comparison of 45 different predictors is conducted based on seven widely used benchmark data sets. Finally, some open problems and perspectives are discussed. Xiaolong Wang 0001, Bin Liu 0014 |
Briefings Bioinform. | 3 |
| 2019 | Protein fold recognition based on multi-view modelingabstractMOTIVATION: Protein fold recognition has attracted increasing attention because it is critical for studies of the 3D structures of proteins and drug design. Researchers have been extensively studying this important task, and several features with high discriminative power have been proposed. However, the development of methods that efficiently combine these features to improve the predictive performance remains a challenging problem. RESULTS: In this study, we proposed two algorithms: MV-fold and MT-fold. MV-fold is a new computational predictor based on the multi-view learning model for fold recognition. Different features of proteins were treated as different views of proteins, including the evolutionary information, secondary structure information and physicochemical properties. These different views constituted the latent space. The ε-dragging technique was employed to enlarge the margins between different protein folds, improving the predictive performance of MV-fold. Then, MV-fold was combined with two template-based methods: HHblits and HMMER. The ensemble method is called MT-fold incorporating the advantages of both discriminative methods and template-based methods. Experimental results on five widely used benchmark datasets (DD, RDD, EDD, TG and LE) showed that the proposed methods outperformed some state-of-the-art methods in this field, indicating that MV-fold and MT-fold are useful computational tools for protein fold recognition and protein homology detection and would be efficient tools for protein sequence analysis. Finally, we constructed an update and rigorous benchmark dataset based on SCOPe (version 2.07) to fairly evaluate the performance of the proposed method, and our method achieved stable performance on this new dataset. This new benchmark dataset will become a widely used benchmark dataset to fairly evaluate the performance of different methods for fold recognition. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ke Yan 0003, Xiaozhao Fang, Yong Xu 0001, Bin Liu 0014 |
Bioinform. | 4 |
| 2019 | Video Highlight Detection via Region-Based Deep Ranking ModelabstractThe video highlight detection task is to localize key elements (moments of user’s major or special interest) in a video. Most of the existing highlight detection approaches extract features from the video segment as a whole without considering the difference of local features spatially. In spatial extent, not all regions are worth watching because some of them only contain the background of the environment without human or other moving objects, especially when there is lots of clutter in the background. To deal with this issue, we propose a novel region-based model which can automatically localize the key elements in a video without any extra supervised annotations. Specifically, the proposed model produces position-sensitive score maps for local regions in the spatial dimension of the video segment, and then aggregates all position-wise scores with position-pooling operation. The regions with higher response values will be extracted as key elements. Thus more effective features of the video segment are obtained to predict the highlight score. The proposed position-sensitive scheme can be easily integrated into an end-to-end fully convolutional network which aims to update parameters via stochastic gradient descent method in the backward propagation to improve the robustness of the model. Extensive experimental results on the YouTube and SumMe datasets demonstrate that the proposed approach achieves significant improvement over state-of-the-art methods. Yifan Jiao, Tianzhu Zhang 0001, Shucheng Huang, Bin Liu 0014, Changsheng Xu |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2019 | Protein Remote Homology Detection and Fold Recognition Based on Sequence-Order Frequency MatrixabstractProtein remote homology detection and fold recognition are two critical tasks for the studies of protein structures and functions. Currently, the profile-based methods achieve the state-of-the-art performance in these fields. However, the widely used sequence profiles, like position-specific frequency matrix (PSFM) and position-specific scoring matrix (PSSM), ignore the sequence-order effects along protein sequence. In this study, we have proposed a novel profile, called sequence-order frequency matrix (SOFM), to extract the sequence-order information of neighboring residues from multiple sequence alignment (MSA). Combined with two profile feature extraction approaches, top-n-grams and the Smith-Waterman algorithm, the SOFMs are applied to protein remote homology detection and fold recognition, and two predictors called SOFM-Top and SOFM-SW are proposed. Experimental results show that SOFM contains more information content than other profiles, and these two predictors outperform other state-of-the-art methods. It is anticipated that SOFM will become a very useful profile in the studies of protein structures and functions. Bin Liu 0014, Junjie Chen 0004, Mingyue Guo 0001, Xiaolong Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2019 | ProtDet-CCH: Protein Remote Homology Detection by Combining Long Short-Term Memory and Ranking MethodsabstractAs one of the most challenging tasks in sequence analysis, protein remote homology detection has been extensively studied. Methods based on discriminative models and ranking approaches have achieved the state-of-the-art performance, and these two kinds of methods are complementary. In this study, three LSTM models have been applied to construct the predictors for protein remote homology detection, including ULSTM, BLSTM, and CNN-BLSTM. They are able to automatically extract the local and global sequence order information. Combined with PSSMs, the CNN-BLSTM achieved the best performance among the three LSTM-based models. We named this method as CNN-BLSTM-PSSM. Finally, a new method called ProtDet-CCH was proposed by combining CNN-BLSTM-PSSM and a ranking method HHblits. Tested on a widely used SCOP benchmark dataset, ProtDet-CCH achieved an ROC score of 0.998, and an ROC50 score of 0.982, significantly outperforming other existing state-of-the-art methods. Experimental results on two updated SCOPe independent datasets showed that ProtDet-CCH can achieve stable performance. Furthermore, our method can provide useful insights for studying the features and motifs of protein families and superfamilies. It is anticipated that ProtDet-CCH will become a very useful tool for protein remote homology detection. Bin Liu 0014 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2018 | A comprehensive review and comparison of different computational methods for protein remote homology detectionabstractProtein remote homology detection is one of the most fundamental and central problems for the studies of protein structures and functions, aiming to detect the distantly evolutionary relationships among proteins via computational methods. During the past decades, many computational approaches have been proposed to solve this important task. These methods have made a substantial contribution to protein remote homology detection. Therefore, it is necessary to give a comprehensive review and comparison on these computational methods. In this article, we divide these computational approaches into three categories, including alignment methods, discriminative methods and ranking methods. Their advantages and disadvantages are discussed in a comprehensive perspective, and their performance is compared on widely used benchmark data sets. Finally, some open questions in this field are further explored and discussed. Junjie Chen 0004, Mingyue Guo 0001, Xiaolong Wang 0001, Bin Liu 0014 |
Briefings Bioinform. | 4 |
| 2018 | iEnhancer-EL: identifying enhancers and their strength with ensemble learning approachabstractMotivation: Identification of enhancers and their strength is important because they play a critical role in controlling gene expression. Although some bioinformatics tools were developed, they are limited in discriminating enhancers from non-enhancers only. Recently, a two-layer predictor called 'iEnhancer-2L' was developed that can be used to predict the enhancer's strength as well. However, its prediction quality needs further improvement to enhance the practical application value. Results: A new predictor called 'iEnhancer-EL' was proposed that contains two layer predictors: the first one (for identifying enhancers) is formed by fusing an array of six key individual classifiers, and the second one (for their strength) formed by fusing an array of ten key individual classifiers. All these key classifiers were selected from 171 elementary classifiers formed by SVM (Support Vector Machine) based on kmer, subsequence profile and PseKNC (Pseudo K-tuple Nucleotide Composition), respectively. Rigorous cross-validations have indicated that the proposed predictor is remarkably superior to the existing state-of-the-art one in this area. Availability and implementation: A web server for the iEnhancer-EL has been established at http://bioinformatics.hitsz.edu.cn/iEnhancer-EL/, by which users can easily get their desired results without the need to go through the mathematical details. Supplementary information: Supplementary data are available at Bioinformatics online. Bin Liu 0014, De-Shuang Huang, Kuo-Chen Chou |
Bioinform. | 1 |
| 2018 | iRO-3wPseKNC: identify DNA replication origins by three-window-based PseKNCabstractMotivation: DNA replication is the key of the genetic information transmission, and it is initiated from the replication origins. Identifying the replication origins is crucial for understanding the mechanism of DNA replication. Although several discriminative computational predictors were proposed to identify DNA replication origins of yeast species, they could only be used to identify very tiny parts (250 or 300 bp) of the replication origins. Besides, none of the existing predictors could successfully capture the 'GC asymmetry bias' of yeast species reported by experimental observations. Hence it would not be surprising why their power is so limited. To grasp the CG asymmetry feature and make the prediction able to cover the entire replication regions of yeast species, we develop a new predictor called 'iRO-3wPseKNC'. Results: Rigorous cross validations on the benchmark datasets from four yeast species (Saccharomyces cerevisiae, Schizosaccharomyces pombe, Kluyveromyces lactis and Pichia pastoris) have indicated that the proposed predictor is really very powerful for predicting the entire DNA duplication origins. Availability and implementation: The web-server for the iRO-3wPseKNC predictor is available at http://bioinformatics.hitsz.edu.cn/iRO-3wPseKNC/, by which users can easily get their desired results without the need to go through the mathematical details. Supplementary information: Supplementary data are available at Bioinformatics online. Bin Liu 0014, Fan Weng, De-Shuang Huang, Kuo-Chen Chou |
Bioinform. | 1 |
| 2018 | iPromoter-2L: a two-layer predictor for identifying promoters and their types by multi-window-based PseKNCabstractMotivation: Being responsible for initiating transaction of a particular gene in genome, promoter is a short region of DNA. Promoters have various types with different functions. Owing to their importance in biological process, it is highly desired to develop computational tools for timely identifying promoters and their types. Such a challenge has become particularly critical and urgent in facing the avalanche of DNA sequences discovered in the postgenomic age. Although some prediction methods were developed, they can only be used to discriminate a specific type of promoters from non-promoters. None of them has the ability to identify the types of promoters. This is due to the facts that different types of promoters may share quite similar consensus sequence pattern, and that the promoters of same type may have considerably different consensus sequences. Results: To overcome such difficulty, using the multi-window-based PseKNC (pseudo K-tuple nucleotide composition) approach to incorporate the short-, middle-, and long-range sequence information, we have developed a two-layer seamless predictor named as 'iPromoter-2 L'. The first layer serves to identify a query DNA sequence as a promoter or non-promoter, and the second layer to predict which of the following six types the identified promoter belongs to: σ24, σ28, σ32, σ38, σ54 and σ70. Availability and implementation: For the convenience of most experimental scientists, a user-friendly and publicly accessible web-server for the powerful new predictor has been established at http://bioinformatics.hitsz.edu.cn/iPromoter-2L/. It is anticipated that iPromoter-2 L will become a very useful high throughput tool for genome analysis. Contact: [email protected] or [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Bin Liu 0014, De-Shuang Huang, Kuo-Chen Chou |
Bioinform. | 1 |
| 2018 | Composing Semantic Collage for Image RetargetingabstractImage retargeting has been applied to display images of any size via devices with various resolutions (e.g., cell phone, TV monitors). To fit an image with the target resolution, certain unimportant regions need to be deleted or distorted and the key problem is to determine the importance of each pixel. Existing methods predict pixel-wise importance in a bottom-up manner via eye fixation estimation or saliency detection. In contrast, the proposed algorithm estimates the pixel-wise importance based on a top-down criterion where the target image maintains the semantic meaning of the original image. To this end, several semantic components corresponding to foreground objects, action contexts, and background regions are extracted. The semantic component maps are integrated by a classification guided fusion network. Specifically, the deep network classifies the original image as object or scene-oriented, and fuses the semantic component maps according to classification results. The network output, referred to as the semantic collage with the same size as the original image, is then fed into any existing optimization method to generate the target image. Extensive experiments are carried out on the RetargetMe dataset and S-Retarget database developed in this work. Experimental results demonstrate the merits of the proposed algorithm over the state-of-the-art image retargeting methods. Si Liu 0001, Zhen Wei 0001, Yao Sun 0004, Xinyu Ou, Bin Liu 0014, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 6 |
| 2018 | Correlation Particle Filter for Visual TrackingabstractIn this paper, we propose a novel correlation particle filter (CPF) for robust visual tracking. Instead of a simple combination of a correlation filter and a particle filter, we exploit and complement the strength of each one. Compared with existing tracking methods based on correlation filters and particle filters, the proposed tracker has four major advantages: 1) it is robust to partial and total occlusions, and can recover from lost tracks by maintaining multiple hypotheses; 2) it can effectively handle large-scale variation via a particle sampling strategy; 3) it can efficiently maintain multiple modes in the posterior density using fewer particles than conventional particle filters, resulting in low computational cost; and 4) it can shepherd the sampled particles toward the modes of the target state distribution using a mixture of correlation filters, resulting in robust tracking performance. Extensive experimental results on challenging benchmark data sets demonstrate that the proposed CPF tracking algorithm performs favorably against the state-of-the-art methods. Tianzhu Zhang 0001, Si Liu 0001, Changsheng Xu, Bin Liu 0014, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 4 |
| 2018 | Three-Dimensional Attention-Based Deep Ranking Model for Video Highlight DetectionabstractThe video highlight detection task is to localize key elements (moments of user's major or special interest) in a video. Most of existing highlight detection approaches extract features from the video segment as a whole without considering the difference of local features both temporally and spatially. Due to the complexity of video content, this kind of mixed features will impact the final highlight prediction. In temporal extent, not all frames are worth watching because some of them only contain the background of the environment without human or other moving objects. In spatial extent, it is similar that not all regions in each frame are highlights especially when there are lots of clutters in the background. To solve the above problem, we propose a novel three-dimensional (3-D) (spatial+temporal) attention model that can automatically localize the key elements in a video without any extra supervised annotations. Specifically, the proposed attention model produces attention weights of local regions along both the spatial and temporal dimensions of the video segment. The regions of key elements in the video will be strengthened with large weights. Thus, the more effective feature of the video segment is obtained to predict the highlight score. The proposed 3-D attention scheme can be easily integrated into a conventional end-to-end deep ranking model that aims to learn a deep neural network to compute the highlight score of each video segment. Extensive experimental results on the YouTube and SumMe datasets demonstrate that the proposed approach achieves significant improvement over state-of-the-art methods. With the proposed 3-D attention model, video highlights can be accurately retrieved in spatial and temporal dimensions without human supervision in several domains, such as gymnastics, parkour, skating, skiing, surfing, and dog activities, on the public datasets. Yifan Jiao, Zhetao Li, Shucheng Huang, Xiaoshan Yang, Bin Liu 0014, Tianzhu Zhang 0001 |
IEEE Trans. Multim. | 5 |
| 2017 | SOFM-Top: Protein Remote Homology Detection and Fold Recognition Based on Sequence-Order Frequency Matrix
Junjie Chen 0004, Mingyue Guo 0001, Xiaolong Wang 0001, Bin Liu 0014 |
ICIC (2) | 4 |
| 2017 | Protein fold recognition based on sparse representation based classification
Ke Yan 0003, Yong Xu 0001, Xiaozhao Fang, Chun-Hou Zheng 0001, Bin Liu 0014 |
Artif. Intell. Medicine | 5 |
| 2017 | ProtDec-LTR2.0: an improved method for protein remote homology detection by combining pseudo protein and supervised Learning to RankabstractSUMMARY: As one of the most important tasks in protein sequence analysis, protein remote homology detection is critical for both basic research and practical applications. Here, we present an effective web server for protein remote homology detection called ProtDec-LTR2.0 by combining ProtDec-Learning to Rank (LTR) and pseudo protein representation. Experimental results showed that the detection performance is obviously improved. The web server provides a user-friendly interface to explore the sequence and structure information of candidate proteins and find their conserved domains by launching a multiple sequence alignment tool. AVAILABILITY AND IMPLEMENTATION: The web server is free and open to all users with no login requirement at http://bioinformatics.hitsz.edu.cn/ProtDec-LTR2.0/. CONTACT: [email protected]. Junjie Chen 0004, Mingyue Guo 0001, Bin Liu 0014 |
Bioinform. | 4 |
| 2017 | iRSpot-EL: identify recombination spots with an ensemble learning approachabstractMOTIVATION: Coexisting in a DNA system, meiosis and recombination are two indispensible aspects for cell reproduction and growth. With the avalanche of genome sequences emerging in the post-genomic age, it is an urgent challenge to acquire the information of DNA recombination spots because it can timely provide very useful insights into the mechanism of meiotic recombination and the process of genome evolution. RESULTS: To address such a challenge, we have developed a predictor, called IRSPOT-EL: , by fusing different modes of pseudo K-tuple nucleotide composition and mode of dinucleotide-based auto-cross covariance into an ensemble classifier of clustering approach. Five-fold cross tests on a widely used benchmark dataset have indicated that the new predictor remarkably outperforms its existing counterparts. Particularly, far beyond their reach, the new predictor can be easily used to conduct the genome-wide analysis and the results obtained are quite consistent with the experimental map. AVAILABILITY AND IMPLEMENTATION: For the convenience of most experimental scientists, a user-friendly web-server for iRSpot-EL has been established at http://bioinformatics.hitsz.edu.cn/iRSpot-EL/, by which users can easily obtain their desired results without the need to go through the complicated mathematical equations involved. CONTACT: [email protected] or [email protected] information: Supplementary data are available at Bioinformatics online. Bin Liu 0014, Shanyi Wang, Ren Long, Kuo-Chen Chou |
Bioinform. | 1 |
| 2017 | Protein remote homology detection based on bidirectional long short-term memoryabstractBACKGROUND: Protein remote homology detection plays a vital role in studies of protein structures and functions. Almost all of the traditional machine leaning methods require fixed length features to represent the protein sequences. However, it is never an easy task to extract the discriminative features with limited knowledge of proteins. On the other hand, deep learning technique has demonstrated its advantage in automatically learning representations. It is worthwhile to explore the applications of deep learning techniques to the protein remote homology detection. RESULTS: In this study, we employ the Bidirectional Long Short-Term Memory (BLSTM) to learn effective features from pseudo proteins, also propose a predictor called ProDec-BLSTM: it includes input layer, bidirectional LSTM, time distributed dense layer and output layer. This neural network can automatically extract the discriminative features by using bidirectional LSTM and the time distributed dense layer. CONCLUSION: Experimental results on a widely-used benchmark dataset show that ProDec-BLSTM outperforms other related methods in terms of both the mean ROC and mean ROC50 scores. This promising result shows that ProDec-BLSTM is a useful tool for protein remote homology detection. Furthermore, the hidden patterns learnt by ProDec-BLSTM can be interpreted and visualized, and therefore, additional useful information can be obtained. Junjie Chen 0004, Bin Liu 0014 |
BMC Bioinform. | 3 |
| 2017 | A weakly supervised method for makeup-invariant face verification
Yao Sun 0004, Lejian Ren, Zhen Wei 0001, Bin Liu 0014, Yanlong Zhai, Si Liu 0001 |
Pattern Recognit. | 4 |
| 2016 | iEnhancer-2L: a two-layer predictor for identifying enhancers and their strength by pseudo k-tuple nucleotide compositionabstractMOTIVATION: Enhancers are of short regulatory DNA elements. They can be bound with proteins (activators) to activate transcription of a gene, and hence play a critical role in promoting gene transcription in eukaryotes. With the avalanche of DNA sequences generated in the post-genomic age, it is a challenging task to develop computational methods for timely identifying enhancers from extremely complicated DNA sequences. Although some efforts have been made in this regard, they were limited at only identifying whether a query DNA element being of an enhancer or not. According to the distinct levels of biological activities and regulatory effects on target genes, however, enhancers should be further classified into strong and weak ones in strength. RESULTS: In view of this, a two-layer predictor called ' IENHANCER-2L: ' was proposed by formulating DNA elements with the 'pseudo k-tuple nucleotide composition', into which the six DNA local parameters were incorporated. To the best of our knowledge, it is the first computational predictor ever established for identifying not only enhancers, but also their strength. Rigorous cross-validation tests have indicated that IENHANCER-2L: holds very high potential to become a useful tool for genome analysis. AVAILABILITY AND IMPLEMENTATION: For the convenience of most experimental scientists, a web server for the two-layer predictor was established at http://bioinformatics.hitsz.edu.cn/iEnhancer-2L/, by which users can easily get their desired results without the need to go through the mathematical details. CONTACT: [email protected], [email protected], [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bin Liu 0014, Longyun Fang, Ren Long, Xun Lan, Kuo-Chen Chou |
Bioinform. | 1 |
| 2016 | iDHS-EL: identifying DNase I hypersensitive sites by fusing three different modes of pseudo nucleotide composition into an ensemble learning frameworkabstractMOTIVATION: Regulatory DNA elements are associated with DNase I hypersensitive sites (DHSs). Accordingly, identification of DHSs will provide useful insights for in-depth investigation into the function of noncoding genomic regions. RESULTS: In this study, using the strategy of ensemble learning framework, we proposed a new predictor called iDHS-EL for identifying the location of DHS in human genome. It was formed by fusing three individual Random Forest (RF) classifiers into an ensemble predictor. The three RF operators were respectively based on the three special modes of the general pseudo nucleotide composition (PseKNC): (i) kmer, (ii) reverse complement kmer and (iii) pseudo dinucleotide composition. It has been demonstrated that the new predictor remarkably outperforms the relevant state-of-the-art methods in both accuracy and stability. AVAILABILITY AND IMPLEMENTATION: For the convenience of most experimental scientists, a web server for iDHS-EL is established at http://bioinformatics.hitsz.edu.cn/iDHS-EL, which is the first web-server predictor ever established for identifying DHSs, and by which users can easily get their desired results without the need to go through the mathematical details. We anticipate that IDHS-EL: will become a very useful high throughput tool for genome analysis. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bin Liu 0014, Ren Long, Kuo-Chen Chou |
Bioinform. | 1 |
| 2016 | iEnhancer-PsedeKNC: Identification of enhancers and their subgroups based on Pseudo degenerate kmer nucleotide composition
Bin Liu 0014 |
Neurocomputing | 1 |
| 2015 | Application of learning to rank to protein remote homology detectionabstractMOTIVATION: Protein remote homology detection is one of the fundamental problems in computational biology, aiming to find protein sequences in a database of known structures that are evolutionarily related to a given query protein. Some computational methods treat this problem as a ranking problem and achieve the state-of-the-art performance, such as PSI-BLAST, HHblits and ProtEmbed. This raises the possibility to combine these methods to improve the predictive performance. In this regard, we are to propose a new computational method called ProtDec-LTR for protein remote homology detection, which is able to combine various ranking methods in a supervised manner via using the Learning to Rank (LTR) algorithm derived from natural language processing. RESULTS: Experimental results on a widely used benchmark dataset showed that ProtDec-LTR can achieve an ROC1 score of 0.8442 and an ROC50 score of 0.9023 outperforming all the individual predictors and some state-of-the-art methods. These results indicate that it is correct to treat protein remote homology detection as a ranking problem, and predictive performance improvement can be achieved by combining different ranking approaches in a supervised manner via using LTR. AVAILABILITY AND IMPLEMENTATION: For users' convenience, the software tools of three basic ranking predictors and Learning to Rank algorithm were provided at http://bioinformatics.hitsz.edu.cn/ProtDec-LTR/home/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bin Liu 0014, Junjie Chen 0004, Xiaolong Wang 0001 |
Bioinform. | 1 |
| 2015 | repDNA: a Python package to generate various modes of feature vectors for DNA sequences by incorporating user-defined physicochemical properties and sequence-order effectsabstractUNLABELLED: In order to develop powerful computational predictors for identifying the biological features or attributes of DNAs, one of the most challenging problems is to find a suitable approach to effectively represent the DNA sequences. To facilitate the studies of DNAs and nucleotides, we developed a Python package called representations of DNAs (repDNA) for generating the widely used features reflecting the physicochemical properties and sequence-order effects of DNAs and nucleotides. There are three feature groups composed of 15 features. The first group calculates three nucleic acid composition features describing the local sequence information by means of kmers; the second group calculates six autocorrelation features describing the level of correlation between two oligonucleotides along a DNA sequence in terms of their specific physicochemical properties; the third group calculates six pseudo nucleotide composition features, which can be used to represent a DNA sequence with a discrete model or vector yet still keep considerable sequence-order information via the physicochemical properties of its constituent oligonucleotides. In addition, these features can be easily calculated based on both the built-in and user-defined properties via using repDNA. AVAILABILITY AND IMPLEMENTATION: The repDNA Python package is freely accessible to the public at http://bioinformatics.hitsz.edu.cn/repDNA/. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bin Liu 0014, Fule Liu, Longyun Fang, Xiaolong Wang 0001, Kuo-Chen Chou |
Bioinform. | 1 |
| 2014 | A Sentence Vector Based Over-Sampling Method for Imbalanced Emotion Classification
Tao Chen 0026, Ruifeng Xu 0001, Qin Lu 0001, Bin Liu 0014, Jun Xu 0007, Zhenyu He 0001 |
CICLing (2) | 4 |
| 2014 | Reader Emotion Prediction Using Concept and Concept Sequence Features in News Headlines
Yuanlin Yao, Ruifeng Xu 0001, Qin Lu 0001, Bin Liu 0014, Jun Xu 0007, Chengtian Zou, Li Yuan 0002, Shuwei Wang, Zhenyu He 0001 |
CICLing (2) | 4 |
| 2014 | Emotion Cause Detection with Linguistic Construction in Chinese Weibo Text
Lin Gui 0003, Li Yuan 0002, Ruifeng Xu 0001, Bin Liu 0014, Qin Lu 0001, Yu Zhou 0025 |
NLPCC | 4 |
| 2014 | Combining evolutionary information extracted from frequency profiles with sequence-based kernels for protein remote homology detectionabstractMOTIVATION: Owing to its importance in both basic research (such as molecular evolution and protein attribute prediction) and practical application (such as timely modeling the 3D structures of proteins targeted for drug development), protein remote homology detection has attracted a great deal of interest. It is intriguing to note that the profile-based approach is promising and holds high potential in this regard. To further improve protein remote homology detection, a key step is how to find an optimal means to extract the evolutionary information into the profiles. RESULTS: Here, we propose a novel approach, the so-called profile-based protein representation, to extract the evolutionary information via the frequency profiles. The latter can be calculated from the multiple sequence alignments generated by PSI-BLAST. Three top performing sequence-based kernels (SVM-Ngram, SVM-pairwise and SVM-LA) were combined with the profile-based protein representation. Various tests were conducted on a SCOP benchmark dataset that contains 54 families and 23 superfamilies. The results showed that the new approach is promising, and can obviously improve the performance of the three kernels. Furthermore, our approach can also provide useful insights for studying the features of proteins in various families. It has not escaped our notice that the current approach can be easily combined with the existing sequence-based methods so as to improve their performance as well. AVAILABILITY AND IMPLEMENTATION: For users' convenience, the source code of generating the profile-based proteins and the multiple kernel learning was also provided at http://bioinformatics.hitsz.edu.cn/main/~binliu/remote/ Bin Liu 0014, Deyuan Zhang, Ruifeng Xu 0001, Jinghao Xu, Xiaolong Wang 0001, Qingcai Chen, Qiwen Dong, Kuo-Chen Chou |
Bioinform. | 1 |
| 2014 | Using distances between Top-n-gram and residue pairs for protein remote homology detectionabstractBACKGROUND: Protein remote homology detection is one of the central problems in bioinformatics, which is important for both basic research and practical application. Currently, discriminative methods based on Support Vector Machines (SVMs) achieve the state-of-the-art performance. Exploring feature vectors incorporating the position information of amino acids or other protein building blocks is a key step to improve the performance of the SVM-based methods. RESULTS: Two new methods for protein remote homology detection were proposed, called SVM-DR and SVM-DT. SVM-DR is a sequence-based method, in which the feature vector representation for protein is based on the distances between residue pairs. SVM-DT is a profile-based method, which considers the distances between Top-n-gram pairs. Top-n-gram can be viewed as a profile-based building block of proteins, which is calculated from the frequency profiles. These two methods are position dependent approaches incorporating the sequence-order information of protein sequences. Various experiments were conducted on a benchmark dataset containing 54 families and 23 superfamilies. Experimental results showed that these two new methods are very promising. Compared with the position independent methods, the performance improvement is obvious. Furthermore, the proposed methods can also provide useful insights for studying the features of protein families. CONCLUSION: The better performance of the proposed methods demonstrates that the position dependant approaches are efficient for protein remote homology detection. Another advantage of our methods arises from the explicit feature space representation, which can be used to analyze the characteristic features of protein families. The source code of SVM-DT and SVM-DR is available at http://bioinformatics.hitsz.edu.cn/DistanceSVM/index.jsp. Bin Liu 0014, Jinghao Xu, Quan Zou 0001, Ruifeng Xu 0001, Xiaolong Wang 0001, Qingcai Chen |
BMC Bioinform. | 1 |
| 2009 | Protein Long Disordered Region Prediction Based on Profile-Level Disorder Propensities and Position-Specific Scoring MatrixesabstractIdentification of long disordered regions in protein sequence is important for understanding protein function. In this work, a class of novel propensities at profile level is presented, namely, the order profile disorder propensities, which use the evolutionary information of profile for protein long disorder prediction. These propensities, combined with position-specific scoring matrices, are inputted to the logistic regression (LR) for the prediction of protein long disordered regions. In 5-fold cross-validation test, our method can achieve an area of 97.5% under the ROC cure. Testing on a blind-test set, our method is significantly more accurate than several existing disorder predictors. Bin Liu 0014, Lei Lin 0001, Xiaolong Wang 0001, Xuan Wang 0002 |
BIBM | 1 |
| 2009 | Prediction of protein binding sites in protein structures using hidden Markov support vector machineabstractBACKGROUND: Predicting the binding sites between two interacting proteins provides important clues to the function of a protein. Recent research on protein binding site prediction has been mainly based on widely known machine learning techniques, such as artificial neural networks, support vector machines, conditional random field, etc. However, the prediction performance is still too low to be used in practice. It is necessary to explore new algorithms, theories and features to further improve the performance. RESULTS: In this study, we introduce a novel machine learning model hidden Markov support vector machine for protein binding site prediction. The model treats the protein binding site prediction as a sequential labelling task based on the maximum margin criterion. Common features derived from protein sequences and structures, including protein sequence profile and residue accessible surface area, are used to train hidden Markov support vector machine. When tested on six data sets, the method based on hidden Markov support vector machine shows better performance than some state-of-the-art methods, including artificial neural networks, support vector machines and conditional random field. Furthermore, its running time is several orders of magnitude shorter than that of the compared methods. CONCLUSION: The improved prediction performance and computational efficiency of the method based on hidden Markov support vector machine can be attributed to the following three factors. Firstly, the relation between labels of neighbouring residues is useful for protein binding site prediction. Secondly, the kernel trick is very advantageous to this field. Thirdly, the complexity of the training step for hidden Markov support vector machine is linear with the number of training samples by using the cutting-plane algorithm. Bin Liu 0014, Xiaolong Wang 0001, Lei Lin 0001, Buzhou Tang, Qiwen Dong, Xuan Wang 0002 |
BMC Bioinform. | 1 |
| 2008 | A discriminative method for protein remote homology detection and fold recognition combining Top-n-grams and latent semantic analysisabstractBACKGROUND: Protein remote homology detection and fold recognition are central problems in bioinformatics. Currently, discriminative methods based on support vector machine (SVM) are the most effective and accurate methods for solving these problems. A key step to improve the performance of the SVM-based methods is to find a suitable representation of protein sequences. RESULTS: In this paper, a novel building block of proteins called Top-n-grams is presented, which contains the evolutionary information extracted from the protein sequence frequency profiles. The protein sequence frequency profiles are calculated from the multiple sequence alignments outputted by PSI-BLAST and converted into Top-n-grams. The protein sequences are transformed into fixed-dimension feature vectors by the occurrence times of each Top-n-gram. The training vectors are evaluated by SVM to train classifiers which are then used to classify the test protein sequences. We demonstrate that the prediction performance of remote homology detection and fold recognition can be improved by combining Top-n-grams and latent semantic analysis (LSA), which is an efficient feature extraction technique from natural language processing. When tested on superfamily and fold benchmarks, the method combining Top-n-grams and LSA gives significantly better results compared to related methods. CONCLUSION: The method based on Top-n-grams significantly outperforms the methods based on many other building blocks including N-grams, patterns, motifs and binary profiles. Therefore, Top-n-gram is a good building block of the protein sequences and can be widely used in many tasks of the computational biology, such as the sequence alignment, the prediction of domain boundary, the designation of knowledge-based potentials and the prediction of protein binding sites. Bin Liu 0014, Xiaolong Wang 0001, Lei Lin 0001, Qiwen Dong, Xuan Wang 0002 |
BMC Bioinform. | 1 |