EDBT 2026 Demo / reviewers in the wild / expert
Ke Yan 0003
dblp:28/7692-3
· DBLP profile ↗
30ranked-venue papers
15as first author
21since 2021 · last 2026
0000-0002-5326-4267ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 12 first-author · 16 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PepLM-GNN: A graph neural network framework leveraging pre-trained language models for peptide-protein binding predictionabstractMOTIVATION: The precise prediction of peptide-protein interaction (PepPI) is a core support for promoting breakthroughs in peptide drug research, as well as understanding the regulatory mechanisms of biomolecules. Researchers have developed several computational methods to predict PepPI. However, existing computational methods also have significant limitations. At the level of data feature characterisation, the problem of PepPI does not conform to the Euclidean axioms, making it difficult for conventional prediction methods to effectively measure the underlying correlations between peptides and proteins. At the level of model generalisation performance, existing approaches are often hampered by insufficient generalisation ability, as manifested by their markedly degraded performance in cold start scenarios involving novel peptides, novel proteins, and novel binding pairs. RESULTS: In this study, we propose a computing framework, PepLM-GNN, that integrates a pre-trained language ProtT5 model with a hybrid graph network for accurate identification of PepPI. This model constructs a graph by using ProtT5-extracted semantic context features of peptides and proteins to form heterogeneous nodes, with edges connecting interacting peptide-protein pairs. The hybrid graph network Graph Convolutional Networks (GCN) provides the comprehensive information of the peptide and protein sequences, while employing the Graph Isomorphism Network (GIN) to capture the global interactions between them. Specifically, the GCN aggregates both the semantic context information of node sequences and local neighbourhood information, effectively representing non-Euclidean data. To capture the global associations, we adopt a GIN strategy to optimize the cross-node feature interaction and transfer process, thereby enhancing the generalisation performance of addressing the cold start scenario. Compared with the existing advanced methods, PepLM-GNN demonstrated highly accurate performance and robustness in predicting the PepPI. We further demonstrated the capabilities of PepLM-GNN in virtual peptide drug screening, which is expected to facilitate the discovery of peptide drugs and the elucidation of protein functions. Ke Yan 0003, Meijing Li, Shutao Chen, Bin Liu 0014 |
PLoS Comput. Biol. | 1 |
| 2025 | Accurate prediction of toxicity peptide and its function using multi-view tensor learning and latent semantic learning frameworkabstractMOTIVATION: Therapeutic peptide is an important ingredient in the treatment of various diseases and drug discovery. The toxicity of peptides is one of the major challenges in peptide drug therapy. With the abundance of therapeutic peptides generated in the post-genomics era, it is a challenge to promptly identify toxicity peptides using computational methods. Although several efforts have been made, few algorithms are designed to identify whether a query peptide exhibits toxicity. Considering the varied levels of biological activities, the toxicity peptides should be further classified into multi-functional peptides. RESULTS: This study introduces a two-level predictor, ToxPre-2L, developed using the multi-view tensor learning and latent semantic learning framework. The proposed method utilized multi-label learning with feature induced labels to avoid the redundancy of information from each view. Then the multi-view tensor learning was employed to establish the latent semantic information among different views, while low-rank constraint learning was leveraged to exploit the correlation information among multi-labels. Finally, we constructed an updated toxicity peptide benchmark dataset to assess the effectiveness of the proposed method. Experimental results demonstrated that ToxPre-2L achieves a better performance than alternative computational methods in the prediction of toxicity peptides and their multi-functional types. AVAILABILITY AND IMPLEMENTATION: The source code and data of ToxPre-2L can be accessed at http://bliulab.net/ToxPre-2L. Ke Yan 0003, Shutao Chen, Bin Liu 0014, Hao Wu 0066 |
Bioinform. | 1 |
| 2025 | Protein Language Pragmatic Analysis and Progressive Transfer Learning for Profiling Peptide-Protein InteractionsabstractProtein complex structural data are growing at an unprecedented pace, but its complexity and diversity pose significant challenges for protein function research. Although deep learning models have been widely used to capture the syntactic structure, word semantics, or semantic meanings of polypeptide and protein sequences, these models often overlook the complex contextual information of sequences. Here, we propose interpretable interaction deep learning (IIDL)-peptide-protein interaction (PepPI), a deep learning model designed to tackle these challenges using data-driven and interpretable pragmatic analysis to profile PepPIs. IIDL-PepPI constructs bidirectional attention modules to represent the contextual information of peptides and proteins, enabling pragmatic analysis. It then adopts a progressive transfer learning framework to simultaneously predict PepPIs and identify binding residues for specific interactions, providing a solution for multilevel in-depth profiling. We validate the performance and robustness of IIDL-PepPI in accurately predicting peptide-protein binary interactions and identifying binding residues compared with the state-of-the-art methods. We further demonstrate the capability of IIDL-PepPI in peptide virtual drug screening and binding affinity assessment, which is expected to advance artificial intelligence-based peptide drug discovery and protein function elucidation. Shutao Chen, Ke Yan 0003, Xuelong Li 0001, Bin Liu 0014 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | TPpred-SC: multi-functional therapeutic peptide prediction based on multi-label supervised contrastive learning
Ke Yan 0003, Hongwu Lv, Jiangyi Shao, Shutao Chen, Bin Liu 0014 |
Sci. China Inf. Sci. | 1 |
| 2024 | Projective Incomplete Multi-View ClusteringabstractDue to the rapid development of multimedia technology and sensor technology, multi-view clustering (MVC) has become a research hotspot in machine learning, data mining, and other fields and has been developed significantly in the past decades. Compared with single-view clustering, MVC improves clustering performance by exploiting complementary and consistent information among different views. Such methods are all based on the assumption of complete views, which means that all the views of all the samples exist. It limits the application of MVC, because there are always missing views in practical situations. In recent years, many methods have been proposed to solve the incomplete MVC (IMVC) problem and a kind of popular method is based on matrix factorization (MF). However, such methods generally cannot deal with new samples and do not take into account the imbalance of information between different views. To address these two issues, we propose a new IMVC method, in which a novel and simple graph regularized projective consensus representation learning model is formulated for incomplete multi-view data clustering task. Compared with the existing methods, our method not only can obtain a set of projections to handle new samples but also can explore information of multiple views in a balanced way by learning the consensus representation in a unified low-dimensional subspace. In addition, a graph constraint is imposed on the consensus representation to mine the structural information inside the data. Experimental results on four datasets show that our method successfully accomplishes the IMVC task and obtain the best clustering performance most of the time. Our implementation is available at https://github.com/Dshijie/PIMVC. Jie Wen 0001, Chengliang Liu 0003, Ke Yan 0003, Gehui Xu, Yong Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Deep Double Incomplete Multi-View Multi-Label Learning With Incomplete Labels and Missing ViewsabstractView missing and label missing are two challenging problems in the applications of multi-view multi-label classification scenery. In the past years, many efforts have been made to address the incomplete multi-view learning or incomplete multi-label learning problem. However, few works can simultaneously handle the challenging case with both the incomplete issues. In this article, we propose a new incomplete multi-view multi-label learning network to address this challenging issue. The proposed method is composed of four major parts: view-specific deep feature extraction network, weighted representation fusion module, classification module, and view-specific deep decoder network. By, respectively, integrating the view missing information and label missing information into the weighted fusion module and classification module, the proposed method can effectively reduce the negative influence caused by two such incomplete issues and sufficiently explore the available data and label information to obtain the most discriminative feature extractor and classifier. Furthermore, our method can be trained in both supervised and semi-supervised manners, which has important implications for flexible deployment. Experimental results on five benchmarks in supervised and semi-supervised cases demonstrate that the proposed method can greatly enhance the classification performance on the difficult incomplete multi-view multi-label classification tasks with missing labels and missing views. Jie Wen 0001, Chengliang Liu 0003, Lunke Fei, Ke Yan 0003, Yong Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | DAmiRLocGNet: miRNA subcellular localization prediction by combining miRNA-disease associations and graph convolutional networksabstractMicroRNAs (miRNAs) are human post-transcriptional regulators in humans, which are involved in regulating various physiological processes by regulating the gene expression. The subcellular localization of miRNAs plays a crucial role in the discovery of their biological functions. Although several computational methods based on miRNA functional similarity networks have been presented to identify the subcellular localization of miRNAs, it remains difficult for these approaches to effectively extract well-referenced miRNA functional representations due to insufficient miRNA-disease association representation and disease semantic representation. Currently, there has been a significant amount of research on miRNA-disease associations, making it possible to address the issue of insufficient miRNA functional representation. In this work, a novel model is established, named DAmiRLocGNet, based on graph convolutional network (GCN) and autoencoder (AE) for identifying the subcellular localizations of miRNA. The DAmiRLocGNet constructs the features based on miRNA sequence information, miRNA-disease association information and disease semantic information. GCN is utilized to gather the information of neighboring nodes and capture the implicit information of network structures from miRNA-disease association information and disease semantic information. AE is employed to capture sequence semantics from sequence similarity networks. The evaluation demonstrates that the performance of DAmiRLocGNet is superior to other competing computational approaches, benefiting from implicit features captured by using GCNs. The DAmiRLocGNet has the potential to be applied to the identification of subcellular localization of other non-coding RNAs. Moreover, it can facilitate further investigation into the functional mechanisms underlying miRNA localization. The source code and datasets are accessed at http://bliulab.net/DAmiRLocGNet. Ke Yan 0003, Bin Liu 0014 |
Briefings Bioinform. | 2 |
| 2023 | PreHom-PCLM: protein remote homology detection by combing motifs and protein cubic language modelabstractProtein remote homology detection is essential for structure prediction, function prediction, disease mechanism understanding, etc. The remote homology relationship depends on multiple protein properties, such as structural information and local sequence patterns. Previous studies have shown the challenges for predicting remote homology relationship by protein features at sequence level (e.g. position-specific score matrix). Protein motifs have been used in structure and function analysis due to their unique sequence patterns and implied structural information. Therefore, designing a usable architecture to fuse multiple protein properties based on motifs is urgently needed to improve protein remote homology detection performance. To make full use of the characteristics of motifs, we employed the language model called the protein cubic language model (PCLM). It combines multiple properties by constructing a motif-based neural network. Based on the PCLM, we proposed a predictor called PreHom-PCLM by extracting and fusing multiple motif features for protein remote homology detection. PreHom-PCLM outperforms the other state-of-the-art methods on the test set and independent test set. Experimental results further prove the effectiveness of multiple features fused by PreHom-PCLM for remote homology detection. Furthermore, the protein features derived from the PreHom-PCLM show strong discriminative power for proteins from different structural classes in the high-dimensional space. Availability and Implementation: http://bliulab.net/PreHom-PCLM. Jiangyi Shao, Ke Yan 0003, Bin Liu 0014 |
Briefings Bioinform. | 3 |
| 2023 | iDRPro-SC: identifying DNA-binding proteins and RNA-binding proteins based on subfunction classifiersabstractNucleic acid-binding proteins are proteins that interact with DNA and RNA to regulate gene expression and transcriptional control. The pathogenesis of many human diseases is related to abnormal gene expression. Therefore, recognizing nucleic acid-binding proteins accurately and efficiently has important implications for disease research. To address this question, some scientists have proposed the method of using sequence information to identify nucleic acid-binding proteins. However, different types of nucleic acid-binding proteins have different subfunctions, and these methods ignore their internal differences, so the performance of the predictor can be further improved. In this study, we proposed a new method, called iDRPro-SC, to predict the type of nucleic acid-binding proteins based on the sequence information. iDRPro-SC considers the internal differences of nucleic acid-binding proteins and combines their subfunctions to build a complete dataset. Additionally, we used an ensemble learning to characterize and predict nucleic acid-binding proteins. The results of the test dataset showed that iDRPro-SC achieved the best prediction performance and was superior to the other existing nucleic acid-binding protein prediction methods. We have established a web server that can be accessed online: http://bliulab.net/iDRPro-SC. Ke Yan 0003, Hao Wu 0066 |
Briefings Bioinform. | 1 |
| 2023 | PreTP-2L: identification of therapeutic peptides and their types using two-layer ensemble learning frameworkabstractMOTIVATION: Therapeutic peptides play an important role in immune regulation. Recently various therapeutic peptides have been used in the field of medical research, and have great potential in the design of therapeutic schedules. Therefore, it is essential to utilize the computational methods to predict the therapeutic peptides. However, the therapeutic peptides cannot be accurately predicted by the existing predictors. Furthermore, chaotic datasets are also an important obstacle of the development of this important field. Therefore, it is still challenging to develop a multi-classification model for identification of therapeutic peptides and their types. RESULTS: In this work, we constructed a general therapeutic peptide dataset. An ensemble-learning method named PreTP-2L was developed for predicting various therapeutic peptide types. PreTP-2L consists of two layers. The first layer predicts whether a peptide sequence belongs to therapeutic peptide, and the second layer predicts if a therapeutic peptide belongs to a particular species. AVAILABILITY AND IMPLEMENTATION: A user-friendly webserver PreTP-2L can be accessed at http://bliulab.net/PreTP-2L. Ke Yan 0003, Bin Liu 0014 |
Bioinform. | 1 |
| 2023 | sAMPpred-GAT: prediction of antimicrobial peptide by graph attention network and predicted peptide structureabstractMOTIVATION: Antimicrobial peptides (AMPs) are essential components of therapeutic peptides for innate immunity. Researchers have developed several computational methods to predict the potential AMPs from many candidate peptides. With the development of artificial intelligent techniques, the protein structures can be accurately predicted, which are useful for protein sequence and function analysis. Unfortunately, the predicted peptide structure information has not been applied to the field of AMP prediction so as to improve the predictive performance. RESULTS: In this study, we proposed a computational predictor called sAMPpred-GAT for AMP identification. To the best of our knowledge, sAMPpred-GAT is the first approach based on the predicted peptide structures for AMP prediction. The sAMPpred-GAT predictor constructs the graphs based on the predicted peptide structures, sequence information and evolutionary information. The Graph Attention Network (GAT) is then performed on the graphs to learn the discriminative features. Finally, the full connection networks are utilized as the output module to predict whether the peptides are AMP or not. Experimental results show that sAMPpred-GAT outperforms the other state-of-the-art methods in terms of AUC, and achieves better or highly comparable performance in terms of the other metrics on the eight independent test datasets, demonstrating that the predicted peptide structure information is important for AMP prediction. AVAILABILITY AND IMPLEMENTATION: A user-friendly webserver of sAMPpred-GAT can be accessed at http://bliulab.net/sAMPpred-GAT and the source code is available at https://github.com/HongWuL/sAMPpred-GAT/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ke Yan 0003, Hongwu Lv, Bin Liu 0014 |
Bioinform. | 1 |
| 2023 | PreTP-Stack: Prediction of Therapeutic Peptides Based on the Stacked Ensemble LearingabstractTherapeutic peptide prediction is critical for drug development and therapeutic therapy. Researchers have developed several computational methods to identify different therapeutic peptide types. However, most computational methods focus on identifying the specific type of therapeutic peptides and fail to accurately predict all types of therapeutic peptides. Moreover, it is still challenging to utilize different properties features to predict the therapeutic peptides. In this study, a novel stacking framework PreTP-Stack is proposed for predicting different types of therapeutic peptide. PreTP-Stack is constructed based on ten different features and four predictors (Random Forest, Linear Discriminant Analysis, XGBoost and Support Vector Machine). Then the proposed method constructs an auto-weighted multi-view learning model as a final meta-classifier to enhance the performance of the basic models. Experimental results showed that the proposed method achieved better or highly comparable performance with the state-of-the-art methods for predicting eight types of therapeutic peptides A user-friendly web-server predictor is available at http://bliulab.net/PreTP-Stack. Ke Yan 0003, Hongwu Lv, Jie Wen 0001, Yong Xu 0001, Bin Liu 0014 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2022 | iDRNA-ITF: identifying DNA- and RNA-binding residues in proteins based on induction and transfer frameworkabstractProtein-DNA and protein-RNA interactions are involved in many biological activities. In the post-genome era, accurate identification of DNA- and RNA-binding residues in protein sequences is of great significance for studying protein functions and promoting new drug design and development. Therefore, some sequence-based computational methods have been proposed for identifying DNA- and RNA-binding residues. However, they failed to fully utilize the functional properties of residues, leading to limited prediction performance. In this paper, a sequence-based method iDRNA-ITF was proposed to incorporate the functional properties in residue representation by using an induction and transfer framework. The properties of nucleic acid-binding residues were induced by the nucleic acid-binding residue feature extraction network, and then transferred into the feature integration modules of the DNA-binding residue prediction network and the RNA-binding residue prediction network for the final prediction. Experimental results on four test sets demonstrate that iDRNA-ITF achieves the state-of-the-art performance, outperforming the other existing sequence-based methods. The webserver of iDRNA-ITF is freely available at http://bliulab.net/iDRNA-ITF. Ning Wang 0054, Ke Yan 0003, Jun Zhang 0078, Bin Liu 0014 |
Briefings Bioinform. | 2 |
| 2022 | TPpred-ATMV: therapeutic peptide prediction by adaptive multi-view tensor learning modelabstractMOTIVATION: Therapeutic peptide prediction is important for the discovery of efficient therapeutic peptides and drug development. Researchers have developed several computational methods to identify different therapeutic peptide types. However, these computational methods focus on identifying some specific types of therapeutic peptides, failing to predict the comprehensive types of therapeutic peptides. Moreover, it is still challenging to utilize different properties to predict the therapeutic peptides. RESULTS: In this study, an adaptive multi-view based on the tensor learning framework TPpred-ATMV is proposed for predicting different types of therapeutic peptides. TPpred-ATMV constructs the class and probability information based on various sequence features. We constructed the latent subspace among the multi-view features and constructed an auto-weighted multi-view tensor learning model to utilize the high correlation based on the multi-view features. Experimental results showed that the TPpred-ATMV is better than or highly comparable with the other state-of-the-art methods for predicting eight types of therapeutic peptides. AVAILABILITY AND IMPLEMENTATION: The code of TPpred-ATMV is accessed at: https://github.com/cokeyk/TPpred-ATMV. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ke Yan 0003, Hongwu Lv, Yongyong Chen, Hao Wu 0066, Bin Liu 0014 |
Bioinform. | 1 |
| 2022 | PreRBP-TL: prediction of species-specific RNA-binding proteins based on transfer learningabstractMOTIVATION: RNA-binding proteins (RBPs) play crucial roles in post-transcriptional regulation. Accurate identification of RBPs helps to understand gene expression, regulation, etc. In recent years, some computational methods were proposed to identify RBPs. However, these methods fail to accurately identify RBPs from some specific species with limited data, such as bacteria. RESULTS: In this study, we introduce a computational method called PreRBP-TL for identifying species-specific RBPs based on transfer learning. The weights of the prediction model were initialized by pretraining with the large general RBP dataset and then fine-tuned with the small species-specific RPB dataset by using transfer learning. The experimental results show that the PreRBP-TL achieves better performance for identifying the species-specific RBPs from Human, Arabidopsis, Escherichia coli and Salmonella, outperforming eight state-of-the-art computational methods. It is anticipated PreRBP-TL will become a useful method for identifying RBPs. AVAILABILITY AND IMPLEMENTATION: For the convenience of researchers to identify RBPs, the web server of PreRBP-TL was established, freely available at http://bliulab.net/PreRBP-TL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jun Zhang 0078, Ke Yan 0003, Qingcai Chen, Bin Liu 0014 |
Bioinform. | 2 |
| 2021 | PreTP-EL: prediction of therapeutic peptides based on ensemble learningabstractTherapeutic peptides are important for understanding the correlation between peptides and their therapeutic diagnostic potential. The therapeutic peptides can be further divided into different types based on therapeutic function sharing different characteristics. Although some computational approaches have been proposed to predict different types of therapeutic peptides, they failed to accurately predict all types of therapeutic peptides. In this study, a predictor called PreTP-EL has been proposed via employing the ensemble learning approach to fuse the different features and machine learning techniques in order to capture the different characteristics of various therapeutic peptides. Experimental results showed that PreTP-EL outperformed other competing methods. Availability and implementation: A user-friendly web-server of PreTP-EL predictor is available at http://bliulab.net/PreTP-EL. Ke Yan 0003, Hongwu Lv, Bin Liu 0014 |
Briefings Bioinform. | 2 |
| 2021 | FoldRec-C2C: protein fold recognition by combining cluster-to-cluster model and protein similarity networkabstractAs a key for studying the protein structures, protein fold recognition is playing an important role in predicting the protein structures associated with COVID-19 and other important structures. However, the existing computational predictors only focus on the protein pairwise similarity or the similarity between two groups of proteins from 2-folds. However, the homology relationship among proteins is in a hierarchical structure. The global protein similarity network will contribute to the performance improvement. In this study, we proposed a predictor called FoldRec-C2C to globally incorporate the interactions among proteins into the prediction. For the FoldRec-C2C predictor, protein fold recognition problem is treated as an information retrieval task in nature language processing. The initial ranking results were generated by a surprised ranking algorithm Learning to Rank, and then three re-ranking algorithms were performed on the ranking lists to adjust the results globally based on the protein similarity network, including seq-to-seq model, seq-to-cluster model and cluster-to-cluster model (C2C). When tested on a widely used and rigorous benchmark dataset LINDAHL dataset, FoldRec-C2C outperforms other 34 state-of-the-art methods in this field. The source code and data of FoldRec-C2C can be downloaded from http://bliulab.net/FoldRec-C2C/download. Jiangyi Shao, Ke Yan 0003, Bin Liu 0014 |
Briefings Bioinform. | 2 |
| 2021 | MLDH-Fold: Protein fold recognition based on multi-view low-rank modeling
Ke Yan 0003, Jie Wen 0001, Yong Xu 0001, Bin Liu 0014 |
Neurocomputing | 1 |
| 2021 | Protein Fold Recognition by Combining Support Vector Machines and Pairwise Sequence Similarity ScoresabstractProtein fold recognition is one of the most essential steps for protein structure prediction, aiming to classify proteins into known protein folds. There are two main computational approaches: one is the template-based method based on the alignment scores between query-template protein pairs and the other is the machine learning method based on the feature representation and classifier. These two approaches have their own advantages and disadvantages. Can we combine these methods to establish more accurate predictors for protein fold recognition? In this study, we made an initial attempt and proposed two novel algorithms: TSVM-fold and ESVM-fold. TSVM-fold was based on the Support Vector Machines (SVMs), which utilizes a set of pairwise sequence similarity scores generated by three complementary template-based methods, including HHblits, SPARKS-X, and DeepFR. These scores measured the global relationships between query sequences and templates. The comprehensive features of the attributes of the sequences were fed into the SVMs for the prediction. Then the TSVM-fold was further combined with the HHblits algorithm so as to improve its generalization ability. The combined method is called ESVM-fold. Experimental results in two rigorous benchmark datasets (LE and YK datasets) showed that the proposed methods outperform some state-of-the-art methods, indicating that the TSVM-fold and ESVM-fold are efficient predictors for protein fold recognition. Ke Yan 0003, Jie Wen 0001, Jin-Xing Liu 0001, Yong Xu 0001, Bin Liu 0014 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | Protein Fold Recognition Based on Auto-Weighted Multi-View Graph Embedding Learning ModelabstractProtein fold recognition is critical for studies of the protein structure prediction and drug design. Several methods have been proposed to obtain discriminative features from the protein sequences for fold recognition. However, the ensemble methods that combine the various features to improve predictive performance remain the challenge problems. In this study, we proposed two novel algorithms: AWMG and EMfold. AWMG used a novel predictor based on the multi-view learning framework for fold recognition. Each view was treated as the intermediate representation of the corresponding data source of proteins, including the evolutionary information and the retrieval information. AWMG calculated the auto-weight for each view respectively and constructed the latent subspace which contains the common information shared by different views. The marginalized constraint was employed to enlarge the margins between different folds, improving the predictive performance of AWMG. Furthermore, we proposed a novel ensemble method called EMfold, which combines two complementary methods AWMG and DeepSS. The later method was a template-based algorithm using the SPARKS-X and DeepFR programs. EMfold integrated the advantages of template-based assignment and machine learning classifier. Experimental results on the two widely datasets (LE and YK) showed that the proposed methods outperformed some state-of-the-art methods, indicating that AWMG and EMfold are useful tools for protein fold recognition. Ke Yan 0003, Jie Wen 0001, Yong Xu 0001, Bin Liu 0014 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | Adaptive Graph Completion Based Incomplete Multi-View ClusteringabstractIn real-world applications, it is often that the collected multi-view data are incomplete, i.e., some views of samples are absent. Existing clustering methods for incomplete multi-view data all focus on obtaining a common representation or graph from the available views but neglect the hidden information of missing views and information imbalance of different views. To solve these problems, a novel method, called adaptive graph completion based incomplete multi-view clustering (AGC_IMC), is proposed in this paper. Specifically, AGC_IMC develops a joint framework for graph completion and consensus representation learning, which mainly contains three components, i.e., within-view preservation, between-view inferring, and consensus representation learning. To reduce the negative influence of information imbalance, AGC_IMC introduces some adaptive weights to balance the importance of different views during the consensus representation learning. Importantly, AGC_IMC has the potential to recover the similarity graphs of all views with the optimal cluster structure, which encourages it to obtain a more discriminative consensus representation. Experimental results on five well-known datasets show that AGC_IMC significantly outperforms the state-of-the-art methods. Jie Wen 0001, Ke Yan 0003, Zheng Zhang 0006, Yong Xu 0001, Junqian Wang, Lunke Fei, Bob Zhang 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | DeepSVM-fold: protein fold recognition by combining support vector machines and pairwise sequence similarity scores generated by deep learning networksabstractProtein fold recognition is critical for studying the structures and functions of proteins. The existing protein fold recognition approaches failed to efficiently calculate the pairwise sequence similarity scores of the proteins in the same fold sharing low sequence similarities. Furthermore, the existing feature vectorization strategies are not able to measure the global relationships among proteins from different protein folds. In this article, we proposed a new computational predictor called DeepSVM-fold for protein fold recognition by introducing a new feature vector based on the pairwise sequence similarity scores calculated from the fold-specific features extracted by deep learning networks. The feature vectors are then fed into a support vector machine to construct the predictor. Experimental results on the benchmark dataset (LE) show that DeepSVM-fold obviously outperforms all the other competing methods. Bin Liu 0014, Chen-Chen Li, Ke Yan 0003 |
Briefings Bioinform. | 3 |
| 2020 | Fold-LTR-TCP: protein fold recognition based on triadic closure principleabstractAs an important task in protein structure and function studies, protein fold recognition has attracted more and more attention. The existing computational predictors in this field treat this task as a multi-classification problem, ignoring the relationship among proteins in the dataset. However, previous studies showed that their relationship is critical for protein homology analysis. In this study, the protein fold recognition is treated as an information retrieval task. The Learning to Rank model (LTR) was employed to retrieve the query protein against the template proteins to find the template proteins in the same fold with the query protein in a supervised manner. The triadic closure principle (TCP) was performed on the ranking list generated by the LTR to improve its accuracy by considering the relationship among the query protein and the template proteins in the ranking list. Finally, a predictor called Fold-LTR-TCP was proposed. The rigorous test on the LE benchmark dataset showed that the Fold-LTR-TCP predictor achieved an accuracy of 73.2%, outperforming all the other competing methods. Bin Liu 0014, Ke Yan 0003 |
Briefings Bioinform. | 3 |
| 2019 | Protein fold recognition based on multi-view modelingabstractMOTIVATION: Protein fold recognition has attracted increasing attention because it is critical for studies of the 3D structures of proteins and drug design. Researchers have been extensively studying this important task, and several features with high discriminative power have been proposed. However, the development of methods that efficiently combine these features to improve the predictive performance remains a challenging problem. RESULTS: In this study, we proposed two algorithms: MV-fold and MT-fold. MV-fold is a new computational predictor based on the multi-view learning model for fold recognition. Different features of proteins were treated as different views of proteins, including the evolutionary information, secondary structure information and physicochemical properties. These different views constituted the latent space. The ε-dragging technique was employed to enlarge the margins between different protein folds, improving the predictive performance of MV-fold. Then, MV-fold was combined with two template-based methods: HHblits and HMMER. The ensemble method is called MT-fold incorporating the advantages of both discriminative methods and template-based methods. Experimental results on five widely used benchmark datasets (DD, RDD, EDD, TG and LE) showed that the proposed methods outperformed some state-of-the-art methods in this field, indicating that MV-fold and MT-fold are useful computational tools for protein fold recognition and protein homology detection and would be efficient tools for protein sequence analysis. Finally, we constructed an update and rigorous benchmark dataset based on SCOPe (version 2.07) to fairly evaluate the performance of the proposed method, and our method achieved stable performance on this new dataset. This new benchmark dataset will become a widely used benchmark dataset to fairly evaluate the performance of different methods for fold recognition. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ke Yan 0003, Xiaozhao Fang, Yong Xu 0001, Bin Liu 0014 |
Bioinform. | 1 |
| 2019 | Robust Sparse Linear Discriminant AnalysisabstractLinear discriminant analysis (LDA) is a very popular supervised feature extraction method and has been extended to different variants. However, classical LDA has the following problems: 1) The obtained discriminant projection does not have good interpretability for features; 2) LDA is sensitive to noise; and 3) LDA is sensitive to the selection of number of projection directions. In this paper, a novel feature extraction method called robust sparse linear discriminant analysis (RSLDA) is proposed to solve the above problems. Specifically, RSLDA adaptively selects the most discriminative features for discriminant analysis by introducing the$l_{2,1}$norm. An orthogonal matrix and a sparse matrix are also simultaneously introduced to guarantee that the extracted features can hold the main energy of the original data and enhance the robustness to noise, and thus RSLDA has the potential to perform better than other discriminant methods. Extensive experiments on six databases demonstrate that the proposed method achieves the competitive performance compared with other state-of-the-art feature extraction methods. Moreover, the proposed method is robust to the noisy data. Jie Wen 0001, Xiaozhao Fang, Jinrong Cui, Lunke Fei, Ke Yan 0003, Yan Chen 0018, Yong Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2019 | Low-Rank Preserving Projection Via Graph Regularized ReconstructionabstractPreserving global and local structures during projection learning is very important for feature extraction. Although various methods have been proposed for this goal, they commonly introduce an extra graph regularization term and the corresponding regularization parameter that needs to be tuned. However, tuning the parameter manually not only is time-consuming, but also is difficult to find the optimal value to obtain a satisfactory performance. This greatly limits their applications. Besides, projections learned by many methods do not have good interpretability and their performances are commonly sensitive to the value of the selected feature dimension. To solve the above problems, a novel method named low-rank preserving projection via graph regularized reconstruction (LRPP_GRR) is proposed. In particular, LRPP_GRR imposes the graph constraint on the reconstruction error of data instead of introducing the extra regularization term to capture the local structure of data, which can greatly reduce the complexity of the model. Meanwhile, a low-rank reconstruction term is exploited to preserve the global structure of data. To improve the interpretability of the learned projection, a sparse term with${l_{2,1}}$norm is imposed on the projection. Furthermore, we introduce an orthogonal reconstruction constraint to make the learned projection hold main energy of data, which enables LRPP_GRR to be more flexible in the selection of feature dimension. Extensive experimental results show the proposed method can obtain competitive performance with other state-of-the-art methods. Jie Wen 0001, Na Han, Xiaozhao Fang, Lunke Fei, Ke Yan 0003, Shanhua Zhan |
IEEE Trans. Cybern. | 5 |
| 2018 | Learning Domain-Invariant Subspace Using Domain Features and Independence MaximizationabstractDomain adaptation algorithms are useful when the distributions of the training and the test data are different. In this paper, we focus on the problem of instrumental variation and time-varying drift in the field of sensors and measurement, which can be viewed as discrete and continuous distributional change in the feature space. We propose maximum independence domain adaptation (MIDA) and semi-supervised MIDA to address this problem. Domain features are first defined to describe the background information of a sample, such as the device label and acquisition time. Then, MIDA learns a subspace which has maximum independence with the domain features, so as to reduce the interdomain discrepancy in distributions. A feature augmentation strategy is also designed to project samples according to their backgrounds so as to improve the adaptation. The proposed algorithms are flexible and fast. Their effectiveness is verified by experiments on synthetic datasets and four real-world ones on sensors, measurement, and computer vision. They can greatly enhance the practicability of sensor systems, as well as extend the application scope of existing domain adaptation algorithms by uniformly handling different kinds of distributional change. Ke Yan 0003, Lu Kou, David Zhang 0001 |
IEEE Trans. Cybern. | 1 |
| 2017 | Protein fold recognition based on sparse representation based classification
Ke Yan 0003, Yong Xu 0001, Xiaozhao Fang, Chun-Hou Zheng 0001, Bin Liu 0014 |
Artif. Intell. Medicine | 1 |
| 2016 | Local multiple directional pattern of palmprint imageabstractLines are the most essential and discriminative features of palmprint images, which motivate researches to propose various line direction based methods for palmprint recognition. Conventional methods usually capture the only one of the most dominant direction of palmprint images. However, a number of points in palmprint images have double or even more than two dominant directions because of a plenty of crossing lines of palmprint images. In this paper, we propose a local multiple directional pattern (LMDP) to effectively characterize the multiple direction features of palmprint images. LMDP can not only exactly denote the number and positions of dominant directions but also effectively reflect the confidence of each dominant direction. Then, a simple and effective coding scheme is designed to represent the LMDP and a block-wise LMDP descriptor is used as the feature space of palmprint images in palmprint recognition. Extensive experimental results demonstrate the superiority of the LMDP over the conventional powerful descriptors and the state-of-the-art direction based methods in palmprint recognition. Lunke Fei, Jie Wen 0001, Zheng Zhang 0006, Ke Yan 0003, Zuofeng Zhong |
ICPR | 4 |
| 2015 | An Improved Denoising Method Based on Wavelet Transform for Processing Bases Sequence Images
Ke Yan 0003, Jin-Xing Liu 0001, Yong Xu 0001 |
ICIC (1) | 1 |