EDBT 2026 Demo / reviewers in the wild / expert
Jian Liu 0040
dblp:35/295-40
· DBLP profile ↗
24ranked-venue papers
2as first author
24since 2021 · last 2026
0000-0001-5516-0157ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 2 first-author · 19 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Feature Selection via Dynamic Feature Graph
Mykola Pechenizkiy, Jinmao Wei 0001, Jian Liu 0040 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | MaskNeo: Advancing Neoantigen Immunogenicity Prediction with Adaptive Feature MaskingabstractNeoantigens, tumor-exclusive mutated peptides presented by human leukocyte antigen (HLA) molecules, are ideal cancer vaccine targets due to their specificity. However, only a tiny fraction of mutated peptides are immunogenic and elicit T-cell responses. Accurately predicting immunogenicity remains a significant challenge. Current indirect predictors model essential steps like HLA binding but fail to guarantee immunogenicity, while direct predictors often impose constraints on inputs (e.g., wild-type peptide requirement) or exhibit suboptimal performance and robustness. Here, we introduce MaskNeo, a deep learning model for direct immunogenicity prediction requiring only neoantigen peptide sequences and HLA alleles. MaskNeo integrates the feature masking mechanism within a pre-trained transformer to selectively suppress task-irrelevant features. Moreover, MaskNeo employs an adaptive balancing strategy to combine complementary features by learning optimal weights. Benchmarking against state-of-the-art methods (BigMHC, DeepImmuno, DeepNeo, NeoaPred) demonstrates MaskNeo's superior predictive ability on balanced data, enhanced robustness on imbalanced data, improved generalizability in cross-domain tests, and consistent performance across diverse peptide lengths and HLA alleles. Ablation studies confirm the contribution of MaskNeo's components. MaskNeo provides a significant advance in neoantigen immunogenicity prediction, positioning it as a powerful tool for prioritizing targets for cancer vaccines. The source code is available at: https://github.com/lyotvincent/MaskNeo. Jian Liu 0040 |
BIBM | 2 |
| 2025 | A machine learning-based method to optimize the immunogenicity of human leukocyte antigen class I-restricted neoantigensabstractNeoantigens, arising from somatic mutations, have the potential to induce immune responses against tumor cells. As natural neoantigens that drive immune responses are uncommon, some methods to optimize neoantigens have been devised, leading to "heteroclitic" neoantigens capable of triggering cross-immunization. Existing methods, however, concentrate exclusively on the affinity of neoantigens for human leukocyte antigen (HLA), while neglecting the enhancement of their immunogenicity. Here, we developed a machine learning-based method, Naso, which integrates simulated annealing search and multi-objective training to optimize the immunogenicity of neoantigens. We also designed a search space optimization strategy to improve search efficiency. Experimental evaluation demonstrated that Naso outperforms existing methods in enhancing the immunogenicity of neoantigens. Specifically, Naso can transform nonimmunoreactive neoantigens into immunoreactive ones with no more than three mutation amino acids. Moreover, Naso can obtain competitive results in other immunological features, such as binding affinity and presentation. Consequently, Naso-optimized neoantigens can elicit an immune response and facilitate cross-immunization, thereby promoting the development of heteroclitic neoantigen-based vaccines. The source code of Naso is available at https://github.com/lyotvincent/Naso. Jinyi Yu, Chaoyang Yan, Jian Liu 0040 |
Briefings Bioinform. | 5 |
| 2025 | HiADN: Lightweight Resolution Enhancement of Hi-C Data Using High Information Attention Distillation NetworkabstractDue to limitations in experimental library preparation and in practical sequencing cost, currently available Hi-C data is often sparse, affecting the precise characterization of complex 3D chromatin structures. Providing efficaciously computational models to elevate the quality of sparse Hi-C sequencing data for restoring the fundamental traits of 3D chromatin is of substantial significance. Herein, we introduce HiADN, a deep learning-based approach to infer dense high-resolution matrices from sparse Hi-C matrices. In particular, we firstly design a specialized architecture HiFM to captures local spatial structures and the patterns of Hi-C data. Then, we develop large kernel convolutional decomposition and attention mechanisms to effectively explore global patterns across longer genomic distances. Using HiADN, it is possible to construct biologically significant regions at high-resolution (e.g., 10 Kb) while only using the 1/100 of original sequencing reads. The experimental results demonstrate that the effect of in silico libraries forecasted by computational models using HiADN is commensurate with that of experimental libraries, surpassing the state-of-the-art (SOTA) models. We further validated the effectiveness of HiADN in reconstructing the three-dimensional spatial structure of chromosomes on the GM12878, K562, and CH12-LX cell line datasets. Pingjing Li, Jiuxin Feng, Jian Liu 0040 |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | A Deep Learning Framework for Chromatin Loop De Novo Prediction With Enhanced Feature ExtractionabstractChromatin encompasses a variety of three-dimensional structures with distinct forms and ranges, among which chromatin loops play a crucial role in gene regulatory mechanisms and the maintenance of cellular homeostasis. Despite the recognition of the critical role of chromatin loops, existing prediction models often overlook the heterogeneity among different feature data, and there is a scarcity of corresponding prediction tools. To fill this gap, we introduce CHASOS2 (CHromatin loop prediction with Anchor Score and OCR Score), a user-friendly toolkit for de novo prediction and evaluation of chromatin loop. CHASOS2 uses convolutional modules with multi-receptive fields to generate features capable of mitigating the heterogeneities among diverse feature data and employs a gradient boosting tree model to predict chromatin loops. Experimental evaluations indicate that CHASOS2 outperforms existing methods, particularly in scenarios involving heterogeneous feature data. A case study applying CHASOS2 de novo prediction toolkit on the K562 cell line demonstrates high consistency with ChIA-PET identified chromatin loops, validating the effectiveness of our method and toolkit. Jian Liu 0040 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | PCLT-PPI: Predicting Multi-Type Interactions Between Proteins Based on Point Cloud Structure and Local Topology PreservationabstractProtein-protein interactions (PPIs) play a crucial role in cellular biochemical reactions. Computationally mining PPI can help us better understand cellular regulatory mechanisms. Most existing methods focus on the linear structure of proteins, ignoring the influence of native spatial structure on their properties. Furthermore, when neural networks are used to learn protein embeddings, the nonlinear transformations may change the topological relationships between proteins. To address the above issues, we propose a PPI prediction method based on protein point cloud structure and local topology preservation, naming it PCLT-PPI. It extracts structural features from protein point cloud structures and relational features through graph neural networks. Throughout the process, PCLT-PPI maintains the local topology of proteins in their origin and embedding spaces. Experimental results show that, under three test set partition modes (Random, BFS, DFS) and four evaluation metrics (F1, AUC, AUPR, Hamming Loss), PCLT-PPI performs better than several state-of-the-art PPI prediction methods, especially when predicting protein PPIs that are not visible during training, exhibiting stronger robustness and higher generalization ability. The results also demonstrate that point cloud structure and local topology preservation can improve PPI prediction performance, which may provide a reference for subsequent related research. Yurui Hou, Jinmao Wei 0001, Jian Liu 0040 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | AREDCI: Assessing Reproducibility and Differential Chromatin Interactions for ChIA-PET Sequencing DataabstractUnderstanding the three-dimensional (3D) architecture of chromatin is pivotal for unraveling gene regulation and cellular processes. Currently, a wealth of data on chromatin interactions, such as ChIA-PET sequencing data, is increasingly accessible. However, challenges persist in comparative analyses of these chromatin interactions. Specifically, accurately identifying differential chromatin interactions (DCIs) remains challenging, yet it is crucial for studying gene expression differences during cellular differentiation. Additionally, assessing the inter-sample reproducibility at the sample level, which is a fundamental indicator of the reliability and consistency of replicate experiments, still lacks a computational method. In this work, we present AREDCI, a novel approach integrating data preprocessing, normalization, reproducibility assessment, and DCI identification. By leveraging multiple normalization techniques and developing a self-similarity-based reproducibility algorithm, AREDCI offers a robust evaluation of sample consistency. Additionally, AREDCI designs algorithms based on Kernel Density Estimation (KDE) and local contextual information to enhance the accuracy and reliability of DCI identification. Experimental evaluations demonstrate that AREDCI outperforms existing methods in both reproducibility assessment and DCI identification. Through experiments on simulated and real-world datasets, AREDCI exhibits commendable precision and recall, showcasing its effectiveness in analyzing chromatin interaction data. Notably, AREDCI successfully identifies significant DCIs during mouse cell differentiation, aligning with known biological processes and molecular functions. Zhihan Ruan, Chaoyang Yan, Jian Liu 0040 |
BIBM | 4 |
| 2024 | CHASOS: A Novel Deep Learning Approach for Chromatin Loop Predictions
Jian Liu 0040 |
ISBRA (1) | 3 |
| 2024 | CRISPR-M: Predicting sgRNA off-target effect using a multi-view deep learning networkabstractUsing the CRISPR-Cas9 system to perform base substitutions at the target site is a typical technique for genome editing with the potential for applications in gene therapy and agricultural productivity. When the CRISPR-Cas9 system uses guide RNA to direct the Cas9 endonuclease to the target site, it may misdirect it to a potential off-target site, resulting in an unintended genome editing. Although several computational methods have been proposed to predict off-target effects, there is still room for improvement in the off-target effect prediction capability. In this paper, we present an effective approach called CRISPR-M with a new encoding scheme and a novel multi-view deep learning model to predict the sgRNA off-target effects for target sites containing indels and mismatches. CRISPR-M takes advantage of convolutional neural networks and bidirectional long short-term memory recurrent neural networks to construct a three-branch network towards multi-views. Compared with existing methods, CRISPR-M demonstrates significant performance advantages running on real-world datasets. Furthermore, experimental analysis of CRISPR-M under multiple metrics reveals its capability to extract features and validates its superiority on sgRNA off-target effect predictions. Jian Liu 0040 |
PLoS Comput. Biol. | 3 |
| 2024 | Unsupervised feature selection by learning exponential weights
Jun Wang 0023, Zhichen Gu, Jinmao Wei 0001, Jian Liu 0040 |
Pattern Recognit. | 5 |
| 2023 | HiBrowser: an interactive and dynamic browser for synchronous Hi-C data visualizationabstractWith the development of chromosome conformation capture technology, the genome-wide investigation of higher-order chromatin structure by using high-throughput chromatin conformation capture (Hi-C) technology is emerging as an important component for understanding the mechanism of gene regulation. Considering genetic and epigenetic differences are typically used to explore the pathological reasons on the chromosome and gene level, visualizing multi-omics data and performing an intuitive analysis by using an interactive browser become a powerful and welcomed way. In this paper, we develop an effective sequence and chromatin interaction data display browser called HiBrowser for visualizing and analyzing Hi-C data and their associated genetic and epigenetic annotations. The advantages of HiBrowser are flexible multi-omics navigation, novel multidimensional synchronization comparisons and dynamic interaction system. In particular, HiBrowser first provides an out of the box web service and allows flexible and dynamic reconstruction of custom annotation tracks on demand during running. In order to conveniently and intuitively analyze the similarities and differences among multiple samples, such as visual comparisons of normal and tumor tissue samples, and pan genomes of multiple (consanguineous) species, HiBrowser develops a clone mode to synchronously display the genome coordinate positions or the same regions of multiple samples on the same page of visualization. HiBrowser also supports a pluralistic and precise search on correlation data of distal cis-regulatory elements and navigation to any region on Hi-C heatmap of interest according to the searching results. HiBrowser is a no-build tool, and could be easily deployed in local server. The source code is available at https://github.com/lyotvincent/HiBrowser. Pingjing Li, Jianguo Lu, Jian Liu 0040 |
Briefings Bioinform. | 5 |
| 2023 | Word-Context Attention for Text Representation
Chengkai Piao, Yapeng Zhu, Jinmao Wei 0001, Jian Liu 0040 |
Neural Process. Lett. | 5 |
| 2023 | A Deep Neural Network-Based Co-Coding Method to Predict Drug-Protein Interactions by Analyzing the Feature Consistency Between Drugs and ProteinsabstractExploring drug-protein interactions (DPIs) through computational methods can effectively reduce the workload and the cost of DPI identification. Previous works try to predict DPIs by integrating and analyzing the unique features of drugs and proteins. They cannot adequately analyze the consistency between the drug features and the protein features due to their different semantics. However, the consistency of their features, such as the correlation originating from their sharing diseases, may reveal some potential DPIs. Here we propose a deep neural network-based co-coding method (DNNCC for short) to predict novel DPIs. DNNCC projects the original features of drugs and proteins to a common embedding space through a co-coding strategy. In this way, the embedding features of drugs and proteins have the same semantics. Therefore, the prediction module can discover the unknown DPIs by exploring the feature consistency between drugs and proteins. The experimental results indicate that the performance of DNNCC is significantly superior to five state-of-the-art DPI prediction methods under several evaluation metrics. The superiority of integrating and analyzing the common features of drugs and proteins is proved by the ablation experiments. The novel DPIs predicted by DNNCC verify that DNNCC is a powerful prior tool that can effectively discover potential DPIs. Chang Sun 0002, Rong Tang 0004, Jipeng Huang, Jinmao Wei 0001, Jian Liu 0040 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2023 | Predicting Drug-Protein Interactions by Self-Adaptively Adjusting the Topological Structure of the Heterogeneous NetworkabstractMany powerful computational methods based on graph neural networks (GNNs) have been proposed to predict drug-protein interactions (DPIs). It can effectively reduce laboratory workload and the cost of drug discovery and drug repurposing. However, many clinical functions of drugs and proteins are unknown due to their unobserved indications. Therefore, it is difficult to establish a reliable drug-protein heterogeneous network that can describe the relationships between drugs and proteins based on the available information. To solve this problem, we propose a DPI prediction method that can self-adaptively adjust the topological structure of the heterogeneous networks, and name it SATS. SATS establishes a representation learning module based on graph attention network to carry out the drug-protein heterogeneous network. It can self-adaptively learn the relationships among the nodes based on their attributes and adjust the topological structure of the network according to the training loss of the model. Finally, SATS predicts the interaction propensity between drugs and proteins based on their embeddings. The experimental results show that SATS can effectively improve the topological structure of the network. The performance of SATS outperforms several state-of-the-art DPI prediction methods under various evaluation metrics. These prove that SATS is useful to deal with incomplete data and unreliable networks. The case studies on the top section of the prediction results further demonstrate that SATS is powerful for discovering novel DPIs. Rong Tang 0004, Chang Sun 0002, Jipeng Huang, Jinmao Wei 0001, Jian Liu 0040 |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | SCDD: a novel single-cell RNA-seq imputation method with diffusion and denoisingabstractSingle-cell sequencing technologies are widely used to discover the evolutionary relationships and the differences in cells. Since dropout events may frustrate the analysis, many imputation approaches for single-cell RNA-seq data have appeared in previous attempts. However, previous imputation attempts usually suffer from the over-smooth problem, which may bring limited improvement or negative effect for the downstream analysis of single-cell RNA-seq data. To solve this difficulty, we propose a novel two-stage diffusion-denoising method called SCDD for large-scale single-cell RNA-seq imputation in this paper. We introduce the diffusion i.e. a direct imputation strategy using the expression of similar cells for potential dropout sites, to perform the initial imputation at first. After the diffusion, a joint model integrated with graph convolutional neural network and contractive autoencoder is developed to generate superposition states of similar cells, from which we restore the original states and remove the noise introduced by the diffusion. The final experimental results indicate that SCDD could effectively suppress the over-smooth problem and remarkably improve the effect of single-cell RNA-seq downstream analysis, including clustering and trajectory analysis. Jian Liu 0040, Yichen Pan, Zhihan Ruan |
Briefings Bioinform. | 1 |
| 2022 | Multi-variable AUC for sifting complementary features and its biomedical applicationabstractAlthough sifting functional genes has been discussed for years, traditional selection methods tend to be ineffective in capturing potential specific genes. First, typical methods focus on finding features (genes) relevant to class while irrelevant to each other. However, the features that can offer rich discriminative information are more likely to be the complementary ones. Next, almost all existing methods assess feature relations in pairs, yielding an inaccurate local estimation and lacking a global exploration. In this paper, we introduce multi-variable Area Under the receiver operating characteristic Curve (AUC) to globally evaluate the complementarity among features by employing Area Above the receiver operating characteristic Curve (AAC). Due to AAC, the class-relevant information newly provided by a candidate feature and that preserved by the selected features can be achieved beyond pairwise computation. Furthermore, we propose an AAC-based feature selection algorithm, named Multi-variable AUC-based Combined Features Complementarity, to screen discriminative complementary feature combinations. Extensive experiments on public datasets demonstrate the effectiveness of the proposed approach. Besides, we provide a gene set about prostate cancer and discuss its potential biological significance from the machine learning aspect and based on the existing biomedical findings of some individual genes. Keyu Du, Jun Wang 0023, Jinmao Wei 0001, Jian Liu 0040 |
Briefings Bioinform. | 5 |
| 2022 | Drug-Protein interaction prediction by correcting the effect of incomplete information in heterogeneous informationabstractMOTIVATION: Large-scale heterogeneous data provide diverse perspectives for predicting drug-protein interactions (DPIs). However, the available information on molecular interactions and clinical associations related to drugs or proteins is incomplete because there may be unproven interactions and associations. This incomplete information in the available data is presented in the form of non-interaction and non-correlation, which may mislead the prediction model. Existing methods fuse incomplete and complete information without considering their integrity, so the negative effects of incomplete information still exist. RESULTS: We develop a network-based DPI prediction method named BRWCP, which uses the complete information network to correct the prediction results acquired by the incomplete information network. By integrating relevant heterogeneous information that may be incomplete, the feature similarities of drugs and proteins are obtained. Combining the feature similarities and known DPIs, an incomplete information-based drug-protein heterogeneous network is constructed. Then, a bidirectional random walk with pruning algorithm is adopted in this heterogeneous network to predict potential DPIs. Next, the predicted DPIs are combined with the chemical fingerprint similarity of drugs and amino acid sequence similarity of proteins to construct the complete information network. The bidirectional random walk with pruning algorithm is applied in the new network to obtain the final prediction results until it converges. Experimental results show that BRWCP is superior to several state-of-the-art DPI prediction methods, and case studies further confirm its ability to tap potential DPIs. AVAILABILITY AND IMPLEMENTATION: The code and data used in BRWCP are available at https://github.com/lyfdomain/BRWCP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chang Sun 0002, Jinmao Wei 0001, Jian Liu 0040 |
Bioinform. | 4 |
| 2022 | Multi-objective data enhancement for deep learning-based ultrasound analysisabstractRecently, Deep Learning based automatic generation of treatment recommendation has been attracting much attention. However, medical datasets are usually small, which may lead to over-fitting and inferior performances of deep learning models. In this paper, we propose multi-objective data enhancement method to indirectly scale up the medical data to avoid over-fitting and generate high quantity treatment recommendations. Specifically, we define a main and several auxiliary tasks on the same dataset and train a specific model for each of these tasks to learn different aspects of knowledge in limited data scale. Meanwhile, a Soft Parameter Sharing method is exploited to share learned knowledge among models. By sharing the knowledge learned by auxiliary tasks to the main task, the proposed method can take different semantic distributions into account during the training process of the main task. We collected an ultrasound dataset of thyroid nodules that contains Findings, Impressions and Treatment Recommendations labeled by professional doctors. We conducted various experiments on the dataset to validate the proposed method and justified its better performance than existing methods. Chengkai Piao, Mengyue Lv, Rongyan Zhou, Jinmao Wei 0001, Jian Liu 0040 |
BMC Bioinform. | 7 |
| 2021 | Adversarial Dual-Channel Variational Graph Autoencoder for Synthetic Lethality Prediction in Human CancersabstractSynthetic Lethality (SL) is a type of vital gene interaction that can lead to various human diseases including cancers. Therefore, SL gene pair prediction can aid in the prevention and treatment of cancer. A number of computational approaches, especially Graph Neural Network (GNN) based methods, have been proposed for this link prediction problem on the graph. However, these GNN-based methods only consider embedding as deterministic vectors and do not take data distribution into account. Here we propose an Adversarial Dual-Channel Variational Graph Autoencoder based on semi-implicit variational inference for SL prediction in human cancers. We consider node embedding as a random variable that has an explicit Gaussian distribution. Then we design a dual-channel GCN encoder to inject stochasticity into the distribution parameters and allow latent embedding to exceed the Gaussian distribution. This hierarchical scheme leads to a more flexible posterior of latent embedding and enhances the model representation capacity. To further obtain a robust and stable representation, an adversarial module is devised for variance regularization. Experimental results compared with other state-of-the-art methods confirm the effectiveness of our proposed method. Moreover, we conduct a case study to demonstrate that our model can be very useful to predict novel SL pairs. Wei Li 0184, Han Zhang 0017, Jian Liu 0040, Yanbin Yin |
BIBM | 4 |
| 2021 | Ontology-based annotation and retrieval for large-scale VCF dataabstractSequencing cost is dramatically reduced by the development of the next-generation sequencing (NGS) technologies. Currently, numerous variant call format (VCF) data and biomedical ontologies, which store mutations data and special biomedical knowledge to applications in the field of biomedical researches such as human genetics, etc., become available in the bioinformatics community. There are some bioinformatics tools developed for the VCF data annotation and analysis. However, most previous works ignore the biomedical ontologies associated with the genetic data, which are usually beneficial to analyze genetic diseases and molecular diagnosis. In particular, annotating information with biomedical ontologies remains an obstacle. In order to effectively integrate biomedical ontologies and enhance the analysis across multiple biomedical sources, we present an automatic workflow called OntoAnnotation for annotating VCF files with biomedical ontologies. Additionally, to facilitate the retrieval of large-scale VCF data for non-bioinformaticians, we develop a web platform called OntoVarSearch, which provides a flexible engine that allows convenient access to genetic variants and ontology-based annotation information stored in the MongoDB database. The OntoAnnotation tool and the OntoVarSearch platform could provide a simple way for users without sufficient programming skills to annotate information with biomedical ontologies and search data stored in VCF files. Jian Liu 0040, Yongzhuang Liu |
BIBM | 1 |
| 2021 | Autoencoder-based drug-target interaction prediction by preserving the consistency of chemical properties and functions of drugsabstractMOTIVATION: Exploring the potential drug-target interactions (DTIs) is a key step in drug discovery and repurposing. In recent years, predicting the probable DTIs through computational methods has gradually become a research hot spot. However, most of the previous studies failed to judiciously take into account the consistency between the chemical properties of drug and its functions. The changes of these relationships may lead to a severely negative effect on the prediction of DTIs. RESULTS: We propose an autoencoder-based method, AEFS, under spatial consistency constraints to predict DTIs. A heterogeneous network is established to integrate the information of drugs, proteins and diseases. The original drug features are projected to an embedding (protein) space by a multi-layer encoder, and further projected into label (disease) space by a decoder. In this process, the clinical information of drugs is introduced to assist the DTI prediction. By maintaining the distribution of drug correlation in the original feature, embedding and label space, AEFS keeps the consistency between chemical properties and functions of drugs. Experimental comparisons indicate that AEFS is more robust for imbalanced data and of significantly superior performance in DTI prediction. Case studies further confirm its ability to mine the latent DTIs. AVAILABILITY AND IMPLEMENTATION: The code of AEFS is available at https://github.com/JackieSun818/AEFS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chang Sun 0002, Yangkun Cao, Jinmao Wei 0001, Jian Liu 0040 |
Bioinform. | 4 |
| 2021 | Improved estimation of model quality using predicted inter-residue distanceabstractMOTIVATION: Protein model quality assessment (QA) is an essential component in protein structure prediction, which aims to estimate the quality of a structure model and/or select the most accurate model out from a pool of structure models, without knowing the native structure. QA remains a challenging task in protein structure prediction. RESULTS: Based on the inter-residue distance predicted by the recent deep learning-based structure prediction algorithm trRosetta, we developed QDistance, a new approach to the estimation of both global and local qualities. QDistance works for both single- and multi-models inputs. We designed several distance-based features to assess the agreement between the predicted and model-derived inter-residue distances. Together with a few widely used features, they are fed into a simple yet powerful linear regression model to infer the global QA scores. The local QA scores for each structure model are predicted based on a comparative analysis with a set of selected reference models. For multi-models input, the reference models are selected from the input based on the predicted global QA scores. For single-model input, the reference models are predicted by trRosetta. With the informative distance-based features, QDistance can predict the global quality with satisfactory accuracy. Benchmark tests on the CASP13 and the CAMEO structure models suggested that QDistance was competitive with other methods. Blind tests in the CASP14 experiments showed that QDistance was robust and ranked among the top predictors. Especially, QDistance was the top 3 local QA method and made the most accurate local QA prediction for unreliable local region. Analysis showed that this superior performance can be attributed to the inclusion of the predicted inter-residue distance. AVAILABILITY AND IMPLEMENTATION: http://yanglab.nankai.edu.cn/QDistance. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lisha Ye, Peikun Wu, Zhen-Ling Peng, Jianzhao Gao, Jian Liu 0040, Jianyi Yang 0002 |
Bioinform. | 5 |
| 2021 | Multi-class feature selection by exploring reliable class correlation
Jinmao Wei 0001, Jian Liu 0040 |
Knowl. Based Syst. | 4 |
| 2021 | Unsupervised Cross-View Feature Selection on incomplete data
Yuanyuan Xu 0002, Jun Wang 0023, Jinmao Wei 0001, Jian Liu 0040, Lina Yao 0001, Wenjie Zhang 0001 |
Knowl. Based Syst. | 5 |