EDBT 2026 Demo / reviewers in the wild / expert
Lijun Quan
dblp:128/5596
· DBLP profile ↗
21ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0003-4551-4198ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Identifying batch-integrated domains from spatial transcriptomics via graph autoencoder with contrastive learning based on cross-modality and data augmentationabstractSpatially resolved transcriptomics (SRT) allows for the comprehensive profiling of gene expression while preserving spatial context, advancing the study of tissue architecture. However, existing computational approaches still face key limitations, particularly the insufficient exploitation of histology information and the lack of cross-modal meaningful contrastive strategies for biological analyses. To overcome these challenges, we propose GCAST, a graph contrastive autoencoder framework for spatial transcriptomics that seamlessly integrates multimodal SRT data. GCAST adopts a self-supervised strategy to derive biologically meaningful representations directly from histology images when available. GCAST constructs dual graph views based on data augmentation and introduces a novel contrastive learning designed to leverage histology-weighted and gene-weighted features and improve biological interpretability. In addition, GCAST employs a block-diagonal graph construction to automatically align multiple datasets, achieving batch-effect correction without manual intervention. The framework not only captures spatial gene expression patterns to identify tissue domains but also adapts to datasets with or without histological images and supports the integration of multiple datasets for joint analyses. Overall, GCAST provides a unified and biologically informed framework that has the potential to facilitate deeper analyses of spatial transcriptomics. Yexuan Mao, Lijun Quan, Guozheng Zhang, Yelu Jiang, Liangpeng Nie, Tingfang Wu, Lingkun Meng, Qiang Lyu |
Briefings Bioinform. | 2 |
| 2025 | QGraphTad: A Novel Quantum-Inspired Method for Conformational B-Cell Epitope PredictionabstractThe recognition of B-cell epitopes plays a crucial role in vaccine design, immunodiagnostic testing, and monoclonal antibody development. While experimental approaches are expensive and time-consuming, computational methods offer a promising alternative for accelerating epitopes recognition. Despite existing prediction methods, there is still a lot of room for improving the prediction accuracy and performance. In this study, a novel quantum-inspired model for predicting conformational B-cell epitopes, named QGraphTad, is proposed. The protein sequence is characterized by sequence embeddings and structural features respectively. These features are subsequently encoded into high-dimensional quantum states and undergo simulated quantum feature encoding (QFE) operations, including rotation, entanglement, and measurement projection, to simulate the structural polymorphism and global interactions of proteins. Then a quantum-enhanced network (QEN) consisting mainly of projection layer, Quantum graph attention (GAT) layer and position-encoded transformer encoders (PTE) is designed to capture complex structure-dependent relationships. Simultaneously, the output from the QFE module is also fed into a three-layer transformer encoder network (TEN) to model global contextual relationships within antigen sequences. After fusing the representations derived from QEN and TEN, the combined features are processed by quantum-enhanced multi-layer perceptrons (QMLP) to predict conformational B-cell epitopes (BCEs). Comprehensive testing on the benchmark dataset demonstrates that QGraphTad significantly outperforms existing state-of-theart methods across multiple evaluation metrics. Buzhong Zhang, Zhaolun Yao, Yajun Xu, Lijun Quan |
BIBM | 4 |
| 2025 | MSOFormer: Multi-scale Transformer with Orthogonal Embedding and Frequency Modeling for Multivariate Time Series ForecastingabstractMultivariate Time Series Forecasting (MTSF) plays a critical role in diverse practical applications. Although Transformer-based models have recently achieved impressive results in this field, their performance is still hindered by three core challenges: complex temporal dependencies, diverse inter-variable correlations, and patterns that span multiple time scales. To address these issues, we propose MSOFormer-a Multi-scale Transformer with Orthogonal Embedding and Frequency Modeling. Specifically, the Dynamic Frequency Filter adaptively weights frequency components across variables based on input characteristics, enabling full-spectrum modeling and precise extraction of key frequency patterns. To improve inter-variable representation, we introduce Orthogonal Embedding, a novel projection strategy for queries and keys that enhances feature diversity in channel-wise self-attention. In addition, Multi-scale Patch Embedding captures temporal features across different scales, providing a comprehensive time series representation. To evaluate MTSF in cloud-native environments, we construct the first three Cloud Kafka cluster datasets, specifically curated for elastic message queue scaling scenarios. Extensive experiments across eleven real-world benchmark datasets demonstrate that MSOFormer consistently outperforms existing state-of-the-art methods, highlighting its effectiveness and broad applicability. Zongtang Hu, Dapeng Sun, Lijun Quan |
CIKM | 6 |
| 2025 | DS-MVP: identifying disease-specific pathogenicity of missense variants by pre-training representationabstractAccurately predicting the pathogenicity of missense variants is crucial for improving disease diagnosis and advancing clinical research. However, existing computational methods primarily focus on general pathogenicity predictions, overlooking assessments of disease-specific conditions. In this study, we propose DS-MVP, a method capable of predicting disease-specific pathogenicity of missense variants in human genomes. DS-MVP first leverages a deep learning model pre-trained on a large general pathogenicity dataset to learn rich representation of missense variants. It then fine-tunes these representations with an XGBoost model on smaller datasets for specific diseases. We evaluated the learned representation by testing it on multiple binary pathogenicity datasets and gene-level statistics, demonstrating that DS-MVP outperforms existing state-of-the-art methods, such as MetaRNN and AlphaMissense. Additionally, DS-MVP excels in multi-label and multi-class classification, effectively classifying disease-specific pathogenic missense variants based on disease conditions. It further enhances predictions by fine-tuning the pre-trained model on disease-specific datasets. Finally, we analyzed the contributions of the pre-trained model and various feature types, with gene description corpus features from large language model and genetic feature fusion contributing the most. These results underscore that DS-MVP represents a broader perspective on pathogenicity prediction and holds potential as an effective tool for disease diagnosis. Qiufeng Chen, Lijun Quan, Lexin Cao, Liangchen Peng, Yelu Jiang, Liangpeng Nie, Tingfang Wu, Qiang Lyu |
Briefings Bioinform. | 2 |
| 2024 | RPEMHC: improved prediction of MHC-peptide binding affinity by a deep learning approach based on residue-residue pair encodingabstractMOTIVATION: Binding of peptides to major histocompatibility complex (MHC) molecules plays a crucial role in triggering T cell recognition mechanisms essential for immune response. Accurate prediction of MHC-peptide binding is vital for the development of cancer therapeutic vaccines. While recent deep learning-based methods have achieved significant performance in predicting MHC-peptide binding affinity, most of them separately encode MHC molecules and peptides as inputs, potentially overlooking critical interaction information between the two. RESULTS: In this work, we propose RPEMHC, a new deep learning approach based on residue-residue pair encoding to predict the binding affinity between peptides and MHC, which encode an MHC molecule and a peptide as a residue-residue pair map. We evaluate the performance of RPEMHC on various MHC-II-related datasets for MHC-peptide binding prediction, demonstrating that RPEMHC achieves better or comparable performance against other state-of-the-art baselines. Moreover, we further construct experiments on MHC-I-related datasets, and experimental results demonstrate that our method can work on both two MHC classes. These extensive validations have manifested that RPEMHC is an effective tool for studying MHC-peptide interactions and can potentially facilitate the vaccine development. AVAILABILITY: The source code of the method along with trained models is freely available at https://github.com/lennylv/RPEMHC. Tingfang Wu, Yelu Jiang, Taoning Chen, Deng Pan 0006, Jingxin Xie, Lijun Quan, Qiang Lyu |
Bioinform. | 8 |
| 2024 | MultiModRLBP: A Deep Learning Approach for Multi-Modal RNA-Small Molecule Ligand Binding Sites PredictionabstractThis study aims to tackle the intricate challenge of predicting RNA-small molecule binding sites to explore the potential value in the field of RNA drug targets. To address this challenge, we propose the MultiModRLBP method, which integrates multi-modal features using deep learning algorithms. These features include 3D structural properties at the nucleotide base level of the RNA molecule, relational graphs based on overall RNA structure, and rich RNA semantic information. In our investigation, we gathered 851 interactions between RNA and small molecule ligand from the RNAglib dataset and RLBind training set. Unlike conventional training sets, this collection broadened its scope by including RNA complexes that have the same RNA sequence but change their respective binding sites due to structural differences or the presence of different ligands. This enhancement enables the MultiModRLBP model to more accurately capture subtle changes at the structural level, ultimately improving its ability to discern nuances among similar RNA conformations. Furthermore, we evaluated MultiModRLBP on two classic test sets, Test18 and Test3, highlighting its performance disparities on small molecules based on metal and non-metal ions. Additionally, we conducted a structural sensitivity analysis on specific complex categories, considering RNA instances with varying degrees of structural changes and whether they share the same ligands. The research results indicate that MultiModRLBP outperforms the current state-of-the-art methods on multiple classic test sets, particularly excelling in predicting binding sites for non-metal ions and instances where the binding sites are widely distributed along the sequence. MultiModRLBP also can be used as a potential tool when the RNA structure is perturbed or the RNA experimental tertiary structure is not available. Most importantly, MultiModRLBP exhibits the capability to distinguish binding characteristics of RNA that are structurally diverse yet exhibit sequence similarity. These advancements hold promise in reducing the costs associated with the development of RNA-targeted drugs. Lijun Quan, Hongjie Wu, Xuhao Ma, Jingxin Xie, Deng Pan 0006, Taoning Chen, Tingfang Wu, Qiang Lyu |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | CAPLA: improved prediction of protein-ligand binding affinity by a deep learning approach based on a cross-attention mechanismabstractMOTIVATION: Accurate and rapid prediction of protein-ligand binding affinity is a great challenge currently encountered in drug discovery. Recent advances have manifested a promising alternative in applying deep learning-based computational approaches for accurately quantifying binding affinity. The structure complementarity between protein-binding pocket and ligand has a great effect on the binding strength between a protein and a ligand, but most of existing deep learning approaches usually extracted the features of pocket and ligand by these two detached modules. RESULTS: In this work, a new deep learning approach based on the cross-attention mechanism named CAPLA was developed for improved prediction of protein-ligand binding affinity by learning features from sequence-level information of both protein and ligand. Specifically, CAPLA employs the cross-attention mechanism to capture the mutual effect of protein-binding pocket and ligand. We evaluated the performance of our proposed CAPLA on comprehensive benchmarking experiments on binding affinity prediction, demonstrating the superior performance of CAPLA over state-of-the-art baseline approaches. Moreover, we provided the interpretability for CAPLA to uncover critical functional residues that contribute most to the binding affinity through the analysis of the attention scores generated by the cross-attention mechanism. Consequently, these results indicate that CAPLA is an effective approach for binding affinity prediction and may contribute to useful help for further consequent applications. AVAILABILITY AND IMPLEMENTATION: The source code of the method along with trained models is freely available at https://github.com/lennylv/CAPLA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tingfang Wu, Taoning Chen, Deng Pan 0006, Jingxin Xie, Lijun Quan, Qiang Lyu |
Bioinform. | 7 |
| 2023 | TransRNAm: Identifying Twelve Types of RNA Modifications by an Interpretable Multi-Label Deep Learning Model Based on TransformerabstractAccurate identification of RNA modification sites is of great significance in understanding the functions and regulatory mechanisms of RNAs. Recent advances have shown great promise in applying computational methods based on deep learning for accurate prediction of RNA modifications. However, those methods generally predicted only a single type of RNA modification. In addition, such methods suffered from the scarcity of the interpretability for their predicted results. In this work, a new Transformer-based deep learning method was proposed to predict multiple RNA modifications simultaneously, referred to as TransRNAm. More specifically, TransRNAm employs Transformer to extract contextual feature and convolutional neural networks to further learn high-latent feature representations of RNA sequences relevant for RNA modifications. Importantly, by integrating the self-attention mechanism in Transformer with convolutional neural network, TransRNAm is capable of not only capturing the critical nucleotide sites that contribute significantly to RNA modification prediction, but also revealing the underlying association among different types of RNA modifications. Consequently, this work provided an accurate and interpretable predictor for multiple RNA modification prediction, which may contribute to uncovering the sequence-based forming mechanism of RNA modification sites. Taoning Chen, Tingfang Wu, Deng Pan 0006, Jinxing Xie, Lijun Quan, Qiang Lyu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2023 | DGCddG: Deep Graph Convolution for Predicting Protein-Protein Binding Affinity Changes Upon MutationsabstractEffectively and accurately predicting the effects of interactions between proteins after amino acid mutations is a key issue for understanding the mechanism of protein function and drug design. In this study, we present a deep graph convolution (DGC) network-based framework, DGCddG, to predict the changes of protein-protein binding affinity after mutation. DGCddG incorporates multi-layer graph convolution to extract a deep, contextualized representation for each residue of the protein complex structure. The mined channels of the mutation sites by DGC is then fitted to the binding affinity with a multi-layer perceptron. Experiments with results on multiple datasets show that our model can achieve relatively good performance for both single and multi-point mutations. For blind tests on datasets related to angiotensin-converting enzyme 2 binding with the SARS-CoV-2 virus, our method shows better results in predicting ACE2 changes, may help in finding favorable antibodies. Code and data availability: https://github.com/lennylv/DGCddG. Yelu Jiang, Lijun Quan, Yiting Zhou, Tingfang Wu, Qiang Lyu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | ctP2ISP: Protein-Protein Interaction Sites Prediction Using Convolution and Transformer With Data AugmentationabstractProteinprotein interactions are the basis of many cellular biological processes, such as cellular organization, signal transduction, and immune response. Identifying proteinprotein interaction sites is essential for understanding the mechanisms of various biological processes, disease development, and drug design. However, it remains a challenging task to make accurate predictions, as the small amount of training data and severe imbalanced classification reduce the performance of computational methods. We design a deep learning method named ctP2ISP to improve the prediction of proteinprotein interaction sites. ctP2ISP employs Convolution and Transformer to extract information and enhance information perception so that semantic features can be mined to identify proteinprotein interaction sites. A weighting loss function with different sample weights is designed to suppress the preference of the model toward multi-category prediction. To efficiently reuse the information in the training set, a preprocessing of data augmentation with an improved sample-oriented sampling strategy is applied. The trained ctP2ISP was evaluated against current state-of-the-art methods on six public datasets. The results show that ctP2ISP outperforms all other competing methods on the balance metrics: F1, MCC, and AUPRC. In particular, our prediction on open tests related to viruses may also be consistent with biological insights. Lijun Quan, Yelu Jiang, Yiting Zhou, Tingfang Wu, Qiang Lyu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | How Deepbics Quantifies Intensities of Transcription Factor-DNA Binding and Facilitates Prediction of Single Nucleotide Variant Pathogenicity With a Deep Learning Model Trained On ChIP-Seq Data SetsabstractThe binding of DNA sequences to cell type-specific transcription factors is essential for regulating gene expression in all organisms. Many variants occurring in these binding regions play crucial roles in human disease by disrupting the cis-regulation of gene expression. We first implemented a sequence-based deep learning model called deepBICS to quantify the intensity of transcription factors-DNA binding. The experimental results not only showed the superiority of deepBICS on ChIP-seq data sets but also suggested deepBICS as a language model could help the classification of disease-related and neutral variants. We then built a language model-based method called deepBICS4SNV to predict the pathogenicity of single nucleotide variants. The good performance of deepBICS4SNV on 2 tests related to Mendelian disorders and viral diseases shows the sequence contextual information derived from language models can improve prediction accuracy and generalization capability. Lijun Quan, Xiaomin Chu, Xiaoyu Sun 0006, Tingfang Wu, Qiang Lyu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2022 | Identifying modifications on DNA-bound histones with joint deep learning of multiple binding sites in DNA sequenceabstractMOTIVATION: Histone modifications are epigenetic markers that impact gene expression by altering the chromatin structure or recruiting histone modifiers. Their accurate identification is key to unraveling the mechanisms by which they regulate gene expression. However, the solutions for this task can be improved by exploiting multiple relationships from dataset and exploring designs of learning models, for example jointly learning technology. RESULTS: This article proposes a deep learning-based multi-objective computational approach, iHMnBS, to identify which of the seven typical histone modifications a DNA sequence may choose to bind, and which parts of the DNA sequence bind to them. iHMnBS employs a customized dataset that allows the marking of modifications contained in histones that may bind to any position in the DNA sequence. iHMnBS tries to mine the information implicit in this richer data by means of deep neural networks. In comprehensive comparisons, iHMnBS outperforms a baseline method, and the probability of binding to modified histones assigned to a representative nucleotide of a DNA sequence can serve as a reference for biological experiments. Since the interaction between transcription factors and histone modifications has an important role in gene expression, we extracted a number of sequence patterns that may bind to transcription factors, and explored their possible impact on disease. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/lennylv/iHMnBS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lijun Quan, Yiting Zhou, Yelu Jiang, Tingfang Wu, Qiang Lyu |
Bioinform. | 2 |
| 2022 | TransPPMP: predicting pathogenicity of frameshift and non-sense mutations by a Transformer based on protein featuresabstractMOTIVATION: Protein structure can be severely disrupted by frameshift and non-sense mutations at specific positions in the protein sequence. Frameshift and non-sense mutation cases can also be found in healthy individuals. A method to distinguish neutral and potentially disease-associated frameshift and non-sense mutations is of practical and fundamental importance. It would allow researchers to rapidly screen out the potentially pathogenic sites from a large number of mutated genes and then use these sites as drug targets to speed up diagnosis and improve access to treatment. The problem of how to distinguish between neutral and potentially disease-associated frameshift and non-sense mutations remains under-researched. RESULTS: We built a Transformer-based neural network model to predict the pathogenicity of frameshift and non-sense mutations on protein features and named it TransPPMP. The feature matrix of contextual sequences computed by the ESM pre-training model, type of mutation residue and the auxiliary features, including structure and function information, are combined as input features, and the focal loss function is designed to solve the sample imbalance problem during the training. In 10-fold cross-validation and independent blind test set, TransPPMP showed good robust performance and absolute advantages in all evaluation metrics compared with four other advanced methods, namely, ENTPRISE-X, VEST-indel, DDIG-in and CADD. In addition, we demonstrate the usefulness of the multi-head attention mechanism in Transformer to predict the pathogenicity of mutations-not only can multiple self-attention heads learn local and global interactions but also functional sites with a large influence on the mutated residue can be captured by attention focus. These could offer useful clues to study the pathogenicity mechanism of human complex diseases for which traditional machine learning methods fall short. AVAILABILITY AND IMPLEMENTATION: TransPPMP is available at https://github.com/lennylv/TransPPMP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Liangpeng Nie, Lijun Quan, Tingfang Wu, Ruji He, Qiang Lyu |
Bioinform. | 2 |
| 2022 | Learning Useful Representations of DNA Sequences From ChIP-Seq Datasets for Exploring Transcription Factor Binding SpecificitiesabstractDeep learning has been successfully applied to surprisingly different domains. Researchers and practitioners are employing trained deep learning models to enrich our knowledge. Transcription factors (TFs)are essential for regulating gene expression in all organisms by binding to specific DNA sequences. Here, we designed a deep learning model named SemanticCS (Semantic ChIP-seq)to predict TF binding specificities. We trained our learning model on an ensemble of ChIP-seq datasets (Multi-TF-cell)to learn useful intermediate features across multiple TFs and cells. To interpret these feature vectors, visualization analysis was used. Our results indicate that these learned representations can be used to train shallow machines for other tasks. Using diverse experimental data and evaluation metrics, we show that SemanticCS outperforms other popular methods. In addition, from experimental data, SemanticCS can help to identify the substitutions that cause regulatory abnormalities and to evaluate the effect of substitutions on the binding affinity for the RXR transcription factor. The online server for SemanticCS is freely available at http://qianglab.scst.suda.edu.cn/semanticCS/. Lijun Quan, Xiaoyu Sun 0006, Liqun Huang, Ruji He, Liangpeng Nie, Yu Chen 0064, Qiang Lyu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | Quantifying Intensities of Transcription Factor-DNA Binding by Learning From an Ensemble of Protein Binding MicroarraysabstractThe control of the coordinated expression of genes is primarily regulated by the interactions between transcription factors (TFs) and their DNA binding sites, which are an integral part of transcriptional regulatory networks. There are many computational tools focused on determining TF binding or unbinding to a DNA sequence. However, other tools focused on further determining the relative preference of such binding are needed. Here, we propose a regression model with deep learning, called SemanticBI, to predict intensities of TF-DNA binding. SemanticBI is a convolutional neural network (CNN)-recurrent neural network (RNN) architecture model that was trained on an ensemble of protein binding microarray data sets that covered multiple TFs. Using this approach, SemanticBI exhibited superior accuracy in predicting binding intensities compared to other popular methods. Moreover, SemanticBI uncovered vectorized sequence-oriented features using its CNN-RNN architecture, which is an abstract representation of the original DNA sequences. Additionally, the use of SemanticBI raises the question of whether motifs are necessary for computational models of TF binding. The online SemanticBI service can be accessed at http://qianglab.scst.suda.edu.cn/semantic/. Lijun Quan, Ruji He, Xiaoyu Sun 0006, Liangpeng Nie, Qiang Lyu |
IEEE J. Biomed. Health Informatics | 1 |
| 2020 | Developing parallel ant colonies filtered by deep learned constrains for predicting RNA secondary structure with pseudo-knots
Lijun Quan, Leixin Cai, Yu Chen 0064, Xiaoyu Sun 0006, Qiang Lyu |
Neurocomputing | 1 |
| 2019 | Sequence-based prediction of protein-protein interaction sites by simplified long short-term memory network
Buzhong Zhang, Jinyan Li 0001, Lijun Quan, Yu Chen 0064 |
Neurocomputing | 3 |
| 2016 | STRUM: structure-based prediction of protein stability changes upon single-point mutationabstractMOTIVATION: Mutations in human genome are mainly through single nucleotide polymorphism, some of which can affect stability and function of proteins, causing human diseases. Several methods have been proposed to predict the effect of mutations on protein stability; but most require features from experimental structure. Given the fast progress in protein structure prediction, this work explores the possibility to improve the mutation-induced stability change prediction using low-resolution structure modeling. RESULTS: We developed a new method (STRUM) for predicting stability change caused by single-point mutations. Starting from wild-type sequences, 3D models are constructed by the iterative threading assembly refinement (I-TASSER) simulations, where physics- and knowledge-based energy functions are derived on the I-TASSER models and used to train STRUM models through gradient boosting regression. STRUM was assessed by 5-fold cross validation on 3421 experimentally determined mutations from 150 proteins. The Pearson correlation coefficient (PCC) between predicted and measured changes of Gibbs free-energy gap, ΔΔG, upon mutation reaches 0.79 with a root-mean-square error 1.2 kcal/mol in the mutation-based cross-validations. The PCC reduces if separating training and test mutations from non-homologous proteins, which reflects inherent correlations in the current mutation sample. Nevertheless, the results significantly outperform other state-of-the-art methods, including those built on experimental protein structures. Detailed analyses show that the most sensitive features in STRUM are the physics-based energy terms on I-TASSER models and the conservation scores from multiple-threading template alignments. However, the ΔΔG prediction accuracy has only a marginal dependence on the accuracy of protein structure models as long as the global fold is correct. These data demonstrate the feasibility to use low-resolution structure modeling for high-accuracy stability change prediction upon point mutations. AVAILABILITY AND IMPLEMENTATION: http://zhanglab.ccmb.med.umich.edu/STRUM/ CONTACT: [email protected] and [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lijun Quan, Yang Zhang 0040 |
Bioinform. | 1 |
| 2014 | Improved packing of protein side chains with parallel ant coloniesabstractINTRODUCTION: The accurate packing of protein side chains is important for many computational biology problems, such as ab initio protein structure prediction, homology modelling, and protein design and ligand docking applications. Many of existing solutions are modelled as a computational optimisation problem. As well as the design of search algorithms, most solutions suffer from an inaccurate energy function for judging whether a prediction is good or bad. Even if the search has found the lowest energy, there is no certainty of obtaining the protein structures with correct side chains. METHODS: We present a side-chain modelling method, pacoPacker, which uses a parallel ant colony optimisation strategy based on sharing a single pheromone matrix. This parallel approach combines different sources of energy functions and generates protein side-chain conformations with the lowest energies jointly determined by the various energy functions. We further optimised the selected rotamers to construct subrotamer by rotamer minimisation, which reasonably improved the discreteness of the rotamer library. RESULTS: We focused on improving the accuracy of side-chain conformation prediction. For a testing set of 442 proteins, 87.19% of X1 and 77.11% of X12 angles were predicted correctly within 40° of the X-ray positions. We compared the accuracy of pacoPacker with state-of-the-art methods, such as CIS-RR and SCWRL4. We analysed the results from different perspectives, in terms of protein chain and individual residues. In this comprehensive benchmark testing, 51.5% of proteins within a length of 400 amino acids predicted by pacoPacker were superior to the results of CIS-RR and SCWRL4 simultaneously. Finally, we also showed the advantage of using the subrotamers strategy. All results confirmed that our parallel approach is competitive to state-of-the-art solutions for packing side chains. CONCLUSIONS: This parallel approach combines various sources of searching intelligence and energy functions to pack protein side chains. It provides a frame-work for combining different inaccuracy/usefulness objective functions by designing parallel heuristic search algorithms. Lijun Quan, Haiou Li, Xiaoyan Xia, Hongjie Wu |
BMC Bioinform. | 1 |
| 2013 | A protein-peptide docking program with modeling receptor flexible areasabstractPredicting the structure of protein-peptide complexes using computational approaches is a difficult problem whose major challenges are properly dealing with molecular flexibility and conformational changes both of the receptor and ligand. Although significant improvements have been achieved in the modeling of side chains, methods for the backbone flexibility in docking still need improvement. In this study a new method is presented for docking peptide into receptor in a full flexible docking manner. It is a parallel approach that combines all the processes during the docking of a folding peptide with a flexible receptor. Haiou Li, Lijun Quan, Xiaoyan Xia |
BIBM | 4 |
| 2013 | Packing protein side-chains by parallel ant coloniesabstractSide-chains are crucial for proteins expressing their biochemical characteristics. Packing protein side-chains is then a necessary task for protein structure prediction, and critical to some descendant and important applications, such as protein design, docking and point mutation analysis. Given all possible candidate rotamers for each residue of protein backbone, packing protein side-chains can be modeled as a combinatorial optimization problem without an accurate energy function. This paper presents a parallel approach, pacoPacker, to pack protein side-chains by ant colony optimization. Each ant colony is used to pack side-chains with the guidance of an energy function. Different colonies use different energy functions. These multiple colonies are running in parallel and cooperate with each other by sharing the pheromone matrix whose role is to tune sampling the rotamer library. In this way, the intelligences embedded in different energy functions can be brought together to find out the best side-chains for the protein backbone. Experimental study has been conducted on two typical benchmarks, and the results show that pacoPacker is competitive to the state-of-art systems. Lijun Quan, Haiou Li, Xiaoyan Xia |
BIBM | 1 |