Qiang Lyu

dblp:163/8033 · DBLP profile ↗
← Back
24ranked-venue papers
1as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 13 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Identifying batch-integrated domains from spatial transcriptomics via graph autoencoder with contrastive learning based on cross-modality and data augmentation
abstract
Spatially resolved transcriptomics (SRT) allows for the comprehensive profiling of gene expression while preserving spatial context, advancing the study of tissue architecture. However, existing computational approaches still face key limitations, particularly the insufficient exploitation of histology information and the lack of cross-modal meaningful contrastive strategies for biological analyses. To overcome these challenges, we propose GCAST, a graph contrastive autoencoder framework for spatial transcriptomics that seamlessly integrates multimodal SRT data. GCAST adopts a self-supervised strategy to derive biologically meaningful representations directly from histology images when available. GCAST constructs dual graph views based on data augmentation and introduces a novel contrastive learning designed to leverage histology-weighted and gene-weighted features and improve biological interpretability. In addition, GCAST employs a block-diagonal graph construction to automatically align multiple datasets, achieving batch-effect correction without manual intervention. The framework not only captures spatial gene expression patterns to identify tissue domains but also adapts to datasets with or without histological images and supports the integration of multiple datasets for joint analyses. Overall, GCAST provides a unified and biologically informed framework that has the potential to facilitate deeper analyses of spatial transcriptomics.
Yexuan Mao, Lijun Quan, Guozheng Zhang, Yelu Jiang, Liangpeng Nie, Tingfang Wu, Lingkun Meng, Qiang Lyu
Briefings Bioinform.11
2025 DS-MVP: identifying disease-specific pathogenicity of missense variants by pre-training representation
abstract
Accurately predicting the pathogenicity of missense variants is crucial for improving disease diagnosis and advancing clinical research. However, existing computational methods primarily focus on general pathogenicity predictions, overlooking assessments of disease-specific conditions. In this study, we propose DS-MVP, a method capable of predicting disease-specific pathogenicity of missense variants in human genomes. DS-MVP first leverages a deep learning model pre-trained on a large general pathogenicity dataset to learn rich representation of missense variants. It then fine-tunes these representations with an XGBoost model on smaller datasets for specific diseases. We evaluated the learned representation by testing it on multiple binary pathogenicity datasets and gene-level statistics, demonstrating that DS-MVP outperforms existing state-of-the-art methods, such as MetaRNN and AlphaMissense. Additionally, DS-MVP excels in multi-label and multi-class classification, effectively classifying disease-specific pathogenic missense variants based on disease conditions. It further enhances predictions by fine-tuning the pre-trained model on disease-specific datasets. Finally, we analyzed the contributions of the pre-trained model and various feature types, with gene description corpus features from large language model and genetic feature fusion contributing the most. These results underscore that DS-MVP represents a broader perspective on pathogenicity prediction and holds potential as an effective tool for disease diagnosis.
Qiufeng Chen, Lijun Quan, Lexin Cao, Liangchen Peng, Yelu Jiang, Liangpeng Nie, Tingfang Wu, Qiang Lyu
Briefings Bioinform.12
2025 Fine-grained recognition of citrus varieties via wavelet channel attention network
Fukai Zhang, Xiao-Bo Jin, Shan An, Qiang Lyu
Knowl. Based Syst.7
2024 RPEMHC: improved prediction of MHC-peptide binding affinity by a deep learning approach based on residue-residue pair encoding
abstract
MOTIVATION: Binding of peptides to major histocompatibility complex (MHC) molecules plays a crucial role in triggering T cell recognition mechanisms essential for immune response. Accurate prediction of MHC-peptide binding is vital for the development of cancer therapeutic vaccines. While recent deep learning-based methods have achieved significant performance in predicting MHC-peptide binding affinity, most of them separately encode MHC molecules and peptides as inputs, potentially overlooking critical interaction information between the two. RESULTS: In this work, we propose RPEMHC, a new deep learning approach based on residue-residue pair encoding to predict the binding affinity between peptides and MHC, which encode an MHC molecule and a peptide as a residue-residue pair map. We evaluate the performance of RPEMHC on various MHC-II-related datasets for MHC-peptide binding prediction, demonstrating that RPEMHC achieves better or comparable performance against other state-of-the-art baselines. Moreover, we further construct experiments on MHC-I-related datasets, and experimental results demonstrate that our method can work on both two MHC classes. These extensive validations have manifested that RPEMHC is an effective tool for studying MHC-peptide interactions and can potentially facilitate the vaccine development. AVAILABILITY: The source code of the method along with trained models is freely available at https://github.com/lennylv/RPEMHC.
Tingfang Wu, Yelu Jiang, Taoning Chen, Deng Pan 0006, Jingxin Xie, Lijun Quan, Qiang Lyu
Bioinform.9
2024 MultiModRLBP: A Deep Learning Approach for Multi-Modal RNA-Small Molecule Ligand Binding Sites Prediction
abstract
This study aims to tackle the intricate challenge of predicting RNA-small molecule binding sites to explore the potential value in the field of RNA drug targets. To address this challenge, we propose the MultiModRLBP method, which integrates multi-modal features using deep learning algorithms. These features include 3D structural properties at the nucleotide base level of the RNA molecule, relational graphs based on overall RNA structure, and rich RNA semantic information. In our investigation, we gathered 851 interactions between RNA and small molecule ligand from the RNAglib dataset and RLBind training set. Unlike conventional training sets, this collection broadened its scope by including RNA complexes that have the same RNA sequence but change their respective binding sites due to structural differences or the presence of different ligands. This enhancement enables the MultiModRLBP model to more accurately capture subtle changes at the structural level, ultimately improving its ability to discern nuances among similar RNA conformations. Furthermore, we evaluated MultiModRLBP on two classic test sets, Test18 and Test3, highlighting its performance disparities on small molecules based on metal and non-metal ions. Additionally, we conducted a structural sensitivity analysis on specific complex categories, considering RNA instances with varying degrees of structural changes and whether they share the same ligands. The research results indicate that MultiModRLBP outperforms the current state-of-the-art methods on multiple classic test sets, particularly excelling in predicting binding sites for non-metal ions and instances where the binding sites are widely distributed along the sequence. MultiModRLBP also can be used as a potential tool when the RNA structure is perturbed or the RNA experimental tertiary structure is not available. Most importantly, MultiModRLBP exhibits the capability to distinguish binding characteristics of RNA that are structurally diverse yet exhibit sequence similarity. These advancements hold promise in reducing the costs associated with the development of RNA-targeted drugs.
Lijun Quan, Hongjie Wu, Xuhao Ma, Jingxin Xie, Deng Pan 0006, Taoning Chen, Tingfang Wu, Qiang Lyu
IEEE J. Biomed. Health Informatics11
2023 Compositional Prototypical Networks for Few-Shot Classification
abstract
It is assumed that pre-training provides the feature extractor with strong class transferability and that high novel class generalization can be achieved by simply reusing the transferable feature extractor. In this work, our motivation is to explicitly learn some fine-grained and transferable meta-knowledge so that feature reusability can be further improved. Concretely, inspired by the fact that humans can use learned concepts or components to help them recognize novel classes, we propose Compositional Prototypical Networks (CPN) to learn a transferable prototype for each human-annotated attribute, which we call a component prototype. We empirically demonstrate that the learned component prototypes have good class transferability and can be reused to construct compositional prototypes for novel classes. Then a learnable weight generator is utilized to adaptively fuse the compositional and visual prototypes. Extensive experiments demonstrate that our method can achieve state-of-the-art results on different datasets and settings. The performance gains are especially remarkable in the 5-way 1-shot setting. The code is available at https://github.com/fikry102/CPN.
Qiang Lyu
AAAI1
2023 WCANet: Wavelet Channel Attention Network for Citrus Variety Identification
abstract
The effective fine-grained identification of citrus varieties plays a vital role in the differential production management of citrus orchards. To our knowledge, there are few studies and publicly available datasets on fine-grained identification of citrus varieties. In this study, we propose Wavelet Channel Attention Network (WCANet) to solve the problem of fine-grained visual classification of citrus varieties and create a Citrus Variety Dataset (CVD) consisting of tree canopy images. WCANet combines global average pooling to extract global features and wavelet transform to capture local features, which greatly improves the capability of channel attention modules for multi-scale feature extraction. Experimental results demonstrate that the WCANet outperforms the state-of-the-art confidence estimation approaches on various benchmarks. Our code and dataset will be open-sourced at https://github.com/fightero/WCANet.
Fukai Zhang, Xiao-Bo Jin, Shan An, Qiang Lyu
ICIP5
2023 CAPLA: improved prediction of protein-ligand binding affinity by a deep learning approach based on a cross-attention mechanism
abstract
MOTIVATION: Accurate and rapid prediction of protein-ligand binding affinity is a great challenge currently encountered in drug discovery. Recent advances have manifested a promising alternative in applying deep learning-based computational approaches for accurately quantifying binding affinity. The structure complementarity between protein-binding pocket and ligand has a great effect on the binding strength between a protein and a ligand, but most of existing deep learning approaches usually extracted the features of pocket and ligand by these two detached modules. RESULTS: In this work, a new deep learning approach based on the cross-attention mechanism named CAPLA was developed for improved prediction of protein-ligand binding affinity by learning features from sequence-level information of both protein and ligand. Specifically, CAPLA employs the cross-attention mechanism to capture the mutual effect of protein-binding pocket and ligand. We evaluated the performance of our proposed CAPLA on comprehensive benchmarking experiments on binding affinity prediction, demonstrating the superior performance of CAPLA over state-of-the-art baseline approaches. Moreover, we provided the interpretability for CAPLA to uncover critical functional residues that contribute most to the binding affinity through the analysis of the attention scores generated by the cross-attention mechanism. Consequently, these results indicate that CAPLA is an effective approach for binding affinity prediction and may contribute to useful help for further consequent applications. AVAILABILITY AND IMPLEMENTATION: The source code of the method along with trained models is freely available at https://github.com/lennylv/CAPLA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tingfang Wu, Taoning Chen, Deng Pan 0006, Jingxin Xie, Lijun Quan, Qiang Lyu
Bioinform.8
2023 Exploring human behavior patterns and socio-demographic factors based on American Time Use Survey
abstract
Summary As human activities in mobile environments are facing an ever‐increasing range of data, it is of great research importance to explore in depth the salient features of data that discriminate human behavior patterns and socio‐demographic factors. However, human activity behavior that consists of a series of complex spatio‐temporal activity processes is difficult to model. In this article, we develop a framework to perform activity behavior pattern mining and recognition, and the proposed framework has been applied to the American Time Use Survey to explore representative activity behavior patterns and socio‐demographic factors. The main contributions are as follows: (1) A method of activity behavior similarity is presented based on daily activities and activity sequences. (2) An activity sequence similarity algorithm with is proposed by line segment tree, greedy algorithm, and dynamic programming. (3) Representative activity behavior patterns and socio‐demographic factors are derived by clustering analysis and mining. (4) The activity behavior pattern is recognized by activity behavior features or socio‐demographic features. Through the experiments, we find that different daily activity behavior patterns are associated with specific socio‐demographic factors.
Hongxin Liu, Shunming Lyu, Xiaofei Niu, Tie Hou, Yuling Ma, Qiang Lyu
Concurr. Comput. Pract. Exp.7
2023 TransRNAm: Identifying Twelve Types of RNA Modifications by an Interpretable Multi-Label Deep Learning Model Based on Transformer
abstract
Accurate identification of RNA modification sites is of great significance in understanding the functions and regulatory mechanisms of RNAs. Recent advances have shown great promise in applying computational methods based on deep learning for accurate prediction of RNA modifications. However, those methods generally predicted only a single type of RNA modification. In addition, such methods suffered from the scarcity of the interpretability for their predicted results. In this work, a new Transformer-based deep learning method was proposed to predict multiple RNA modifications simultaneously, referred to as TransRNAm. More specifically, TransRNAm employs Transformer to extract contextual feature and convolutional neural networks to further learn high-latent feature representations of RNA sequences relevant for RNA modifications. Importantly, by integrating the self-attention mechanism in Transformer with convolutional neural network, TransRNAm is capable of not only capturing the critical nucleotide sites that contribute significantly to RNA modification prediction, but also revealing the underlying association among different types of RNA modifications. Consequently, this work provided an accurate and interpretable predictor for multiple RNA modification prediction, which may contribute to uncovering the sequence-based forming mechanism of RNA modification sites.
Taoning Chen, Tingfang Wu, Deng Pan 0006, Jinxing Xie, Lijun Quan, Qiang Lyu
IEEE ACM Trans. Comput. Biol. Bioinform.8
2023 DGCddG: Deep Graph Convolution for Predicting Protein-Protein Binding Affinity Changes Upon Mutations
abstract
Effectively and accurately predicting the effects of interactions between proteins after amino acid mutations is a key issue for understanding the mechanism of protein function and drug design. In this study, we present a deep graph convolution (DGC) network-based framework, DGCddG, to predict the changes of protein-protein binding affinity after mutation. DGCddG incorporates multi-layer graph convolution to extract a deep, contextualized representation for each residue of the protein complex structure. The mined channels of the mutation sites by DGC is then fitted to the binding affinity with a multi-layer perceptron. Experiments with results on multiple datasets show that our model can achieve relatively good performance for both single and multi-point mutations. For blind tests on datasets related to angiotensin-converting enzyme 2 binding with the SARS-CoV-2 virus, our method shows better results in predicting ACE2 changes, may help in finding favorable antibodies. Code and data availability: https://github.com/lennylv/DGCddG.
Yelu Jiang, Lijun Quan, Yiting Zhou, Tingfang Wu, Qiang Lyu
IEEE ACM Trans. Comput. Biol. Bioinform.7
2023 ctP2ISP: Protein-Protein Interaction Sites Prediction Using Convolution and Transformer With Data Augmentation
abstract
Proteinprotein interactions are the basis of many cellular biological processes, such as cellular organization, signal transduction, and immune response. Identifying proteinprotein interaction sites is essential for understanding the mechanisms of various biological processes, disease development, and drug design. However, it remains a challenging task to make accurate predictions, as the small amount of training data and severe imbalanced classification reduce the performance of computational methods. We design a deep learning method named ctP2ISP to improve the prediction of proteinprotein interaction sites. ctP2ISP employs Convolution and Transformer to extract information and enhance information perception so that semantic features can be mined to identify proteinprotein interaction sites. A weighting loss function with different sample weights is designed to suppress the preference of the model toward multi-category prediction. To efficiently reuse the information in the training set, a preprocessing of data augmentation with an improved sample-oriented sampling strategy is applied. The trained ctP2ISP was evaluated against current state-of-the-art methods on six public datasets. The results show that ctP2ISP outperforms all other competing methods on the balance metrics: F1, MCC, and AUPRC. In particular, our prediction on open tests related to viruses may also be consistent with biological insights.
Lijun Quan, Yelu Jiang, Yiting Zhou, Tingfang Wu, Qiang Lyu
IEEE ACM Trans. Comput. Biol. Bioinform.7
2023 How Deepbics Quantifies Intensities of Transcription Factor-DNA Binding and Facilitates Prediction of Single Nucleotide Variant Pathogenicity With a Deep Learning Model Trained On ChIP-Seq Data Sets
abstract
The binding of DNA sequences to cell type-specific transcription factors is essential for regulating gene expression in all organisms. Many variants occurring in these binding regions play crucial roles in human disease by disrupting the cis-regulation of gene expression. We first implemented a sequence-based deep learning model called deepBICS to quantify the intensity of transcription factors-DNA binding. The experimental results not only showed the superiority of deepBICS on ChIP-seq data sets but also suggested deepBICS as a language model could help the classification of disease-related and neutral variants. We then built a language model-based method called deepBICS4SNV to predict the pathogenicity of single nucleotide variants. The good performance of deepBICS4SNV on 2 tests related to Mendelian disorders and viral diseases shows the sequence contextual information derived from language models can improve prediction accuracy and generalization capability.
Lijun Quan, Xiaomin Chu, Xiaoyu Sun 0006, Tingfang Wu, Qiang Lyu
IEEE ACM Trans. Comput. Biol. Bioinform.5
2022 Identifying modifications on DNA-bound histones with joint deep learning of multiple binding sites in DNA sequence
abstract
MOTIVATION: Histone modifications are epigenetic markers that impact gene expression by altering the chromatin structure or recruiting histone modifiers. Their accurate identification is key to unraveling the mechanisms by which they regulate gene expression. However, the solutions for this task can be improved by exploiting multiple relationships from dataset and exploring designs of learning models, for example jointly learning technology. RESULTS: This article proposes a deep learning-based multi-objective computational approach, iHMnBS, to identify which of the seven typical histone modifications a DNA sequence may choose to bind, and which parts of the DNA sequence bind to them. iHMnBS employs a customized dataset that allows the marking of modifications contained in histones that may bind to any position in the DNA sequence. iHMnBS tries to mine the information implicit in this richer data by means of deep neural networks. In comprehensive comparisons, iHMnBS outperforms a baseline method, and the probability of binding to modified histones assigned to a representative nucleotide of a DNA sequence can serve as a reference for biological experiments. Since the interaction between transcription factors and histone modifications has an important role in gene expression, we extracted a number of sequence patterns that may bind to transcription factors, and explored their possible impact on disease. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/lennylv/iHMnBS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lijun Quan, Yiting Zhou, Yelu Jiang, Tingfang Wu, Qiang Lyu
Bioinform.7
2022 TransPPMP: predicting pathogenicity of frameshift and non-sense mutations by a Transformer based on protein features
abstract
MOTIVATION: Protein structure can be severely disrupted by frameshift and non-sense mutations at specific positions in the protein sequence. Frameshift and non-sense mutation cases can also be found in healthy individuals. A method to distinguish neutral and potentially disease-associated frameshift and non-sense mutations is of practical and fundamental importance. It would allow researchers to rapidly screen out the potentially pathogenic sites from a large number of mutated genes and then use these sites as drug targets to speed up diagnosis and improve access to treatment. The problem of how to distinguish between neutral and potentially disease-associated frameshift and non-sense mutations remains under-researched. RESULTS: We built a Transformer-based neural network model to predict the pathogenicity of frameshift and non-sense mutations on protein features and named it TransPPMP. The feature matrix of contextual sequences computed by the ESM pre-training model, type of mutation residue and the auxiliary features, including structure and function information, are combined as input features, and the focal loss function is designed to solve the sample imbalance problem during the training. In 10-fold cross-validation and independent blind test set, TransPPMP showed good robust performance and absolute advantages in all evaluation metrics compared with four other advanced methods, namely, ENTPRISE-X, VEST-indel, DDIG-in and CADD. In addition, we demonstrate the usefulness of the multi-head attention mechanism in Transformer to predict the pathogenicity of mutations-not only can multiple self-attention heads learn local and global interactions but also functional sites with a large influence on the mutated residue can be captured by attention focus. These could offer useful clues to study the pathogenicity mechanism of human complex diseases for which traditional machine learning methods fall short. AVAILABILITY AND IMPLEMENTATION: TransPPMP is available at https://github.com/lennylv/TransPPMP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Liangpeng Nie, Lijun Quan, Tingfang Wu, Ruji He, Qiang Lyu
Bioinform.5
2022 Asynchronous spiking neural P systems with local synchronization of rules
Tingfang Wu, Qiang Lyu
Inf. Sci.3
2022 Learning Useful Representations of DNA Sequences From ChIP-Seq Datasets for Exploring Transcription Factor Binding Specificities
abstract
Deep learning has been successfully applied to surprisingly different domains. Researchers and practitioners are employing trained deep learning models to enrich our knowledge. Transcription factors (TFs)are essential for regulating gene expression in all organisms by binding to specific DNA sequences. Here, we designed a deep learning model named SemanticCS (Semantic ChIP-seq)to predict TF binding specificities. We trained our learning model on an ensemble of ChIP-seq datasets (Multi-TF-cell)to learn useful intermediate features across multiple TFs and cells. To interpret these feature vectors, visualization analysis was used. Our results indicate that these learned representations can be used to train shallow machines for other tasks. Using diverse experimental data and evaluation metrics, we show that SemanticCS outperforms other popular methods. In addition, from experimental data, SemanticCS can help to identify the substitutions that cause regulatory abnormalities and to evaluate the effect of substitutions on the binding affinity for the RXR transcription factor. The online server for SemanticCS is freely available at http://qianglab.scst.suda.edu.cn/semanticCS/.
Lijun Quan, Xiaoyu Sun 0006, Liqun Huang, Ruji He, Liangpeng Nie, Yu Chen 0064, Qiang Lyu
IEEE ACM Trans. Comput. Biol. Bioinform.9
2021 Evolution-Communication Spiking Neural P Systems
abstract
Spiking neural P systems (SNP systems) are a class of distributed and parallel computation models, which are inspired by the way in which neurons process information through spikes, where the integrate-and-fire behavior of neurons and the distribution of produced spikes are achieved by spiking rules. In this work, a novel mechanism for separately describing the integrate-and-fire behavior of neurons and the distribution of produced spikes, and a novel variant of the SNP systems, named evolution-communication SNP (ECSNP) systems, is proposed. More precisely, the integrate-and-fire behavior of neurons is achieved by spike-evolution rules, and the distribution of produced spikes is achieved by spike-communication rules. Then, the computational power of ECSNP systems is examined. It is demonstrated that ECSNP systems are Turing universal as number-generating devices. Furthermore, the computational power of ECSNP systems with a restricted form, i.e. the quantity of spikes in each neuron throughout a computation does not exceed some constant, is also investigated, and it is shown that such restricted ECSNP systems can only characterize the family of semilinear number sets. These results manifest that the capacity of neurons for information storage (i.e. the quantity of spikes) has a critical impact on the ECSNP systems to achieve a desired computational power.
Tingfang Wu, Qiang Lyu, Linqiang Pan
Int. J. Neural Syst.2
2021 Quantifying Intensities of Transcription Factor-DNA Binding by Learning From an Ensemble of Protein Binding Microarrays
abstract
The control of the coordinated expression of genes is primarily regulated by the interactions between transcription factors (TFs) and their DNA binding sites, which are an integral part of transcriptional regulatory networks. There are many computational tools focused on determining TF binding or unbinding to a DNA sequence. However, other tools focused on further determining the relative preference of such binding are needed. Here, we propose a regression model with deep learning, called SemanticBI, to predict intensities of TF-DNA binding. SemanticBI is a convolutional neural network (CNN)-recurrent neural network (RNN) architecture model that was trained on an ensemble of protein binding microarray data sets that covered multiple TFs. Using this approach, SemanticBI exhibited superior accuracy in predicting binding intensities compared to other popular methods. Moreover, SemanticBI uncovered vectorized sequence-oriented features using its CNN-RNN architecture, which is an abstract representation of the original DNA sequences. Additionally, the use of SemanticBI raises the question of whether motifs are necessary for computational models of TF binding. The online SemanticBI service can be accessed at http://qianglab.scst.suda.edu.cn/semantic/.
Lijun Quan, Ruji He, Xiaoyu Sun 0006, Liangpeng Nie, Qiang Lyu
IEEE J. Biomed. Health Informatics7
2020 Constructing Node-Independent Spanning Trees in Augmented Cubes
abstract
For a network, edge/node-independent spanning trees (ISTs) can not only tolerate faulty edges/nodes, but also be used to distribute secure messages. As important node-symmetric variants of the hypercubes, the augmented cubes have received much attention from researchers. The n-dimensional augmented cube AQn is both (2n ‒ 1)-edge-connected and (2n ‒ 1)-nodeconnected (n ≢ 3), thus the well-known edge conjecture and node conjecture of ISTs are both interesting questions in AQn. So far, the edge conjecture on augmented cubes was proved to be true. However, the node conjecture on AQn is still open. In this paper, we further study the construction principle of the node-ISTs by using the double neighbors of every node in the higher dimension. We prove the existence of 2k − 1 node-ISTs rooted at node 0 in A Q n ( 00...0 ︸ n−k )(n≥k≥4) by proposing an ingenious way of construction and propose a corresponding O(NlogN) time algorithm, where N = 2k is the number of nodes in A Q n ( 00...0 ︸ n−k ) .
Baolei Cheng, Jianxi Fan, Qiang Lyu, Cheng-Kuan Lin
Fundam. Informaticae3
2020 Developing parallel ant colonies filtered by deep learned constrains for predicting RNA secondary structure with pseudo-knots
Lijun Quan, Leixin Cai, Yu Chen 0064, Xiaoyu Sun 0006, Qiang Lyu
Neurocomputing6
2018 Constructing independent spanning trees with height n on the n-dimensional crossed cube
Baolei Cheng, Jianxi Fan, Qiang Lyu, Jingya Zhou
Future Gener. Comput. Syst.3
2017 Deep learning methods for protein torsion angle prediction
abstract
BACKGROUND: Deep learning is one of the most powerful machine learning methods that has achieved the state-of-the-art performance in many domains. Since deep learning was introduced to the field of bioinformatics in 2012, it has achieved success in a number of areas such as protein residue-residue contact prediction, secondary structure prediction, and fold recognition. In this work, we developed deep learning methods to improve the prediction of torsion (dihedral) angles of proteins. RESULTS: We design four different deep learning architectures to predict protein torsion angles. The architectures including deep neural network (DNN) and deep restricted Boltzmann machine (DRBN), deep recurrent neural network (DRNN) and deep recurrent restricted Boltzmann machine (DReRBM) since the protein torsion angle prediction is a sequence related problem. In addition to existing protein features, two new features (predicted residue contact number and the error distribution of torsion angles extracted from sequence fragments) are used as input to each of the four deep learning architectures to predict phi and psi angles of protein backbone. The mean absolute error (MAE) of phi and psi angles predicted by DRNN, DReRBM, DRBM and DNN is about 20-21° and 29-30° on an independent dataset. The MAE of phi angle is comparable to the existing methods, but the MAE of psi angle is 29°, 2° lower than the existing methods. On the latest CASP12 targets, our methods also achieved the performance better than or comparable to a state-of-the art method. CONCLUSIONS: Our experiment demonstrates that deep learning is a valuable method for predicting protein torsion angles. The deep recurrent network architecture performs slightly better than deep feed-forward architecture, and the predicted residue contact number and the error distribution of torsion angles extracted from sequence fragments are useful features for improving prediction accuracy.
Haiou Li, Jie Hou 0001, Badri Adhikari, Qiang Lyu, Jianlin Cheng
BMC Bioinform.4
2017 Deep Conditional Random Field Approach to Transmembrane Topology Prediction and Application to GPCR Three-Dimensional Structure Modeling
abstract
Transmembrane proteins play important roles in cellular energy production, signal transmission, and metabolism. Many shallow machine learning methods have been applied to transmembrane topology prediction, but the performance was limited by the large size of membrane proteins and the complex biological evolution information behind the sequence. In this paper, we proposed a novel deep approach based on conditional random fields named as dCRF-TM for predicting the topology of transmembrane proteins. Conditional random fields take into account more complicated interrelation between residue labels in full-length sequence than HMM and SVM-based methods. Three widely-used datasets were employed in the benchmark. DCRF-TM had the accuracy 95 percent over helix location prediction and the accuracy 78 percent over helix number prediction. DCRF-TM demonstrated a more robust performance on large size proteins (>350 residues) against 11 state-of-the-art predictors. Further dCRF-TM was applied to ab initio modeling three-dimensional structures of seven-transmembrane receptors, also known as G protein-coupled receptors. The predictions on 24 solved G protein-coupled receptors and unsolved vasopressin V2 receptor illustrated that dCRF-TM helped abGPCR-I-TASSER to improve TM-score 34.3 percent rather than using the random transmembrane definition. Two out of five predicted models caught the experimental verified disulfide bonds in vasopressin V2 receptor.
Hongjie Wu, Kun Wang 0005, Liyao Lu, Yu Xue 0003, Qiang Lyu, Min Jiang 0009
IEEE ACM Trans. Comput. Biol. Bioinform.5