VLDB 2026 Research / reviewers in the wild / expert
Qi Liu 0019
dblp:95/2446-19
· DBLP profile ↗
32ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0003-2578-1221ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 29 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ProTCR: a protein language model-driven framework for decoding TCR-antigen recognition toward precision immunotherapiesabstractThe ability of T-cell receptors (TCRs) to recognize neoantigens is fundamental to the initiation and maintenance of adaptive immune responses. In TCR-based immunotherapies, elucidating the recognition patterns of TCRs for peptides and accurately identifying therapeutically relevant TCR-peptide pairs remain critical challenges. Here, we present a novel dual-pathway network model, ProTCR, which integrates the protein language model ProtT5 with deep learning methods. By incorporating both global and local feature extraction mechanisms, ProTCR enables efficient representation of amino acid sequences, thereby enhancing the model's generalizability across diverse data distributions and improving its biological interpretability. ProTCR demonstrates robust performance and broad applicability across various datasets, including neoantigens, previously unseen peptides, and MHC class II-restricted epitopes, overcoming the reliance on known peptide-TCR pairs observed in previous studies. It also offers new insights for predicting diverse classes of antigenic peptides. We applied ProTCR to several clinically relevant scenarios, including immunotherapeutic target identification in acute myeloid leukemia, neoantigen-targeted immunotherapy in solid tumours, and antigen-specific T cell recognition against pathogens such as influenza and severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). Across these complex settings, ProTCR consistently maintained high accuracy and stability, demonstrating strong cross-task adaptability and broad potential for clinical application. This work not only provides a powerful tool for elucidating immune response mechanisms but also offers a solid computational foundation for the design of neoantigen or TCR based precision immunotherapy strategies. Minrui Xu, Manman Lu, Siwen Zhang, Lanming Chen, Qi Liu 0019, Lu Xie |
Briefings Bioinform. | 6 |
| 2025 | PrimeNet: rational design of Prime editing pegRNAs by deep learningabstractThe rapid development of gene editing technology has revolutionized life science research and biotechnology applications. Prime editing, a precise gene editing tool, has shown promise in various applications, including disease research and therapeutic interventions. However, its suboptimal editing efficiency for extensive fragments and lack of predictive models have hindered its widespread adoption. Existing models exhibit low prediction accuracy and limitations, such as neglecting epigenetic factors that impact gene editing effects. To address these challenges, we developed PrimeNet, a novel prediction model that integrates significant epigenetic factors, including chromatin accessibility and DNA methylation. By incorporating data from multiple cell lines and introducing multiscale convolution and attention mechanisms, PrimeNet enhances the accuracy of predictions and generalization performance. Our results show that PrimeNet achieves a Spearman correlation coefficient of 0.94 and 0.82 on two datasets originated from HEK293T and K562 cell lines, respectively, outperforming existing models. This novel model has the potential to guide experimental design, enhance the success rate of gene editing, and reduce unnecessary experimental costs, thereby advancing the application of gene editing technology in genetic disease treatment and related fields. Xichen Liao, Qi Liu 0019, Guohui Chuai |
Briefings Bioinform. | 2 |
| 2025 | Fusing Micro- and Macro-Scale Information to Predict Anticancer Synergistic Drug CombinationsabstractDrug combination therapy is highly regarded in cancer treatment. Computational methods offer a time- and cost-effective opportunity to explore the vast combination space. Although deep learning-based prediction methods lead the field, their generalization ability remains unsatisfactory. Few previous studies have the ability to finely characterize drugs and cell lines at both the micro-scale and macro-scale. Furthermore, the interaction of cross-scale information is often overlooked. These two points limit models' ability of predicting the synergism of drug combinations in cell lines. To address the issues, we propose a novel anticancer synergistic drug combination prediction method termed MMFSynergy in this article. The construction of MMFSynergy involves three phases. First, MMFSynergy pretrains two micro encoders and a macro graph encoder, which can capture micro- or macro-scale information from large volumes of unlabeled data and generate generic features for drugs and proteins. Second, it represents drugs and proteins by fusing cross-scale information through a self-supervised task. Finally, it employs a Transformer Encoder-based model to predict synergy scores, taking representations of drugs in the combinations and the associated proteins of cell lines as input. We compared our method with eight advanced methods across three typical scenarios based on two public datasets. The results consistently demonstrated that the proposed method's generalization ability outperforms six advanced methods'. We also conducted experiments including but not limited to ablation study and case study to further exhibit the effectiveness of MMFSynergy. Xiaowen Wang 0003, Hongming Zhu, Qi Liu 0019, Qin Liu 0004 |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Computational prediction and characterization of cell-type-specific and shared binding sitesabstractMOTIVATION: Cell-type-specific gene expression is maintained in large part by transcription factors (TFs) selectively binding to distinct sets of sites in different cell types. Recent research works have provided evidence that such cell-type-specific binding is determined by TF's intrinsic sequence preferences, cooperative interactions with co-factors, cell-type-specific chromatin landscapes and 3D chromatin interactions. However, computational prediction and characterization of cell-type-specific and shared binding sites is rarely studied. RESULTS: In this article, we propose two computational approaches for predicting and characterizing cell-type-specific and shared binding sites by integrating multiple types of features, in which one is based on XGBoost and another is based on convolutional neural network (CNN). To validate the performance of our proposed approaches, ChIP-seq datasets of 10 binding factors were collected from the GM12878 (lymphoblastoid) and K562 (erythroleukemic) human hematopoietic cell lines, each of which was further categorized into cell-type-specific (GM12878- and K562-specific) and shared binding sites. Then, multiple types of features for these binding sites were integrated to train the XGBoost- and CNN-based models. Experimental results show that our proposed approaches significantly outperform other competing methods on three classification tasks. Moreover, we identified independent feature contributions for cell-type-specific and shared sites through SHAP values and explored the ability of the CNN-based model to predict cell-type-specific and shared binding sites by excluding or including DNase signals. Furthermore, we investigated the generalization ability of our proposed approaches to different binding factors in the same cellular environment. AVAILABILITY AND IMPLEMENTATION: The source code is available at: https://github.com/turningpoint1988/CSSBS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qinhu Zhang, Pengrui Teng, Siguo Wang, Zhenghao Guo, Chang-an Yuan 0001, Qi Liu 0019, De-Shuang Huang |
Bioinform. | 9 |
| 2023 | Predicting anticancer synergistic drug combinations based on multi-task learningabstractBACKGROUND: The discovery of anticancer drug combinations is a crucial work of anticancer treatment. In recent years, pre-screening drug combinations with synergistic effects in a large-scale search space adopting computational methods, especially deep learning methods, is increasingly popular with researchers. Although achievements have been made to predict anticancer synergistic drug combinations based on deep learning, the application of multi-task learning in this field is relatively rare. The successful practice of multi-task learning in various fields shows that it can effectively learn multiple tasks jointly and improve the performance of all the tasks. METHODS: In this paper, we propose MTLSynergy which is based on multi-task learning and deep neural networks to predict synergistic anticancer drug combinations. It simultaneously learns two crucial prediction tasks in anticancer treatment, which are synergy prediction of drug combinations and sensitivity prediction of monotherapy. And MTLSynergy integrates the classification and regression of prediction tasks into the same model. Moreover, autoencoders are employed to reduce the dimensions of input features. RESULTS: Compared with the previous methods listed in this paper, MTLSynergy achieves the lowest mean square error of 216.47 and the highest Pearson correlation coefficient of 0.76 on the drug synergy prediction task. On the corresponding classification task, the area under the receiver operator characteristics curve and the area under the precision-recall curve are 0.90 and 0.62, respectively, which are equivalent to the comparison methods. Through the ablation study, we verify that multi-task learning and autoencoder both have a positive effect on prediction performance. In addition, the prediction results of MTLSynergy in many cases are also consistent with previous studies. CONCLUSION: Our study suggests that multi-task learning is significantly beneficial for both drug synergy prediction and monotherapy sensitivity prediction when combining these two tasks into one model. The ability of MTLSynergy to discover new anticancer synergistic drug combinations noteworthily outperforms other state-of-the-art methods. MTLSynergy promises to be a powerful tool to pre-screen anticancer synergistic drug combinations. Danyi Chen, Xiaowen Wang 0003, Hongming Zhu, Yizhi Jiang, Qi Liu 0019, Qin Liu 0004 |
BMC Bioinform. | 6 |
| 2022 | PRODeepSyn: predicting anticancer synergistic drug combinations by embedding cell lines with protein-protein interaction networkabstractAlthough drug combinations in cancer treatment appear to be a promising therapeutic strategy with respect to monotherapy, it is arduous to discover new synergistic drug combinations due to the combinatorial explosion. Deep learning technology holds immense promise for better prediction of in vitro synergistic drug combinations for certain cell lines. In methods applying such technology, omics data are widely adopted to construct cell line features. However, biological network data are rarely considered yet, which is worthy of in-depth study. In this study, we propose a novel deep learning method, termed PRODeepSyn, for predicting anticancer synergistic drug combinations. By leveraging the Graph Convolutional Network, PRODeepSyn integrates the protein-protein interaction (PPI) network with omics data to construct low-dimensional dense embeddings for cell lines. PRODeepSyn then builds a deep neural network with the Batch Normalization mechanism to predict synergy scores using the cell line embeddings and drug features. PRODeepSyn achieves the lowest root mean square error of 15.08 and the highest Pearson correlation coefficient of 0.75, outperforming two deep learning methods and four machine learning methods. On the classification task, PRODeepSyn achieves an area under the receiver operator characteristics curve of 0.90, an area under the precision-recall curve of 0.63 and a Cohen's Kappa of 0.53. In the ablation study, we find that using the multi-omics data and the integrated PPI network's information both can improve the prediction results. Additionally, the case study demonstrates the consistency between PRODeepSyn and previous studies. Xiaowen Wang 0003, Hongming Zhu, Yizhi Jiang, Yunjie Li, Qi Liu 0019, Qin Liu 0004 |
Briefings Bioinform. | 8 |
| 2022 | iCRISEE: an integrative analysis of CRISPR screen by reducing false positive hitsabstractClustered regularly interspaced short palindromic repeats associated protein 9 (CRISPR/Cas9) technology has become a popular tool for the study of genome function, and the use of this technology can achieve large-scale screening studies of specific phenotypes. Several analysis tools for CRISPR/Cas9 screening data have been developed, while high false positive rate remains a great challenge. To this end, we developed iCRISEE, an integrative analysis of CRISPR ScrEEn by reducing false positive hits. iCRISEE can dramatically reduce false positive hits and it is robust to different single guide RNA (sgRNA) library by introducing precise data filter and normalization, model selection and valid sgRNA number correction in data preprocessing, sgRNA ranking and gene ranking. Furthermore, a powerful web server has been presented to automatically complete the whole CRISPR/Cas9 screening analysis, where we integrated the main hypothesis in multiple algorithms as a full workflow, including quality control, sgRNA extracting, sgRNA alignment, sgRNA ranking, gene ranking and pathway enrichment. In addition, output of iCRISEE, including result mapping, sample clustering, sgRNA ranking and gene ranking, can be easily visualized and downloaded for publication. Taking together, iCRISEE presents to be the state-of-the-art and user-friendly tool for CRISPR screening data analysis. iCRISEE is available at https://www.icrisee.com. Tengbo Zhang, Yaxu Li, Linjun Weng, Jiali Zhu, Jieling Qin, Qi Liu 0019 |
Briefings Bioinform. | 8 |
| 2022 | Base-resolution prediction of transcription factor binding signals by a deep learning frameworkabstractTranscription factors (TFs) play an important role in regulating gene expression, thus the identification of the sites bound by them has become a fundamental step for molecular and cellular biology. In this paper, we developed a deep learning framework leveraging existing fully convolutional neural networks (FCN) to predict TF-DNA binding signals at the base-resolution level (named as FCNsignal). The proposed FCNsignal can simultaneously achieve the following tasks: (i) modeling the base-resolution signals of binding regions; (ii) discriminating binding or non-binding regions; (iii) locating TF-DNA binding regions; (iv) predicting binding motifs. Besides, FCNsignal can also be used to predict opening regions across the whole genome. The experimental results on 53 TF ChIP-seq datasets and 6 chromatin accessibility ATAC-seq datasets show that our proposed framework outperforms some existing state-of-the-art methods. In addition, we explored to use the trained FCNsignal to locate all potential TF-DNA binding regions on a whole chromosome and predict DNA sequences of arbitrary length, and the results show that our framework can find most of the known binding regions and accept sequences of arbitrary length. Furthermore, we demonstrated the potential ability of our framework in discovering causal disease-associated single-nucleotide polymorphisms (SNPs) through a series of experiments. Qinhu Zhang, Siguo Wang, Zhen-Hao Guo, Qi Liu 0019, De-Shuang Huang |
PLoS Comput. Biol. | 7 |
| 2021 | iDMer: an integrative and mechanism-driven response system for identifying compound interventions for sudden virus outbreakabstractEmerging viral infections seriously threaten human health globally. Several challenges exist in identifying effective compounds against viral infections: (1) at the initial stage of a new virus outbreak, little information, except for its genome information, may be available; (2) although the identified compounds may be effective, they may be toxic in vivo and (3) cytokine release syndrome (CRS) triggered by viral infections is the primary cause of mortality. Currently, an integrative tool that takes all those aspects into consideration for identifying effective compounds to prevent viral infections is absent. In this study, we developed iDMer, as an integrative and mechanism-driven response system for addressing these challenges during the sudden virus outbreaks. iDMer comprises three mechanism-driven compound identification modules, that is, a virus-host interaction-oriented module, an autophagy-oriented module and a CRS-oriented module. As a one-stop integrative platform, iDMer incorporates compound toxicity evaluation and compound combination identification for virus treatment with clear mechanisms. iDMer was successfully tested on five viruses, including the current severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). Our results indicated that, for all five tested viruses, compounds that were reported in the literature or experimentally validated for virus treatment were enriched at the top, demonstrating the generalized effectiveness of iDMer. Finally, we demonstrated that combinations of the individual modules successfully identified combinations of compounds effective for virus intervention with clear mechanisms. Zhiting Wei, Yuli Gao, Fangliangzi Meng, Yukang Gong, Chenyu Zhu, Bin Ju, Chao Zhang 0112, Zhongmin Liu, Qi Liu 0019 |
Briefings Bioinform. | 10 |
| 2021 | Locating transcription factor binding sites by fully convolutional neural networkabstractTranscription factors (TFs) play an important role in regulating gene expression, thus identification of the regions bound by them has become a fundamental step for molecular and cellular biology. In recent years, an increasing number of deep learning (DL) based methods have been proposed for predicting TF binding sites (TFBSs) and achieved impressive prediction performance. However, these methods mainly focus on predicting the sequence specificity of TF-DNA binding, which is equivalent to a sequence-level binary classification task, and fail to identify motifs and TFBSs accurately. In this paper, we developed a fully convolutional network coupled with global average pooling (FCNA), which by contrast is equivalent to a nucleotide-level binary classification task, to roughly locate TFBSs and accurately identify motifs. Experimental results on human ChIP-seq datasets show that FCNA outperforms other competing methods significantly. Besides, we find that the regions located by FCNA can be used by motif discovery tools to further refine the prediction performance. Furthermore, we observe that FCNA can accurately identify TF-DNA binding motifs across different cell lines and infer indirect TF-DNA bindings. Qinhu Zhang, Siguo Wang, Qi Liu 0019, De-Shuang Huang |
Briefings Bioinform. | 5 |
| 2021 | FL-QSAR: a federated learning-based QSAR prototype for collaborative drug discoveryabstractMOTIVATION: Quantitative structure-activity relationship (QSAR) analysis is commonly used in drug discovery. Collaborations among pharmaceutical institutions can lead to a better performance in QSAR prediction, however, intellectual property and related financial interests remain substantially hindering inter-institutional collaborations in QSAR modeling for drug discovery. RESULTS: For the first time, we verified the feasibility of applying the horizontal federated learning (HFL), which is a recently developed collaborative and privacy-preserving learning framework to perform QSAR analysis. A prototype platform of federated-learning-based QSAR modeling for collaborative drug discovery, i.e. FL-QSAR, is presented accordingly. We first compared the HFL framework with a classic privacy-preserving computation framework, i.e. secure multiparty computation to indicate its difference from various perspective. Then we compared FL-QSAR with the public collaboration in terms of QSAR modeling. Our extensive experiments demonstrated that (i) collaboration by FL-QSAR outperforms a single client using only its private data, and (ii) collaboration by FL-QSAR achieves almost the same performance as that of collaboration via cleartext learning algorithms using all shared information. Taking together, our results indicate that FL-QSAR under the HFL framework provides an efficient solution to break the barriers between pharmaceutical institutions in QSAR modeling, therefore promote the development of collaborative and privacy-preserving drug discovery with extendable ability to other privacy-related biomedical areas. AVAILABILITY AND IMPLEMENTATION: The source codes of FL-QSAR are available on the GitHub: https://github.com/bm2-lab/FL-QSAR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shaoqi Chen, Dongyu Xue, Guohui Chuai, Qiang Yang 0001, Qi Liu 0019 |
Bioinform. | 5 |
| 2020 | Data imbalance in CRISPR off-target predictionabstractFor genome-wide CRISPR off-target cleavage sites (OTS) prediction, an important issue is data imbalance-the number of true OTS recognized by whole-genome off-target detection techniques is much smaller than that of all possible nucleotide mismatch loci, making the training of machine learning model very challenging. Therefore, computational models proposed for OTS prediction and scoring should be carefully designed and properly evaluated in order to avoid bias. In our study, two tools are taken as examples to further emphasize the data imbalance issue in CRISPR off-target prediction to achieve better sensitivity and specificity for optimized CRISPR gene editing. We would like to indicate that (1) the benchmark of CRISPR off-target prediction should be properly evaluated and not overestimated by considering data imbalance issue; (2) incorporation of efficient computational techniques (including ensemble learning and data synthesis techniques) can help to address the data imbalance issue and improve the performance of CRISPR off-target prediction. Taking together, we call for more efforts to address the data imbalance issue in CRISPR off-target prediction to facilitate clinical utility of CRISPR-based gene editing techniques. Yuli Gao, Guohui Chuai, Weichuan Yu, Shen Qu, Qi Liu 0019 |
Briefings Bioinform. | 5 |
| 2019 | MPdeep: Medical Procession with Deep Learning
Qi Liu 0019, Wenzheng Bao |
ICIC (2) | 1 |
| 2019 | Advanced Machine Learning Techniques for BioinformaticsabstractThe papers in this special section focus on the machine learning methods, and applications of these methods to computational biology. Quan Zou 0001, Qi Liu 0019 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2018 | Benchmarking CRISPR on-target sgRNA designabstractCRISPR (Clustered Regularly Interspaced Short Palindromic Repeats)-based gene editing has been widely implemented in various cell types and organisms. A major challenge in the effective application of the CRISPR system is the need to design highly efficient single-guide RNA (sgRNA) with minimal off-target cleavage. Several tools are available for sgRNA design, while limited tools were compared. In our opinion, benchmarking the performance of the available tools and indicating their applicable scenarios are important issues. Moreover, whether the reported sgRNA design rules are reproducible across different sgRNA libraries, cell types and organisms remains unclear. In our study, a systematic and unbiased benchmark of the sgRNA predicting efficacy was performed on nine representative on-target design tools, based on six benchmark data sets covering five different cell types. The benchmark study presented here provides novel quantitative insights into the available CRISPR tools. Jifang Yan, Guohui Chuai, Chi Zhou 0003, Chen-Yu Zhu, Chao Zhang 0112, Qi Liu 0019 |
Briefings Bioinform. | 10 |
| 2014 | Reconsideration of in silico siRNA design from a perspective of heterogeneous data integration: problems and solutionsabstractThe success of RNA interference (RNAi) depends on the interaction between short interference RNAs (siRNAs) and mRNAs. Design of highly efficient and specific siRNAs has become a challenging issue in applications of RNAi. Here, we present a detailed survey on the state-of-the-art siRNAs design, focusing on several key issues with the current in silico RNAi studies, including: (i) inconsistencies among the proposed guidelines for siRNAs design and the incomplete list of siRNAs features, (ii) improper integration of the heterogeneous cross-platform siRNAs data, (iii) inadequate consideration of the binding specificity of the target mRNAs and (iv) reduction in the 'off-target' effect in siRNAs design. With these considerations, the popular in silico siRNAs design rules are reexamined and several inconsistent viewpoints toward siRNAs feature identifications are clarified. In addition, novel computational models for siRNAs design using state-of-art machine learning techniques are discussed, which focus on heterogeneous data integration, joint feature selection and customized siRNAs screening toward highly specific targets. We believe that addressing such issues in siRNA study will provide new clues for further improved design of more efficient and specific siRNAs in RNAi. Qi Liu 0019, Ruixin Zhu, Ying Xu 0001 |
Briefings Bioinform. | 1 |
| 2014 | iPEAP: integrating multiple omics and genetic data for pathway enrichment analysisabstractUNLABELLED: A challenge in biodata analysis is to understand the underlying phenomena among many interactions in signaling pathways. Such study is formulated as the pathway enrichment analysis, which identifies relevant pathways functional enriched in high-throughput data. The question faced here is how to analyze different data types in a unified and integrative way by characterizing pathways that these data simultaneously reveal. To this end, we developed integrative Pathway Enrichment Analysis Platform, iPEAP, which handles transcriptomics, proteomics, metabolomics and GWAS data under a unified aggregation schema. iPEAP emphasizes on the ability to aggregate various pathway enrichment results generated in different high-throughput experiments, as well as the quantitative measurements of different ranking results, thus providing the first benchmark platform for integration, comparison and evaluation of multiple types of data and enrichment methods. AVAILABILITY AND IMPLEMENTATION: iPEAP is freely available at http://www.tongji.edu.cn/∼qiliu/ipeap.html. Haoqi Sun, Haiping Wang 0001, Ruixin Zhu, Kailin Tang, Qin Gong, Juan Cui, Qi Liu 0019 |
Bioinform. | 8 |
| 2013 | Towards a bioinformatics analysis of anti-Alzheimer's herbal medicines from a target network perspectiveabstractWith the growth of aging population all over the world, a rising incidence of Alzheimer's disease (AD) has been recently observed. In contrast to FDA-approved western drugs, herbal medicines, featured as abundant ingredients and multi-targeting, have been acknowledged with notable anti-AD effects although the mechanism of action (MOA) is unknown. Investigating the possible MOA for these herbs can not only refresh but also extend the current knowledge of AD pathogenesis. In this study, clinically tested anti-AD herbs, their ingredients as well as their corresponding target proteins were systematically reviewed together with applicable bioinformatics resources and methodologies. Based on above information and resources, we present a systematically target network analysis framework to explore the mechanism of anti-AD herb ingredients. Our results indicated that, in addition to the binding of those symptom-relieving targets as the FDA-approved drugs usually do, ingredients of anti-AD herbs also interact closely with a variety of successful therapeutic targets related to other diseases, such as inflammation, cancer and diabetes, suggesting the possible cross-talks between these complicated diseases. Furthermore, pathways of Ca(2+) equilibrium maintaining upstream of cell proliferation and inflammation were densely targeted by the anti-AD herbal ingredients with rigorous statistic evaluation. In addition to the holistic understanding of the pathogenesis of AD, the integrated network analysis on the MOA of herbal ingredients may also suggest new clues for the future disease modifying strategies. Ruixin Zhu, Kailin Tang, Jing Zhao 0003, Qi Liu 0019 |
Briefings Bioinform. | 7 |
| 2013 | Transfer across Completely Different Feature Spaces via Spectral EmbeddingabstractIn many applications, it is very expensive or time consuming to obtain a lot of labeled examples. One practically important problem is: can the labeled data from other related sources help predict the target task, even if they have 1) different feature spaces (e.g., image versus text data), 2) different data distributions, and 3) different output spaces? This paper proposes a solution and discusses the conditions where this is highly likely to produce better results. It first unifies the feature spaces of the target and source data sets by spectral embedding, even when they are with completely different feature spaces. The principle is to devise an optimization objective that preserves the original structure of the data, while at the same time, maximizes the similarity between the two. A linear projection model, as well as a nonlinear approach are derived on the basis of this principle with closed forms. Second, a judicious sample selection strategy is applied to select only those related source examples. At last, a Bayesian-based approach is applied to model the relationship between different output spaces. The three steps can bridge related heterogeneous sources in order to learn the target task. Among the 20 experiment data sets, for example, the images with wavelet-transformed-based features are used to predict another set of images whose features are constructed from color-histogram space; documents are used to help image classification, etc. By using these extracted examples from heterogeneous sources, the models can reduce the error rate by as much as 50 percent, compared with the methods using only the examples from the target task. Xiaoxiao Shi, Qi Liu 0019, Wei Fan 0001, Philip S. Yu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2012 | CBrowse: a SAM/BAM-based contig browser for transcriptome assembly visualization and analysisabstractSUMMARY: To address the impending need for exploring rapidly increased transcriptomics data generated for non-model organisms, we developed CBrowse, an AJAX-based web browser for visualizing and analyzing transcriptome assemblies and contigs. Designed in a standard three-tier architecture with a data pre-processing pipeline, CBrowse is essentially a Rich Internet Application that offers many seamlessly integrated web interfaces and allows users to navigate, sort, filter, search and visualize data smoothly. The pre-processing pipeline takes the contig sequence file in FASTA format and its relevant SAM/BAM file as the input; detects putative polymorphisms, simple sequence repeats and sequencing errors in contigs and generates image, JSON and database-compatible CSV text files that are directly utilized by different web interfaces. CBowse is a generic visualization and analysis tool that facilitates close examination of assembly quality, genetic polymorphisms, sequence repeats and/or sequencing errors in transcriptome sequencing projects. AVAILABILITY: CBrowse is distributed under the GNU General Public License, available at http://bioinfolab.muohio.edu/CBrowse/ CONTACT: [email protected] or [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Guoli Ji, Emily Schmidt, Douglas Lenox, Qi Liu 0019, Lin Liu 0001, Chun Liang |
Bioinform. | 7 |
| 2012 | Integrated QSAR study for inhibitors of hedgehog signal pathway against multiple cell lines: a collaborative filtering methodabstractBACKGROUND: The Hedgehog Signaling Pathway is one of signaling pathways that are very important to embryonic development. The participation of inhibitors in the Hedgehog Signal Pathway can control cell growth and death, and searching novel inhibitors to the functioning of the pathway are in a great demand. As the matter of fact, effective inhibitors could provide efficient therapies for a wide range of malignancies, and targeting such pathway in cells represents a promising new paradigm for cell growth and death control. Current research mainly focuses on the syntheses of the inhibitors of cyclopamine derivatives, which bind specifically to the Smo protein, and can be used for cancer therapy. While quantitatively structure-activity relationship (QSAR) studies have been performed for these compounds among different cell lines, none of them have achieved acceptable results in the prediction of activity values of new compounds. In this study, we proposed a novel collaborative QSAR model for inhibitors of the Hedgehog Signaling Pathway by integration the information from multiple cell lines. Such a model is expected to substantially improve the QSAR ability from single cell lines, and provide useful clues in developing clinically effective inhibitors and modifications of parent lead compounds for target on the Hedgehog Signaling Pathway. RESULTS: In this study, we have presented: (1) a collaborative QSAR model, which is used to integrate information among multiple cell lines to boost the QSAR results, rather than only a single cell line QSAR modeling. Our experiments have shown that the performance of our model is significantly better than single cell line QSAR methods; and (2) an efficient feature selection strategy under such collaborative environment, which can derive the commonly important features related to the entire given cell lines, while simultaneously showing their specific contributions to a specific cell-line. Based on feature selection results, we have proposed several possible chemical modifications to improve the inhibitor affinity towards multiple targets in the Hedgehog Signaling Pathway. CONCLUSIONS: Our model with the feature selection strategy presented here is efficient, robust, and flexible, and can be easily extended to model large-scale multiple cell line/QSAR data. The data and scripts for collaborative QSAR modeling are available in the Additional file 1. Dongsheng Che, Vincent Wenchen Zheng, Ruixin Zhu, Qi Liu 0019 |
BMC Bioinform. | 5 |
| 2012 | Screening of selective histone deacetylase inhibitors by proteochemometric modelingabstractBACKGROUND: Histone deacetylase (HDAC) is a novel target for the treatment of cancer and it can be classified into three classes, i.e., classes I, II, and IV. The inhibitors selectively targeting individual HDAC have been proved to be the better candidate antitumor drugs. To screen selective HDAC inhibitors, several proteochemometric (PCM) models based on different combinations of three kinds of protein descriptors, two kinds of ligand descriptors and multiplication cross-terms were constructed in our study. RESULTS: The results show that structure similarity descriptors are better than sequence similarity descriptors and geometry descriptors in the leftacterization of HDACs. Furthermore, the predictive ability was not improved by introducing the cross-terms in our models. Finally, a best PCM model based on protein structure similarity descriptors and 32-dimensional general descriptors was derived (R2 = 0.9897, Qtest2 = 0.7542), which shows a powerful ability to screen selective HDAC inhibitors. CONCLUSIONS: Our best model not only predict the activities of inhibitors for each HDAC isoform, but also screen and distinguish class-selective inhibitors and even more isoform-selective inhibitors, thus it provides a potential way to discover or design novel candidate antitumor drugs with reduced side effect. Dingfeng Wu, Qi Liu 0019, Ruixin Zhu |
BMC Bioinform. | 5 |
| 2012 | Quantitatively integrating molecular structure and bioactivity profile evidence into drug-target relationship analysisabstractBACKGROUND: Public resources of chemical compound are in a rapid growth both in quantity and the types of data-representation. To comprehensively understand the relationship between the intrinsic features of chemical compounds and protein targets is an essential task to evaluate potential protein-binding function for virtual drug screening. In previous studies, correlations were proposed between bioactivity profiles and target networks, especially when chemical structures were similar. With the lack of effective quantitative methods to uncover such correlation, it is demanding and necessary for us to integrate the information from multiple data sources to produce an comprehensive assessment of the similarity between small molecules, as well as quantitatively uncover the relationship between compounds and their targets by such integrated schema. RESULTS: In this study a multi-view based clustering algorithm was introduced to quantitatively integrate compound similarity from both bioactivity profiles and structural fingerprints. Firstly, a hierarchy clustering was performed with the fused similarity on 37 compounds curated from PubChem. Compared to clustering in a single view, the overall common target number within fused classes has been improved by using the integrated similarity, which indicated that the present multi-view based clustering is more efficient by successfully identifying clusters with its members sharing more number of common targets. Analysis in certain classes reveals that mutual complement of the two views for compound description helps to discover missing similar compound when only single view was applied. Then, a large-scale drug virtual screen was performed on 1267 compounds curated from Connectivity Map (CMap) dataset based on the fused similarity, which obtained a better ranking result compared to that of single-view. These comprehensive tests indicated that by combining different data representations; an improved assessment of target-specific compound similarity can be achieved. CONCLUSIONS: Our study presented an efficient, extendable and quantitative computational model for integration of different compound representations, and expected to provide new clues to improve the virtual drug screening from various pharmacological properties. Scripts, supplementary materials and data used in this study are publicly available at http://lifecenter.sgst.cn/fusion/. Tianlei Xu, Ruixin Zhu, Qi Liu 0019 |
BMC Bioinform. | 3 |
| 2011 | An Integrative Approach for Genomic Island Prediction in Prokaryotic Genomes
John Fazekas, Matthew Booth, Qi Liu 0019, Dongsheng Che |
ISBRA | 4 |
| 2011 | A new protein-ligand binding sites prediction method based on the integration of protein sequence conservation informationabstractBACKGROUND: Prediction of protein-ligand binding sites is an important issue for protein function annotation and structure-based drug design. Nowadays, although many computational methods for ligand-binding prediction have been developed, there is still a demanding to improve the prediction accuracy and efficiency. In addition, most of these methods are purely geometry-based, if the prediction methods improvement could be succeeded by integrating physicochemical or sequence properties of protein-ligand binding, it may also be more helpful to address the biological question in such studies. RESULTS: In our study, in order to investigate the contribution of sequence conservation in binding sites prediction and to make up the insufficiencies in purely geometry based methods, a simple yet efficient protein-binding sites prediction algorithm is presented, based on the geometry-based cavity identification integrated with sequence conservation information. Our method was compared with the other three classical tools: PocketPicker, SURFNET, and PASS, and evaluated on an existing comprehensive dataset of 210 non-redundant protein-ligand complexes. The results demonstrate that our approach correctly predicted the binding sites in 59% and 75% of cases among the TOP1 candidates and TOP3 candidates in the ranking list, respectively, which performs better than those of SURFNET and PASS, and achieves generally a slight better performance with PocketPicker. CONCLUSIONS: Our work has successfully indicated the importance of the sequence conservation information in binding sites prediction as well as provided a more accurate way for binding sites identification. Tianli Dai, Qi Liu 0019, Ruixin Zhu |
BMC Bioinform. | 2 |
| 2011 | Multi-target QSAR Modelling in the Analysis and Design of HIV-HCV Co-Inhibitors : An In-silico StudyabstractBACKGROUND: HIV and HCV infections have become the leading global public-health threats. Even more remarkable, HIV-HCV co-infection is rapidly emerging as a major cause of morbidity and mortality throughout the world, due to the common rapid mutation characteristics of the two viruses as well as their similar complex influence to immunology system. Although considerable progresses have been made on the study of the infection of HIV and HCV respectively, few researches have been conducted on the investigation of the molecular mechanism of their co-infection and designing of the multi-target co-inhibitors for the two viruses simultaneously. RESULTS: In our study, a multi-target Quantitative Structure-Activity Relationship (QSAR) study of the inhibitors for HIV-HCV co-infection were addressed with an in-silico machine learning technique, i.e. multi-task learning, to help to guide the co-inhibitor design. Firstly, an integrated dataset with 3 HIV inhibitor subsets targeted on protease, integrase and reverse transcriptase respectively, together with another 6 subsets of 2 HCV inhibitors targeted on NS3 serine protease and NS5B polymerase respectively were compiled. Secondly, an efficient multi-target QSAR modelling of HIV-HCV co-inhibitors was performed by applying an accelerated gradient method based multi-task learning on the whole 9 datasets. Furthermore, by solving the L-1-infinity regularized optimization, the Drug-like index features for compound description were ranked according to their joint importance in multi-target QSAR modelling of HIV and HCV. Finally, a drug structure-activity simulation for investigating the relationships between compound structures and binding affinities was presented based on our multiple target analysis, which is then providing several novel clues for the design of multi-target HIV-HCV co-inhibitors with increasing likelihood of successful therapies on HIV, HCV and HIV-HCV co-infection. CONCLUSIONS: The framework presented in our study provided an efficient way to identify and design inhibitors that simultaneously and selectively bind to multiple targets from multiple viruses with high affinity, and will definitely shed new lights on the future work of inhibitor synthesis for multi-target HIV, HCV, and HIV-HCV co-infection treatments. Qi Liu 0019, Lin Liu 0001, Ruixin Zhu |
BMC Bioinform. | 1 |
| 2010 | Transfer Learning on Heterogenous Feature Spaces via Spectral TransformationabstractLabeled examples are often expensive and time-consuming to obtain. One practically important problem is: can the labeled data from other related sources help predict the target task, even if they have (a) different feature spaces (e.g., image vs. text data), (b) different data distributions, and (c) different output spaces? This paper proposes a solution and discusses the conditions where this is possible and highly likely to produce better results. It works by first using spectral embedding to unify the different feature spaces of the target and source data sets, even when they have completely different feature spaces. The principle is to cast into an optimization objective that preserves the original structure of the data, while at the same time, maximizes the similarity between the two. Second, a judicious sample selection strategy is applied to select only those related source examples. At last, a Bayesian-based approach is applied to model the relationship between different output spaces. The three steps can bridge related heterogeneous sources in order to learn the target task. Among the 12 experiment data sets, for example, the images with wavelet-transformed-based features are used to predict another set of images whose features are constructed from color-histogram space. By using these extracted examples from heterogeneous sources, the models can reduce the error rate by as much as ~50\%, compared with the methods using only the examples from the target task. Xiaoxiao Shi, Qi Liu 0019, Wei Fan 0001, Philip S. Yu, Ruixin Zhu |
ICDM | 2 |
| 2010 | Predictive Modeling with Heterogeneous SourcesabstractLack of labeled training examples is a common problem for many applications. At the same time, there is often an abundance of labeled data from related tasks, although they have different distributions and outputs (e.g., different class labels, and different scales of regression values). In the medical domain, for example, we may have a limited number of vaccine efficacy examples against a new swine flu H1N1 epidemic, whereas there exists a large amount of labeled vaccine data from previous years' flu. However, it is difficult to directly apply the older flu vaccine data as training examples because of the difference in data distribution and efficacy output criteria between different viruses. To increase the sources of labeled data, we propose a method to utilize these examples whose marginal distribution and output criteria can be different. The idea is to first select a subset of source examples similar in distribution to the target data; all the selected instances are then “re-scaled” and assigned new output values from the labeled space of the target task. A new predictive model is built on the enlarged training set. We derive a generalization bound that specifically considers distribution difference and further evaluate the model on a number of applications. For an siRNA efficacy prediction problem, we extract examples from 4 heterogeneous regression tasks and 2 classification tasks to learn the target model, and achieve an average improvement of 30% in accuracy. Xiaoxiao Shi, Qi Liu 0019, Wei Fan 0001, Qiang Yang 0001, Philip S. Yu |
SDM | 2 |
| 2010 | In-silico prediction of blood-secretory human proteins using a ranking algorithmabstractBACKGROUND: Computational identification of blood-secretory proteins, especially proteins with differentially expressed genes in diseased tissues, can provide highly useful information in linking transcriptomic data to proteomic studies for targeted disease biomarker discovery in serum. RESULTS: A new algorithm for prediction of blood-secretory proteins is presented using an information-retrieval technique, called manifold ranking. On a dataset containing 305 known blood-secretory human proteins and a large number of other proteins that are either not blood-secretory or unknown, the new method performs better than the previous published method, measured in terms of the area under the recall-precision curve (AUC). A key advantage of the presented method is that it does not explicitly require a negative training set, which could often be noisy or difficult to derive for most biological problems, hence making our method more applicable than classification-based data mining methods in general biological studies. CONCLUSION: We believe that our program will prove to be very useful to biomedical researchers who are interested in finding serum markers, especially when they have candidate proteins derived through transcriptomic or proteomic analyses of diseased tissues. A computer program is developed for prediction of blood-secretory proteins based on manifold ranking, which is accessible at our website http://csbl.bmb.uga.edu/publications/materials/qiliu/blood_secretory_protein.html. Qi Liu 0019, Juan Cui, Qiang Yang 0001, Ying Xu 0001 |
BMC Bioinform. | 1 |
| 2010 | Multi-task learning for cross-platform siRNA efficacy prediction: an in-silico studyabstractBACKGROUND: Gene silencing using exogenous small interfering RNAs (siRNAs) is now a widespread molecular tool for gene functional study and new-drug target identification. The key mechanism in this technique is to design efficient siRNAs that incorporated into the RNA-induced silencing complexes (RISC) to bind and interact with the mRNA targets to repress their translations to proteins. Although considerable progress has been made in the computational analysis of siRNA binding efficacy, few joint analysis of different RNAi experiments conducted under different experimental scenarios has been done in research so far, while the joint analysis is an important issue in cross-platform siRNA efficacy prediction. A collective analysis of RNAi mechanisms for different datasets and experimental conditions can often provide new clues on the design of potent siRNAs. RESULTS: An elegant multi-task learning paradigm for cross-platform siRNA efficacy prediction is proposed. Experimental studies were performed on a large dataset of siRNA sequences which encompass several RNAi experiments recently conducted by different research groups. By using our multi-task learning method, the synergy among different experiments is exploited and an efficient multi-task predictor for siRNA efficacy prediction is obtained. The 19 most popular biological features for siRNA according to their jointly importance in multi-task learning were ranked. Furthermore, the hypothesis is validated out that the siRNA binding efficacy on different messenger RNAs(mRNAs) have different conditional distribution, thus the multi-task learning can be conducted by viewing tasks at an "mRNA"-level rather than at the "experiment"-level. Such distribution diversity derived from siRNAs bound to different mRNAs help indicate that the properties of target mRNA have important implications on the siRNA binding efficacy. CONCLUSIONS: The knowledge gained from our study provides useful insights on how to analyze various cross-platform RNAi data for uncovering of their complex mechanism. Qi Liu 0019, Qian Xu 0005, Vincent Wenchen Zheng, Hong Xue 0001, Qiang Yang 0001 |
BMC Bioinform. | 1 |
| 2009 | Analyses of domains and domain fusions in human proto-oncogenesabstractBACKGROUND: Understanding the constituent domains of oncogenes, their origins and their fusions may shed new light about the initiation and the development of cancers. RESULTS: We have developed a computational pipeline for identification of functional domains of human genes, prediction of the origins of these domains and their major fusion events during evolution through integration of existing and new tools of our own. An application of the pipeline to 124 well-characterized human oncogenes has led to the identification of a collection of domains and domain pairs that occur substantially more frequently in oncogenes than in human genes on average. Most of these enriched domains and domain pairs are related to tyrosine kinase activities. In addition, our analyses indicate that a substantial portion of the domain-fusion events of oncogenes took place in metazoans during evolution. CONCLUSION: We expect that the computational pipeline for domain identification, domain origin and domain fusion prediction will prove to be useful for studying other groups of genes. Qi Liu 0019, Jinling Huang, Huiqing Liu, Ping Wan, Xiuzi Ye, Ying Xu 0001 |
BMC Bioinform. | 1 |
| 2008 | Computational prediction of human proteins that can be secreted into the bloodstreamabstractWe present a novel computational method for predicting which proteins from highly and abnormally expressed genes in diseased human tissues, such as cancers, can be secreted into the bloodstream, suggesting possible marker proteins for follow-up serum proteomic studies. A main challenging issue in tackling this problem is that our understanding about the downstream localization after proteins are secreted outside the cells is very limited and not sufficient to provide useful hints about secretion to the bloodstream. To bypass this difficulty, we have taken a data mining approach by first collecting, through extensive literature searches, human proteins that are known to be secreted into the bloodstream due to various pathological conditions as detected by previous proteomic studies, and then asking the question: 'what do these secreted proteins have in common in terms of their physical and chemical properties, amino acid sequence and structural features that can be used to predict them?' We have identified a list of features, such as signal peptides, transmembrane domains, glycosylation sites, disordered regions, secondary structural content, hydrophobicity and polarity measures that show relevance to protein secretion. Using these features, we have trained a support vector machine-based classifier to predict protein secretion to the bloodstream. On a large test set containing 98 secretory proteins and 6601 non-secretory proteins of human, our classifier achieved approximately 90% prediction sensitivity and approximately 98% prediction specificity. Several additional datasets are used to further assess the performance of our classifier. On a set of 122 proteins that were found to be of abnormally high abundance in human blood due to various cancers, our program predicted 62 as blood-secreted proteins. By applying our program to abnormally highly expressed genes in gastric cancer and lung cancer tissues detected through microarray gene expression studies, we predicted 13 and 31 as blood secreted, respectively, suggesting that they could serve as potential biomarkers for these two cancers, respectively. Our study demonstrated that our method can provide highly useful information to link genomic and proteomic studies for disease biomarker discovery. Our software can be accessed at http://csbl1.bmb.uga.edu/cgi-bin/Secretion/secretion.cgi. Juan Cui, Qi Liu 0019, David Puett, Ying Xu 0001 |
Bioinform. | 2 |