EDBT 2026 Demo / reviewers in the wild / expert
Qi Liu 0024
dblp:95/2446-24
· DBLP profile ↗
23ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0001-8892-7078ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 4Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | stImage: a versatile framework for optimizing spatial transcriptomic analysis through customizable deep histology and location informed integrationabstractSpatial transcriptomics (ST) integrates gene expression data with the spatial organization of cells and their associated histology, offering unprecedented insights into tissue biology. While existing methods incorporate either location-based or histology-informed information, none fully synergize gene expression, histological features, and precise spatial coordinates within a unified framework. Moreover, these methods often exhibit inconsistent performance across diverse datasets and conditions. Here, we introduce stImage, an open-source R package that provides a comprehensive and flexible solution for ST analysis. By generating deep learning-derived histology features and offering 54 integrative strategies, stImage seamlessly combines transcriptional profiles, histology images, and spatial information. We demonstrate stImage's effectiveness across multiple datasets, underscoring its ability to guide users toward the most suitable integration strategy using diagnostic graph. Our results highlight how stImage can optimize ST, consistently improving biological insights and advancing our understanding of tissue architecture. stImage is freely available at https://github.com/YuWang-VUMC/stImage. Haichun Yang, Ruining Deng, Yuankai Huo, Qi Liu 0024, Shyr Yu, Shilin Zhao |
Briefings Bioinform. | 5 |
| 2024 | A distribution-free and analytic method for power and sample size calculation in single-cell differential expressionabstractMOTIVATION: Differential expression analysis in single-cell transcriptomics unveils cell type-specific responses to various treatments or biological conditions. To ensure the robustness and reliability of the analysis, it is essential to have a solid experimental design with ample statistical power and sample size. However, existing methods for power and sample size calculation often assume a specific distribution for single-cell transcriptomics data, potentially deviating from the true data distribution. Moreover, they commonly overlook cell-cell correlations within individual samples, posing challenges in accurately representing biological phenomena. Additionally, due to the complexity of deriving an analytic formula, most methods employ time-consuming simulation-based strategies. RESULTS: We propose an analytic-based method named scPS for calculating power and sample sizes based on generalized estimating equations. scPS stands out by making no assumptions about the data distribution and considering cell-cell correlations within individual samples. scPS is a rapid and powerful approach for designing experiments in single-cell differential expression analysis. AVAILABILITY AND IMPLEMENTATION: scPS is freely available at https://github.com/cyhsuTN/scPS and Zenodo https://zenodo.org/records/13375996. Chih-Yuan Hsu, Qi Liu 0024, Shyr Yu |
Bioinform. | 2 |
| 2024 | scKWARN: Kernel-weighted-average robust normalization for single-cell RNA-seq dataabstractMOTIVATION: Single-cell RNA-seq normalization is an essential step to correct unwanted biases caused by sequencing depth, capture efficiency, dropout, and other technical factors. Existing normalization methods primarily reduce biases arising from sequencing depth by modeling count-depth relationship and/or assuming a specific distribution for read counts. However, these methods may lead to over or under-correction due to presence of technical biases beyond sequencing depth and the restrictive assumption on models and distributions. RESULTS: We present scKWARN, a Kernel Weighted Average Robust Normalization designed to correct known or hidden technical confounders without assuming specific data distributions or count-depth relationships. scKWARN generates a pseudo expression profile for each cell by borrowing information from its fuzzy technical neighbors through a kernel smoother. It then compares this profile against the reference derived from cells with the same bimodality patterns to determine the normalization factor. As demonstrated in both simulated and real datasets, scKWARN outperforms existing methods in removing a variety of technical biases while preserving true biological heterogeneity. AVAILABILITY AND IMPLEMENTATION: scKWARN is freely available at https://github.com/cyhsuTN/scKWARN. Chih-Yuan Hsu, Qi Liu 0024, Shyr Yu |
Bioinform. | 3 |
| 2024 | Cross-scale multi-instance learning for pathological image diagnosisabstractAnalyzing high resolution whole slide images (WSIs) with regard to information across multiple scales poses a significant challenge in digital pathology. Multi-instance learning (MIL) is a common solution for working with high resolution images by classifying bags of objects (i.e. sets of smaller image patches). However, such processing is typically performed at a single scale (e.g., 20× magnification) of WSIs, disregarding the vital inter-scale information that is key to diagnoses by human pathologists. In this study, we propose a novel cross-scale MIL algorithm to explicitly aggregate inter-scale relationships into a single MIL network for pathological image diagnosis. The contribution of this paper is three-fold: (1) A novel cross-scale MIL (CS-MIL) algorithm that integrates the multi-scale information and the inter-scale relationships is proposed; (2) A toy dataset with scale-specific morphological features is created and released to examine and visualize differential cross-scale attention; (3) Superior performance on both in-house and public datasets is demonstrated by our simple cross-scale MIL strategy. The official implementation is publicly available at https://github.com/hrlblab/CS-MIL. Ruining Deng, Can Cui 0006, Lucas W. Remedios, Shunxing Bao, R. Michael Womick, Sophie Chiron, Jia Li 0027, Joseph T. Roland, Ken S. Lau, Qi Liu 0024, Keith T. Wilson, Yaohong Wang, Lori A. Coburn, Bennett A. Landman, Yuankai Huo |
Medical Image Anal. | 10 |
| 2024 | FindAdapt: A python package for fast and accurate adapter detection in small RNA sequencingabstractAdapter trimming is an essential step for analyzing small RNA sequencing data, where reads are generally longer than target RNAs ranging from 18 to 30 bp. Most adapter trimming tools require adapter information as input. However, adapter information is hard to access, specified incorrectly, or not provided with publicly available datasets, hampering their reproducibility and reusability. Manual identification of adapter patterns from raw reads is labor-intensive and error-prone. Moreover, the use of randomized adapters to reduce ligation biases during library preparation makes adapter detection even more challenging. Here, we present FindAdapt, a Python package for fast and accurate detection of adapter patterns without relying on prior information. We demonstrated that FindAdapt was far superior to existing approaches. It identified adapters successfully in 180 simulation datasets with diverse read structures and 3,184 real datasets covering a variety of commercial and customized small RNA library preparation kits. FindAdapt is stand-alone software that can be easily integrated into small RNA sequencing analysis pipelines. Hua-Chang Chen, Jing Wang 0026, Shyr Yu, Qi Liu 0024 |
PLoS Comput. Biol. | 4 |
| 2022 | Comprehensive evaluation of noise reduction methods for single-cell RNA sequencing dataabstractNormalization and batch correction are critical steps in processing single-cell RNA sequencing (scRNA-seq) data, which remove technical effects and systematic biases to unmask biological signals of interest. Although a number of computational methods have been developed, there is no guidance for choosing appropriate procedures in different scenarios. In this study, we assessed the performance of 28 scRNA-seq noise reduction procedures in 55 scenarios using simulated and real datasets. The scenarios accounted for multiple biological and technical factors that greatly affect the denoising performance, including relative magnitude of batch effects, the extent of cell population imbalance, the complexity of cell group structures, the proportion and the similarity of nonoverlapping cell populations, dropout rates and variable library sizes. We used multiple quantitative metrics and visualization of low-dimensional cell embeddings to evaluate the performance on batch mixing while preserving the original cell group and gene structures. Based on our results, we specified technical or biological factors affecting the performance of each method and recommended proper methods in different scenarios. In addition, we highlighted one challenging scenario where most methods failed and resulted in overcorrection. Our studies not only provided a comprehensive guideline for selecting suitable noise reduction procedures but also pointed out unsolved issues in the field, especially the urgent need of developing metrics for assessing batch correction on imperceptible cell-type mixing. Shih-Kai Chu, Shilin Zhao, Shyr Yu, Qi Liu 0024 |
Briefings Bioinform. | 4 |
| 2022 | Quantifying and correcting slide-to-slide variation in multiplexed immunofluorescence imagesabstractMOTIVATION: Multiplexed imaging is a nascent single-cell assay with a complex data structure susceptible to technical variability that disrupts inference. These in situ methods are valuable in understanding cell-cell interactions, but few standardized processing steps or normalization techniques of multiplexed imaging data are available. RESULTS: We implement and compare data transformations and normalization algorithms in multiplexed imaging data. Our methods adapt the ComBat and functional data registration methods to remove slide effects in this domain, and we present an evaluation framework to compare the proposed approaches. We present clear slide-to-slide variation in the raw, unadjusted data and show that many of the proposed normalization methods reduce this variation while preserving and improving the biological signal. Furthermore, we find that dividing multiplexed imaging data by its slide mean, and the functional data registration methods, perform the best under our proposed evaluation framework. In summary, this approach provides a foundation for better data quality and evaluation criteria in multiplexed imaging. AVAILABILITY AND IMPLEMENTATION: Source code is provided at: https://github.com/statimagcoll/MultiplexedNormalization and an R package to implement these methods is available here: https://github.com/ColemanRHarris/mxnorm. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Coleman R. Harris, Eliot T. McKinley, Joseph T. Roland, Qi Liu 0024, Martha J. Shrubsole, Ken S. Lau, Robert J. Coffey, Julia Wrobel, Simon N. Vandekar |
Bioinform. | 4 |
| 2022 | Dysregulated ligand-receptor interactions from single-cell transcriptomicsabstractMOTIVATION: Intracellular communication is crucial to many biological processes, such as differentiation, development, homeostasis and inflammation. Single-cell transcriptomics provides an unprecedented opportunity for studying cell-cell communications mediated by ligand-receptor interactions. Although computational methods have been developed to infer cell type-specific ligand-receptor interactions from one single-cell transcriptomics profile, there is lack of approaches considering ligand and receptor simultaneously to identifying dysregulated interactions across conditions from multiple single-cell profiles. RESULTS: We developed scLR, a statistical method for examining dysregulated ligand-receptor interactions between two conditions. scLR models the distribution of the product of ligands and receptors expressions and accounts for inter-sample variances and small sample sizes. scLR achieved high sensitivity and specificity in simulation studies. scLR revealed important cytokine signaling between macrophages and proliferating T cells during severe acute COVID-19 infection, and activated TGF-β signaling from alveolar type II cells in the pathogenesis of pulmonary fibrosis. AVAILABILITY AND IMPLEMENTATION: scLR is freely available at https://github.com/cyhsuTN/scLR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qi Liu 0024, Chih-Yuan Hsu, Jia Li 0027, Shyr Yu |
Bioinform. | 1 |
| 2019 | scRNABatchQC: multi-samples quality control for single cell RNA-seq dataabstractSUMMARY: Single cell RNA sequencing is a revolutionary technique to characterize inter-cellular transcriptomics heterogeneity. However, the data are noise-prone because gene expression is often driven by both technical artifacts and genuine biological variations. Proper disentanglement of these two effects is critical to prevent spurious results. While several tools exist to detect and remove low-quality cells in one single cell RNA-seq dataset, there is lack of approach to examining consistency between sample sets and detecting systematic biases, batch effects and outliers. We present scRNABatchQC, an R package to compare multiple sample sets simultaneously over numerous technical and biological features, which gives valuable hints to distinguish technical artifact from biological variations. scRNABatchQC helps identify and systematically characterize sources of variability in single cell transcriptome data. The examination of consistency across datasets allows visual detection of biases and outliers. AVAILABILITY AND IMPLEMENTATION: scRNABatchQC is freely available at https://github.com/liuqivandy/scRNABatchQC as an R package. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qi Liu 0024, Quanhu Sheng, Jie Ping, Marisol Adelina Ramirez, Ken S. Lau, Robert J. Coffey, Shyr Yu |
Bioinform. | 1 |
| 2019 | Systems-level network modeling of Small Cell Lung Cancer subtypes identifies master regulators and destabilizersabstractAdopting a systems approach, we devise a general workflow to define actionable subtypes in human cancers. Applied to small cell lung cancer (SCLC), the workflow identifies four subtypes based on global gene expression patterns and ontologies. Three correspond to known subtypes (SCLC-A, SCLC-N, and SCLC-Y), while the fourth is a previously undescribed ASCL1+ neuroendocrine variant (NEv2, or SCLC-A2). Tumor deconvolution with subtype gene signatures shows that all of the subtypes are detectable in varying proportions in human and mouse tumors. To understand how multiple stable subtypes can arise within a tumor, we infer a network of transcription factors and develop BooleaBayes, a minimally-constrained Boolean rule-fitting approach. In silico perturbations of the network identify master regulators and destabilizers of its attractors. Specific to NEv2, BooleaBayes predicts ELF3 and NR0B1 as master regulators of the subtype, and TCF3 as a master destabilizer. Since the four subtypes exhibit differential drug sensitivity, with NEv2 consistently least sensitive, these findings may lead to actionable therapeutic strategies that consider SCLC intratumoral heterogeneity. Our systems-level approach should generalize to other cancer types. David J. Wooten, Sarah M. Groves, Darren R. Tyson, Qi Liu 0024, Jing S. Lim, Réka Albert, Carlos F. Lopez, Julien Sage, Vito Quaranta |
PLoS Comput. Biol. | 4 |
| 2017 | Investigating MicroRNA and transcription factor co-regulatory networks in colorectal cancerabstractBACKGROUND: Colorectal cancer (CRC) is one of the most common malignancies worldwide with poor prognosis. Studies have showed that abnormal microRNA (miRNA) expression can affect CRC pathogenesis and development through targeting critical genes in cellular system. However, it is unclear about which miRNAs play central roles in CRC's pathogenesis and how they interact with transcription factors (TFs) to regulate the cancer-related genes. RESULTS: To address this issue, we systematically explored the major regulation motifs, namely feed-forward loops (FFLs), that consist of miRNAs, TFs and CRC-related genes through the construction of a miRNA-TF regulatory network in CRC. First, we compiled CRC-related miRNAs, CRC-related genes, and human TFs from multiple data sources. Second, we identified 13,123 3-node FFLs including 25 miRNA-FFLs, 13,005 TF-FFLs and 93 composite-FFLs, and merged the 3-node FFLs to construct a CRC-related regulatory network. The network consists of three types of regulatory subnetworks (SNWs): miRNA-SNW, TF-SNW, and composite-SNW. To enhance the accuracy of the network, the results were filtered by using The Cancer Genome Atlas (TCGA) expression data in CRC, whereby we generated a core regulatory network consisting of 58 significant FFLs. We then applied a hub identification strategy to the significant FFLs and found 5 significant components, including two miRNAs (hsa-miR-25 and hsa-miR-31), two genes (ADAMTSL3 and AXIN1) and one TF (BRCA1). The follow up prognosis analysis indicated all of the 5 significant components having good prediction of overall survival of CRC patients. CONCLUSIONS: In summary, we generated a CRC-specific miRNA-TF regulatory network, which is helpful to understand the complex CRC regulatory mechanisms and guide clinical treatment. The discovered 5 regulators might have critical roles in CRC pathogenesis and warrant future investigation. Jiamao Luo, Huilin Niu, Jing Wang 0026, Qi Liu 0024, Zhongming Zhao, Hua Xu 0001, Yanqing Ding, Jingchun Sun, Qingling Zhang 0003 |
BMC Bioinform. | 6 |
| 2016 | Proceedings of the 15th Annual UT-KBRIN Bioinformatics Summit 2016: Cadiz, KY, USA. 8-10 April 2016abstractI1 Proceedings of the Fifteenth Annual UT- KBRIN Bioinformatics Summit 2016 Eric C. Rouchka, Julia H. Chariker, Benjamin J. Harrison, Juw Won Park P1 CC-PROMISE: Projection onto the Most Interesting Statistical Evidence (PROMISE) with Canonical Correlation to integrate gene expression and methylation data with multiple pharmacologic and clinical endpoints Xueyuan Cao, Stanley Pounds, Susana Raimondi, James Downing, Raul Ribeiro, Jeffery Rubnitz, Jatinder Lamba P2 Integration of microRNA-mRNA interaction networks with gene expression data to increase experimental power Bernie J Daigle, Jr. P3 Designing and writing software for in silico subtractive hybridization of large eukaryotic genomes Deborah Burgess, Stephanie Gehrlich, John C Carmen P4 Tracking the molecular evolution of Pax gene Nicholas Johnson; Chandrakanth Emani P5 Identifying genetic differences in thermally dimorphic and state specific fungi using in silico genomic comparison Stephanie Gehrlich, Deborah Burgess, John C Carmen P6 Identification of conserved genomic regions and variation therein amongst Cetartiodactyla species using next generation sequencing Kalpani De Silva, Michael P Heaton, Theodore S Kalbfleisch P7 Mining physiological data to identify patients with similar medical events and phenotypes Teeradache Viangteeravat, Rahul Mudunuri, Oluwaseun Ajayi, Fatih Şen, Eunice Y Huang P8 Smart brief for home health monitoring Mohammad Mohebbi, Luaire Florian, Douglas J Jackson, John F Naber P9 Side-effect term matching for computational adverse drug reaction predictions AKM Sabbir, Sally R Ellingson P10 Enrichment vs robustness: A comparison of transcriptomic data clustering metrics Yuping Lu, Charles A Phillips, Michael A Langston P11 Deep neural networks for transcriptome-based cancer classification Rahul K Sevakula, Raghuveer Thirukovalluru, Nishchal K. Verma, Yan Cui P12 Motif discovery using K-means clustering Mohammed Sayed, Juw Won Park P13 Large scale discovery of active enhancers from nascent RNA sequencing Jing Wang, Qi Liu, Yu Shyr P14 Computationally characterizing genomic pipelines and benchmarking results using GATK best practices on the high performance computing cluster at the University of Kentucky Xiaofei Zhang, Sally R Ellingson P15 Development of approaches enabling the identification of abnormal gene expression from RNA-Seq in personalized oncology Naresh Prodduturi, Gavin R Oliver, Diane Grill, Jie Na, Jeanette Eckel-Passow, Eric W Klee P16 Processing RNA-Seq data of plants infected with coffee ringspot virus Michael M Goodin, Mark Farman, Harrison Inocencio, Chanyong Jang, Jerzy W Jaromczyk, Neil Moore, Kelly Sovacool P17 Comparative transcriptomics of three Acinetobacter baumanii clinical isolates with different antibiotic resistance patterns Leon Dent, Mike Izban, Sammed Mandape, Shruti Sakhare, Siddharth Pratap, Dana Marshall P18 Metagenomic assessment of possible microbial contamination in the equine reference genome assembly M Scotty DePriest, James N MacLeod, Theodore S Kalbfleisch P19 Molecular evolution of cancer driver genes Chandrakanth Emani, Hanady Adam, Ethan Blandford, Joel Campbell, Joshua Castlen, Brittany Dixon, Ginger Gilbert, Aaron Hall, Philip Kreisle, Jessica Lasher, Bethany Oakes, Allison Speer, Maximilian Valentine P20 Biorepository Laboratory Information Management System Naga Satya V Rao Nagisetty, Rony Jose, Teeradache Viangteeravat, Robert Rooney, David Hains Eric C. Rouchka, Julia H. Chariker, Benjamin J. Harrison, Juw Won Park, Xueyuan Cao, Stan Pounds, Susana C. Raimondi, James R. Downing, Raul C. Ribeiro, Jeffrey Rubnitz, Jatinder Lamba, Bernie J. Daigle Jr., Deborah Burgess, Stephanie Gehrlich, John C. Carmen, Chandrakanth Emani, Kalpani De Silva, Michael P. Heaton, Ted Kalbfleisch, Teeradache Viangteeravat, Rahul Mudunuri, Oluwaseun Ajayi, Fatih Sen, Eunice Y. Huang, Mohammad Mohebbi, Luaire Florian, Douglas J. Jackson, John F. Naber, Akm Sabbir, Sally R. Ellingson, Yuping Lu, Charles A. Phillips, Michael A. Langston, Rahul Kumar Sevakula, Raghuveer Thirukovalluru, Nishchal K. Verma, Yan Cui 0001, Mohammed Sayed, Jing Wang 0026, Qi Liu 0024, Shyr Yu, Naresh Prodduturi, Gavin R. Oliver, Diane Grill, Jie Na, Jeanette Eckel-Passow, Eric W. Klee, Michael M. Goodin, Mark L. Farman, Harrison Inocencio, Chanyong Jang, Jerzy W. Jaromczyk, Neil Moore, Kelly L. Sovacool, Leon Dent, Mike Izban, Sammed N. Mandape, Shruti S. Sakhare, Siddharth Pratap, Dana Marshall, M. Scotty Depriest, James N. MacLeod, Hanady Adam, Ethan Blandford, Joel Campbell, Joshua Castlen, Brittany Dixon, Ginger Gilbert, Aaron Hall, Philip Kreisle, Jessica Lasher, Bethany Oakes, Allison Speer, Maximilian Valentine, Naga Satya Venkateswara Ra Nagisetty, Rony Jose, Robert W. Rooney, David Hains |
BMC Bioinform. | 41 |
| 2016 | An improved method to construct basic probability assignment based on the confusion matrix for classification problem
Xinyang Deng, Qi Liu 0024, Yong Deng 0001, Sankaran Mahadevan |
Inf. Sci. | 2 |
| 2015 | Newborns prediction based on a belief Markov chain model
Xinyang Deng, Qi Liu 0024, Yong Deng 0001 |
Appl. Intell. | 2 |
| 2014 | CLIP-EZ: a computational tool for HITS-CLIP data analysisabstractBackground Mapping the binding regions of mRNA-binding proteins is critical to the understanding of their regulatory roles in cellular processes. Recent development in experimental technologies combines high throughput sequencing with crosslink immunoprecipitation (HITS-CLIP), which has the merit of detecting RNA-protein interaction sites at a high resolution to single nucleotide level. Analysis of such data typically involves many steps, and teasing out true signals from noise requires crosslink induced mutations (CIMS) analysis, peak identification, integration of the two signal types, and some downstream analysis such as motif finding and conservation evaluation. To our knowledge, there is a lack of a single computational tool that can perform all the tasks as mentioned in an easily accessible manner. Despite the fact that there are several tools available, each performs an individual task. Xue Zhong, Qi Liu 0024, Shyr Yu |
BMC Bioinform. | 2 |
| 2010 | TF-centered downstream gene set enrichment analysis: Inference of causal regulators by integrating TF-DNA interactions and protein post-translational modifications informationabstractBACKGROUND: Inference of causal regulators responsible for gene expression changes under different conditions is of great importance but remains rather challenging. To date, most approaches use direct binding targets of transcription factors (TFs) to associate TFs with expression profiles. However, the low overlap between binding targets of a TF and the affected genes of the TF knockout limits the power of those methods. RESULTS: We developed a TF-centered downstream gene set enrichment analysis approach to identify potential causal regulators responsible for expression changes. We constructed hierarchical and multi-layer regulation models to derive possible downstream gene sets of a TF using not only TF-DNA interactions, but also, for the first time, post-translational modifications (PTM) information. We verified our method in one expression dataset of large-scale TF knockout and another dataset involving both TF knockout and TF overexpression. Compared with the flat model using TF-DNA interactions alone, our method correctly identified five more actual perturbed TFs in large-scale TF knockout data and six more perturbed TFs in overexpression data. Potential regulatory pathways downstream of three perturbed regulators- SNF1, AFT1 and SUT1 -were given to demonstrate the power of multilayer regulation models integrating TF-DNA interactions and PTM information. Additionally, our method successfully identified known important TFs and inferred some novel potential TFs involved in the transition from fermentative to glycerol-based respiratory growth and in the pheromone response. Downstream regulation pathways of SUT1 and AFT1 were also supported by the mRNA and/or phosphorylation changes of their mediating TFs and/or "modulator" proteins. CONCLUSIONS: The results suggest that in addition to direct transcription, indirect transcription and post-translational regulation are also responsible for the effects of TFs perturbation, especially for TFs overexpression. Many TFs inferred by our method are supported by literature. Multiple TF regulation models could lead to new hypotheses for future experiments. Our method provides a valuable framework for analyzing gene expression data to identify causal regulators in the context of TF-DNA interactions and PTM information. Qi Liu 0024, Yejun Tan, Tao Huang 0004, Guohui Ding 0001, Zhidong Tu, Hongyue Dai, Lu Xie |
BMC Bioinform. | 1 |
| 2010 | Erratum to "Combining belief functions based on distance of evidence" [Decision Support Systems Volume(38/3)489-493]
Deqiang Han, Yong Deng 0001, Qi Liu 0024 |
Decis. Support Syst. | 3 |
| 2007 | InPrePPI: an integrated evaluation method based on genomic context for predicting protein-protein interactions in prokaryotic genomesabstractBACKGROUND: Although many genomic features have been used in the prediction of protein-protein interactions (PPIs), frequently only one is used in a computational method. After realizing the limited power in the prediction using only one genomic feature, investigators are now moving toward integration. So far, there have been few integration studies for PPI prediction; one failed to yield appreciable improvement of prediction and the others did not conduct performance comparison. It remains unclear whether an integration of multiple genomic features can improve the PPI prediction and, if it can, how to integrate these features. RESULTS: In this study, we first performed a systematic evaluation on the PPI prediction in Escherichia coli (E. coli) by four genomic context based methods: the phylogenetic profile method, the gene cluster method, the gene fusion method, and the gene neighbor method. The number of predicted PPIs and the average degree in the predicted PPI networks varied greatly among the four methods. Further, no method outperformed the others when we tested using three well-defined positive datasets from the KEGG, EcoCyc, and DIP databases. Based on these comparisons, we developed a novel integrated method, named InPrePPI. InPrePPI first normalizes the AC value (an integrated value of the accuracy and coverage) of each method using three positive datasets, then calculates a weight for each method, and finally uses the weight to calculate an integrated score for each protein pair predicted by the four genomic context based methods. We demonstrate that InPrePPI outperforms each of the four individual methods and, in general, the other two existing integrated methods: the joint observation method and the integrated prediction method in STRING. These four methods and InPrePPI are implemented in a user-friendly web interface. CONCLUSION: This study evaluated the PPI prediction by four genomic context based methods, and presents an integrated evaluation method that shows better performance in E. coli. Jingchun Sun, Guohui Ding 0001, Qi Liu 0024, Youyu He, Tieliu Shi, Zhongming Zhao |
BMC Bioinform. | 4 |
| 2006 | Insights into the Coupling of Duplication Events and Macroevolution from an Age Profile of Animal Transmembrane Gene FamiliesabstractThe evolution of new gene families subsequent to gene duplication may be coupled to the fluctuation of population and environment variables. Based upon that, we presented a systematic analysis of the animal transmembrane gene duplication events on a macroevolutionary scale by integrating the palaeontology repository. The age of duplication events was calculated by maximum likelihood method, and the age distribution was estimated by density histogram and normal kernel density estimation. We showed that the density of the duplicates displays a positive correlation with the estimates of maximum number of cell types of common ancestors, and the oxidation events played a key role in the major transitions of this density trace. Next, we focused on the Phanerozoic phase, during which more macroevolution data are available. The pulse mass extinction timepoints coincide with the local peaks of the age distribution, suggesting that the transmembrane gene duplicates fixed frequently when the environment changed dramatically. Moreover, a 61-million-year cycle is the most possible cycle in this phase by spectral analysis, which is consistent with the cycles recently detected in biodiversity. Our data thus elucidate a strong coupling of duplication events and macroevolution; furthermore, our method also provides a new way to address these questions. Guohui Ding 0001, Jiuhong Kang, Qi Liu 0024, Tieliu Shi, Gang Pei |
PLoS Comput. Biol. | 3 |
| 2005 | Refined phylogenetic profiles method for predicting protein-protein interactionsabstractMOTIVATION: The increasing availability of complete genome sequences provides excellent opportunity for the further development of tools for functional studies in proteomics. Several experimental approaches and in silico algorithms have been developed to cluster proteins into networks of biological significance that may provide new biological insights, especially into understanding the functions of many uncharacterized proteins. Among these methods, the phylogenetic profiles method has been widely used to predict protein-protein interactions. It involves the selection of reference organisms and identification of homologous proteins. Up to now, no published report has systematically studied the effects of the reference genome selection and the identification of homologous proteins upon the accuracy of this method. RESULTS: In this study, we optimized the phylogenetic profiles method by integrating phylogenetic relationships among reference organisms and sequence homology information to improve prediction accuracy. Our results revealed that the selection of the reference organisms set and the criteria for homology identification significantly are two critical factors for the prediction accuracy of this method. Our refined phylogenetic profiles method shows greater performance and potentially provides more reliable functional linkages compared with previous methods. Jingchun Sun, Jinlin Xu, Qi Liu 0024, Aimin Zhao, Tieliu Shi |
Bioinform. | 4 |
| 2005 | A topsis-based centroid-index ranking method of fuzzy numbers and its application in decision-makingabstractRanking fuzzy numbers plays a very important role in decision-making problems. Existing centroid-index ranking methods have some drawbacks. In this article, a new centroid-index ranking method of fuzzy numbers is proposed. The proposed method is using the ideal of Technique for Order Preference by Similarity to an Ideal Solution (TOPSIS). Some numerical examples show that the new method can overcome the drawbacks of the existing methods. Finally, a human selection problem is used to illustrate the efficiency of the proposed fuzzy ranking method. Yong Deng 0001, Qi Liu 0024 |
Cybern. Syst. | 2 |
| 2004 | Combining belief functions based on distance of evidence
Yong Deng 0001, Wenkang Shi, Zhenfu Zhu, Qi Liu 0024 |
Decis. Support Syst. | 4 |
| 2004 | A new similarity measure of generalized fuzzy numbers and its application to pattern recognition
Yong Deng 0001, Wenkang Shi, Qi Liu 0024 |
Pattern Recognit. Lett. | 4 |