VLDB 2026 Research / reviewers in the wild / expert
Eric W. Klee
dblp:89/249
· DBLP profile ↗
15ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0003-2946-5795ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A clinical knowledge graph-based framework to prioritize candidate genes for facilitating diagnosis of Mendelian diseases and rare genetic conditionsabstractBACKGROUND: Diagnosing Mendelian and rare genetic conditions requires identifying phenotype-associated genetic findings and prioritizing likely disease-causing genes. This task is labor-intensive for molecular and clinical geneticists, who must review extensive literature and databases to link patient phenotypes with causal genotypes. The challenge is further complicated by the large number of genetic variants detected through next-generation sequencing, which impacts both diagnosis timelines and patient care strategies. To address this, in silico methods that prioritize causal genes based on patient-derived phenotypes offer an effective solution, reducing the time involved in diagnostic case reviews and enhancing the efficiency of clinical diagnosis. RESULTS: We developed the phenotype prioritization and analysis for rare diseases (PPAR) to rank genes based on human phenotype ontology (HPO) terms, with the specific goal of aiding the interpretation of genetic testing for Mendelian and rare diseases. PPAR leverages embeddings from a knowledge graph and incorporates knowledge from connections between genes, HPO terms, and gene ontology annotations. When applied on a clinical rare disease cohort and the publicly available deciphering developmental disorders (DDD) dataset. PPAR ranked the causal gene in the top 10 for 27% of cases in the clinical cohort and for 85% of cases in the DDD dataset, outperforming other established HPO-based methods. CONCLUSION: Our findings demonstrate that PPAR, a method developed from the clinical knowledge graph, effectively ranks causal genes based on patient-derived HPO terms in rare and Mendelian disease contexts. PPAR has shown superior performance compared to other well-established HPO-only methods and provides an efficient, accessible solution for clinical geneticists. The Python-based tool is publicly available at https://github.com/dimi-lab/PPAR , offering a user-friendly platform for gene prioritization. Rohan Gnanaolivu, Gavin R. Oliver, W. Garrett Jenkinson, Emily Blake, Wenan Chen, Nicholas Chia, Eric W. Klee |
BMC Bioinform. | 7 |
| 2024 | A supervised learning method for classifying methylation disordersabstractBACKGROUND: DNA methylation is one of the most stable and well-characterized epigenetic alterations in humans. Accordingly, it has already found clinical utility as a molecular biomarker in a variety of disease contexts. Existing methods for clinical diagnosis of methylation-related disorders focus on outlier detection in a small number of CpG sites using standardized cutoffs which differentiate healthy from abnormal methylation levels. The standardized cutoff values used in these methods do not take into account methylation patterns which are known to differ between the sexes and with age. RESULTS: Here we profile genome-wide DNA methylation from blood samples drawn from within a cohort composed of healthy controls of different age and sex alongside patients with Prader-Willi syndrome (PWS), Beckwith-Wiedemann syndrome, Fragile-X syndrome, Angelman syndrome, and Silver-Russell syndrome. We propose a Generalized Additive Model to perform age and sex adjusted outlier analysis of around 700,000 CpG sites throughout the human genome. Utilizing z-scores among the cohort for each site, we deployed an ensemble based machine learning pipeline and achieved a combined prediction accuracy of 0.96 (Binomial 95% Confidence Interval 0.868[Formula: see text]0.995). CONCLUSION: We demonstrate a method for age and sex adjusted outlier detection of differentially methylated loci based on a large cohort of healthy individuals. We present a custom machine learning pipeline utilizing this outlier analysis to classify samples for potential methylation associated congenital disorders. These methods are able to achieve high accuracy when used with machine learning methods to classify abnormal methylation patterns. Jesse R. Walsh, Guangchao Sun, Jagadheshwar Balan, Jayson Hardcastle, Jason Vollenweider, Calvin Jerde, Kandelaria Rumilla, Christy Koellner, Alaa Koleilat, Linda Hasadsri, Benjamin Kipp, W. Garrett Jenkinson, Eric W. Klee |
BMC Bioinform. | 13 |
| 2023 | Enhancing Patient Care in Rare Genetic Diseases: An HPO-based Phenotyping PipelineabstractIdentifying rare genetic diseases poses challenges, often resulting in delayed referrals to genetic specialty care. Consequently, a large number of rare diseases remain undiagnosed for years, with many of these patients often dying without an accurate diagnosis. To help address these diagnostic odysseys, we propose a two-step pipeline using natural language processing (NLP) and machine learning (ML) techniques leveraging raw clinical text narratives of patients from longitudinal clinical data in electronic health records (EHR). Our approach employs two state-of-the-art concept recognition algorithms, ClinPhen and PhenoTagger, to extract human phenotype ontology (HPO) terms from the clinical notes and ML algorithms to identify patients who may benefit from referral. By identifying key phrases indicative of rare diseases, our system can prompt healthcare providers to seek clinical genomics consultations, improving diagnosis and patient outcomes. We evaluated the predictive performance of extracted HPO terms using different ML models. Our strategy and methodological approach are intended to serve as complementary steps in aiding clinicians and researchers to improve the diagnosis efficiency of rare genetic diseases. Moein Enayati, Gavin M. Schaeferle, Brendan C. Lanpher, Eric W. Klee, Che Ngufor |
BIBM | 5 |
| 2021 | HELLO: improved neural network architectures and methodologies for small variant callingabstractBACKGROUND: Modern Next Generation- and Third Generation- Sequencing methods such as Illumina and PacBio Circular Consensus Sequencing platforms provide accurate sequencing data. Parallel developments in Deep Learning have enabled the application of Deep Neural Networks to variant calling, surpassing the accuracy of classical approaches in many settings. DeepVariant, arguably the most popular among such methods, transforms the problem of variant calling into one of image recognition where a Deep Neural Network analyzes sequencing data that is formatted as images, achieving high accuracy. In this paper, we explore an alternative approach to designing Deep Neural Networks for variant calling, where we use meticulously designed Deep Neural Network architectures and customized variant inference functions that account for the underlying nature of sequencing data instead of converting the problem to one of image recognition. RESULTS: Results from 27 whole-genome variant calling experiments spanning Illumina, PacBio and hybrid Illumina-PacBio settings suggest that our method allows vastly smaller Deep Neural Networks to outperform the Inception-v3 architecture used in DeepVariant for indel and substitution-type variant calls. For example, our method reduces the number of indel call errors by up to 18%, 55% and 65% for Illumina, PacBio and hybrid Illumina-PacBio variant calling respectively, compared to a similarly trained DeepVariant pipeline. In these cases, our models are between 7 and 14 times smaller. CONCLUSIONS: We believe that the improved accuracy and problem-specific customization of our models will enable more accurate pipelines and further method development in the field. HELLO is available at https://github.com/anands-repo/hello. Anand Ramachandran 0001, Steven S. Lumetta, Eric W. Klee, Deming Chen |
BMC Bioinform. | 3 |
| 2020 | LeafCutterMD: an algorithm for outlier splicing detection in rare diseasesabstractMOTIVATION: Next-generation sequencing is rapidly improving diagnostic rates in rare Mendelian diseases, but even with whole genome or whole exome sequencing, the majority of cases remain unsolved. Increasingly, RNA sequencing is being used to solve many cases that evade diagnosis through sequencing alone. Specifically, the detection of aberrant splicing in many rare disease patients suggests that identifying RNA splicing outliers is particularly useful for determining causal Mendelian disease genes. However, there is as yet a paucity of statistical methodologies to detect splicing outliers. RESULTS: We developed LeafCutterMD, a new statistical framework that significantly improves the previously published LeafCutter in the context of detecting outlier splicing events. Through simulations and analysis of real patient data, we demonstrate that LeafCutterMD has better power than the state-of-the-art methodology while controlling false-positive rates. When applied to a cohort of disease-affected probands from the Mayo Clinic Center for Individualized Medicine, LeafCutterMD recovered all aberrantly spliced genes that had previously been identified by manual curation efforts. AVAILABILITY AND IMPLEMENTATION: The source code for this method is available under the opensource Apache 2.0 license in the latest release of the LeafCutter software package available online at http://davidaknowles.github.io/leafcutter. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. W. Garrett Jenkinson, Yang I Li, Shubham Basu, Margot A. Cousin, Gavin R. Oliver, Eric W. Klee |
Bioinform. | 6 |
| 2019 | A recurrent Markov state-space generative model for sequencesabstractWhile the Hidden Markov Model (HMM) is a versatile generative model of sequences capable of performing many exact inferences efficiently, it is not suited for capturing complex long-term structure in the data. Advanced state-space models based on Deep Neural Networks (DNN) overcome this limitation but cannot perform exact inferences. In this article, we present a new generative model for sequences that combines both aspects, the ability to perform exact inferences and the ability to model long-term structure, by augmenting the HMM with a deterministic, continuous state variable modeled through a Recurrent Neural Network. We empirically study the performance of the model on (i) synthetic data comparing it to the HMM, (ii) a supervised learning task in bioinformatics where it outperforms two DNN-based regressors and (iii) in the generative modeling of music where it outperforms many prominent DNN-based generative models. Anand Ramachandran 0001, Steven S. Lumetta, Eric W. Klee, Deming Chen |
AISTATS | 3 |
| 2019 | Recommendations for performance optimizations when using GATK3.8 and GATK4abstractBACKGROUND: Use of the Genome Analysis Toolkit (GATK) continues to be the standard practice in genomic variant calling in both research and the clinic. Recently the toolkit has been rapidly evolving. Significant computational performance improvements have been introduced in GATK3.8 through collaboration with Intel in 2017. The first release of GATK4 in early 2018 revealed rewrites in the code base, as the stepping stone toward a Spark implementation. As the software continues to be a moving target for optimal deployment in highly productive environments, we present a detailed analysis of these improvements, to help the community stay abreast with changes in performance. RESULTS: We re-evaluated multiple options, such as threading, parallel garbage collection, I/O options and data-level parallelization. Additionally, we considered the trade-offs of using GATK3.8 and GATK4. We found optimized parameter values that reduce the time of executing the best practices variant calling procedure by 29.3% for GATK3.8 and 16.9% for GATK4. Further speedups can be accomplished by splitting data for parallel analysis, resulting in run time of only a few hours on whole human genome sequenced to the depth of 20X, for both versions of GATK. Nonetheless, GATK4 is already much more cost-effective than GATK3.8. Thanks to significant rewrites of the algorithms, the same analysis can be run largely in a single-threaded fashion, allowing users to process multiple samples on the same CPU. CONCLUSIONS: In time-sensitive situations, when a patient has a critical or rapidly developing condition, it is useful to minimize the time to process a single sample. In such cases we recommend using GATK3.8 by splitting the sample into chunks and computing across multiple nodes. The resultant walltime will be nnn.4 hours at the cost of $41.60 on 4 c5.18xlarge instances of Amazon Cloud. For cost-effectiveness of routine analyses or for large population studies, it is useful to maximize the number of samples processed per unit time. Thus we recommend GATK4, running multiple samples on one node. The total walltime will be ∼34.1 hours on 40 samples, with 1.18 samples processed per hour at the cost of $2.60 per sample on c5.18xlarge instance of Amazon Cloud. Jacob R Heldenbrand, Saurabh Baheti, Matthew A. Bockol, Travis M. Drucker, Steven N. Hart, Matthew E. Hudson, Ravishankar K. Iyer, Michael Kalmbach, Katherine Irene Kendig, Eric W. Klee, Nathan R Mattson, Eric D. Wieben, Mathieu Wiepert, Derek E. Wildman, Liudmila S. Mainzer |
BMC Bioinform. | 10 |
| 2019 | Correction to: Recommendations for performance optimizations when using GATK3.8 and GATK4abstractFollowing publication of the original article [1], the author explained that Table 2 is displayed incorrectly. The correct Table 2 is given below. The original article has been corrected. Jacob R Heldenbrand, Saurabh Baheti, Matthew A. Bockol, Travis M. Drucker, Steven N. Hart, Matthew E. Hudson, Ravishankar K. Iyer, Michael Kalmbach, Katherine Irene Kendig, Eric W. Klee, Nathan R Mattson, Eric D. Wieben, Mathieu Wiepert, Derek E. Wildman, Liudmila S. Mainzer |
BMC Bioinform. | 10 |
| 2018 | Deep Learning for Better Variant Calling for Cancer Diagnosis and TreatmentabstractHigh-throughput techniques have revolutionized the study of genomics and molecular biology in recent years. These methods provide a large quantity of sequence data, and have applications in different areas of bioinformatics. One can sequence parts or whole of an organism's DNA to determine genetic information about an individual or a population, measure expression levels of different genes under different conditions, and determine binding affinity of proteins to DNA segments revealing details regarding gene regulation, at a higher resolution than before. However, different high-throughput methods that target even a single application have different underlying error models. Robust analytic pipelines are necessary to extract necessary information from the raw data. In this paper, we discuss future research directions for developing such analytics using techniques from Machine Learning and Deep Neural Networks. We focus on two applications that will affect the diagnosis and treatment of cancer. Anand Ramachandran 0001, Huiren Li, Eric W. Klee, Steven S. Lumetta, Deming Chen |
ASP-DAC | 3 |
| 2018 | Comparative analysis of de novo assemblers for variation discovery in personal genomesabstractCurrent variant discovery approaches often rely on an initial read mapping to the reference sequence. Their effectiveness is limited by the presence of gaps, potential misassemblies, regions of duplicates with a high-sequence similarity and regions of high-sequence divergence in the reference. Also, mapping-based approaches are less sensitive to large INDELs and complex variations and provide little phase information in personal genomes. A few de novo assemblers have been developed to identify variants through direct variant calling from the assembly graph, micro-assembly and whole-genome assembly, but mainly for whole-genome sequencing (WGS) data. We developed SGVar, a de novo assembly workflow for haplotype-based variant discovery from whole-exome sequencing (WES) data. Using simulated human exome data, we compared SGVar with five variation-aware de novo assemblers and with BWA-MEM together with three haplotype- or local de novo assembly-based callers. SGVar outperforms the other assemblers in sensitivity and tolerance of sequencing errors. We recapitulated the findings on whole-genome and exome data from a Utah residents with Northern and Western European ancestry (CEU) trio, showing that SGVar had high sensitivity both in the highly divergent human leukocyte antigen (HLA) region and in non-HLA regions of chromosome 6. In particular, SGVar is robust to sequencing error, k-mer selection, divergence level and coverage depth. Unlike mapping-based approaches, SGVar is capable of resolving long-range phase and identifying large INDELs from WES, more prominently from WGS. We conclude that SGVar represents an ideal platform for WES-based variant discovery in highly divergent regions and across the whole genome. Shulan Tian, Huihuang Yan, Eric W. Klee, Michael Kalmbach, Susan L. Slager |
Briefings Bioinform. | 3 |
| 2016 | Proceedings of the 15th Annual UT-KBRIN Bioinformatics Summit 2016: Cadiz, KY, USA. 8-10 April 2016abstractI1 Proceedings of the Fifteenth Annual UT- KBRIN Bioinformatics Summit 2016 Eric C. Rouchka, Julia H. Chariker, Benjamin J. Harrison, Juw Won Park P1 CC-PROMISE: Projection onto the Most Interesting Statistical Evidence (PROMISE) with Canonical Correlation to integrate gene expression and methylation data with multiple pharmacologic and clinical endpoints Xueyuan Cao, Stanley Pounds, Susana Raimondi, James Downing, Raul Ribeiro, Jeffery Rubnitz, Jatinder Lamba P2 Integration of microRNA-mRNA interaction networks with gene expression data to increase experimental power Bernie J Daigle, Jr. P3 Designing and writing software for in silico subtractive hybridization of large eukaryotic genomes Deborah Burgess, Stephanie Gehrlich, John C Carmen P4 Tracking the molecular evolution of Pax gene Nicholas Johnson; Chandrakanth Emani P5 Identifying genetic differences in thermally dimorphic and state specific fungi using in silico genomic comparison Stephanie Gehrlich, Deborah Burgess, John C Carmen P6 Identification of conserved genomic regions and variation therein amongst Cetartiodactyla species using next generation sequencing Kalpani De Silva, Michael P Heaton, Theodore S Kalbfleisch P7 Mining physiological data to identify patients with similar medical events and phenotypes Teeradache Viangteeravat, Rahul Mudunuri, Oluwaseun Ajayi, Fatih Şen, Eunice Y Huang P8 Smart brief for home health monitoring Mohammad Mohebbi, Luaire Florian, Douglas J Jackson, John F Naber P9 Side-effect term matching for computational adverse drug reaction predictions AKM Sabbir, Sally R Ellingson P10 Enrichment vs robustness: A comparison of transcriptomic data clustering metrics Yuping Lu, Charles A Phillips, Michael A Langston P11 Deep neural networks for transcriptome-based cancer classification Rahul K Sevakula, Raghuveer Thirukovalluru, Nishchal K. Verma, Yan Cui P12 Motif discovery using K-means clustering Mohammed Sayed, Juw Won Park P13 Large scale discovery of active enhancers from nascent RNA sequencing Jing Wang, Qi Liu, Yu Shyr P14 Computationally characterizing genomic pipelines and benchmarking results using GATK best practices on the high performance computing cluster at the University of Kentucky Xiaofei Zhang, Sally R Ellingson P15 Development of approaches enabling the identification of abnormal gene expression from RNA-Seq in personalized oncology Naresh Prodduturi, Gavin R Oliver, Diane Grill, Jie Na, Jeanette Eckel-Passow, Eric W Klee P16 Processing RNA-Seq data of plants infected with coffee ringspot virus Michael M Goodin, Mark Farman, Harrison Inocencio, Chanyong Jang, Jerzy W Jaromczyk, Neil Moore, Kelly Sovacool P17 Comparative transcriptomics of three Acinetobacter baumanii clinical isolates with different antibiotic resistance patterns Leon Dent, Mike Izban, Sammed Mandape, Shruti Sakhare, Siddharth Pratap, Dana Marshall P18 Metagenomic assessment of possible microbial contamination in the equine reference genome assembly M Scotty DePriest, James N MacLeod, Theodore S Kalbfleisch P19 Molecular evolution of cancer driver genes Chandrakanth Emani, Hanady Adam, Ethan Blandford, Joel Campbell, Joshua Castlen, Brittany Dixon, Ginger Gilbert, Aaron Hall, Philip Kreisle, Jessica Lasher, Bethany Oakes, Allison Speer, Maximilian Valentine P20 Biorepository Laboratory Information Management System Naga Satya V Rao Nagisetty, Rony Jose, Teeradache Viangteeravat, Robert Rooney, David Hains Eric C. Rouchka, Julia H. Chariker, Benjamin J. Harrison, Juw Won Park, Xueyuan Cao, Stan Pounds, Susana C. Raimondi, James R. Downing, Raul C. Ribeiro, Jeffrey Rubnitz, Jatinder Lamba, Bernie J. Daigle Jr., Deborah Burgess, Stephanie Gehrlich, John C. Carmen, Chandrakanth Emani, Kalpani De Silva, Michael P. Heaton, Ted Kalbfleisch, Teeradache Viangteeravat, Rahul Mudunuri, Oluwaseun Ajayi, Fatih Sen, Eunice Y. Huang, Mohammad Mohebbi, Luaire Florian, Douglas J. Jackson, John F. Naber, Akm Sabbir, Sally R. Ellingson, Yuping Lu, Charles A. Phillips, Michael A. Langston, Rahul Kumar Sevakula, Raghuveer Thirukovalluru, Nishchal K. Verma, Yan Cui 0001, Mohammed Sayed, Jing Wang 0026, Qi Liu 0024, Shyr Yu, Naresh Prodduturi, Gavin R. Oliver, Diane Grill, Jie Na, Jeanette Eckel-Passow, Eric W. Klee, Michael M. Goodin, Mark L. Farman, Harrison Inocencio, Chanyong Jang, Jerzy W. Jaromczyk, Neil Moore, Kelly L. Sovacool, Leon Dent, Mike Izban, Sammed N. Mandape, Shruti S. Sakhare, Siddharth Pratap, Dana Marshall, M. Scotty Depriest, James N. MacLeod, Hanady Adam, Ethan Blandford, Joel Campbell, Joshua Castlen, Brittany Dixon, Ginger Gilbert, Aaron Hall, Philip Kreisle, Jessica Lasher, Bethany Oakes, Allison Speer, Maximilian Valentine, Naga Satya Venkateswara Ra Nagisetty, Rony Jose, Robert W. Rooney, David Hains |
BMC Bioinform. | 49 |
| 2012 | TREAT: a bioinformatics tool for variant annotations and visualizations in targeted and exome sequencing dataabstractUNLABELLED: TREAT (Targeted RE-sequencing Annotation Tool) is a tool for facile navigation and mining of the variants from both targeted resequencing and whole exome sequencing. It provides a rich integration of publicly available as well as in-house developed annotations and visualizations for variants, variant-hosting genes and host-gene pathways. AVAILABILITY AND IMPLEMENTATION: TREAT is freely available to non-commercial users as either a stand-alone annotation and visualization tool, or as a comprehensive workflow integrating sequencing alignment and variant calling. The executables, instructions and the Amazon Cloud Images of TREAT can be downloaded at the website: http://ndc.mayo.edu/mayo/research/biostat/stand-alone-packages.cfm. Yan W. Asmann, Sumit Middha, Asif Hossain, Saurabh Baheti, High-Seng Chai, Zhifu Sun, Patrick H. Duffy, Ahmed A. Hadad, Asha A. Nair, Yuji Zhang 0001, Eric W. Klee, Krishna Rani Kalari, Jean-Pierre A. Kocher |
Bioinform. | 13 |
| 2012 | SAAP-RRBS: streamlined analysis and annotation pipeline for reduced representation bisulfite sequencingabstractUNLABELLED: Reduced representation bisulfite sequencing (RRBS) is a cost-effective approach for genome-wide methylation pattern profiling. Analyzing RRBS sequencing data is challenging and specialized alignment/mapping programs are needed. Although such programs have been developed, a comprehensive solution that provides researchers with good quality and analyzable data is still lacking. To address this need, we have developed a Streamlined Analysis and Annotation Pipeline for RRBS data (SAAP-RRBS) that integrates read quality assessment/clean-up, alignment, methylation data extraction, annotation, reporting and visualization. This package facilitates a rapid transition from sequencing reads to a fully annotated CpG methylation report to biological interpretation. AVAILABILITY AND IMPLEMENTATION: SAAP-RRBS is freely available to non-commercial users at the web site http://ndc.mayo.edu/mayo/research/biostat/stand-alone-packages.cfm. Zhifu Sun, Saurabh Baheti, Sumit Middha, Rahul Kanwar, Yuji Zhang 0001, Andreas S. Beutler, Eric W. Klee, Yan W. Asmann, E. Aubrey Thompson, Jean-Pierre A. Kocher |
Bioinform. | 8 |
| 2007 | Quantitating tissue specificity of human genes to facilitate biomarker discoveryabstractWe describe a method to identify candidate cancer biomarkers by analyzing numeric approximations of tissue specificity of human genes. These approximations were calculated by analyzing predicted tissue expression distributions of genes derived from mapping expressed sequence tags (ESTs) to the human genome sequence using a binary indexing algorithm. Tissue-specificity values facilitated high-throughput analysis of the human genes and enabled the identification of genes highly specific to different tissues. Tissue expression distributions for several genes were compared to estimates obtained from other public gene expression datasets and experimentally validated using quantitative RT-PCR on RNA isolated from several human tissues. Our results demonstrate that most human genes ( approximately 98%) are expressed in many tissues (low specificity), and only a small number of genes possess very specific tissue expression profiles. These genes comprise a rich dataset from which novel therapeutic targets and novel diagnostic serum biomarkers may be selected. George Vasmatzis, Eric W. Klee, Dagmar M. Kube, Terry M. Therneau, Farhad Kosari |
Bioinform. | 2 |
| 2005 | Evaluating eukaryotic secreted protein predictionabstractBACKGROUND: Improvements in protein sequence annotation and an increase in the number of annotated protein databases has fueled development of an increasing number of software tools to predict secreted proteins. Six software programs capable of high throughput and employing a wide range of prediction methods, SignalP 3.0, SignalP 2.0, TargetP 1.01, PrediSi, Phobius, and ProtComp 6.0, are evaluated. RESULTS: Prediction accuracies were evaluated using 372 unbiased, eukaryotic, SwissProt protein sequences. TargetP, SignalP 3.0 maximum S-score and SignalP 3.0 D-score were the most accurate single scores (90-91% accurate). The combination of a positive TargetP prediction, SignalP 2.0 maximum Y-score, and SignalP 3.0 maximum S-score increased accuracy by six percent. CONCLUSION: Single predictive scores could be highly accurate, but almost all accuracies were slightly less than those reported by program authors. Predictive accuracy could be substantially improved by combining scores from multiple methods into a single composite prediction. Eric W. Klee, Lynda B. M. Ellis |
BMC Bioinform. | 1 |