VLDB 2026 Research / reviewers in the wild / expert
Rachel Karchin
dblp:80/263
· DBLP profile ↗
18ranked-venue papers
6as first author
1since 2021 · last 2022
0000-0002-5069-1239ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 6 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
10 papers |
Bioinformatics and computational biology · 100% |
Topics — the 22 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
cancer genomics |
0.9 | 3 | 2022 | Estimation of cancer cell fractions and clone trees from multi-region sequencing of tumors · Bioinform. 2022 CRAVAT: cancer-related analysis of variants toolkit · Bioinform. 2013 CHASM and SNVBox: toolkit for detecting biologically important single nucleotide mutations in cancer · Bioinform. 2011 |
Bioinformatics and computational biology › cancer genomics
tumor evolution |
0.6 | 1 | 2022 | Estimation of cancer cell fractions and clone trees from multi-region sequencing of tumors · Bioinform. 2022 |
Bioinformatics and computational biology › protein function prediction
functional impact prediction |
0.2 | 2 | 2011 | Using bioinformatics to predict the functional impact of SNVs · Bioinform. 2011 LS-SNP: large-scale annotation of coding non-synonymous SNPs based on multiple information sources · Bioinform. 2005 |
Bioinformatics and computational biology › genomics
variant annotation |
0.2 | 1 | 2013 | CRAVAT: cancer-related analysis of variants toolkit · Bioinform. 2013 |
Bioinformatics and computational biology › statistical genetics
variant prioritization |
0.2 | 1 | 2013 | CRAVAT: cancer-related analysis of variants toolkit · Bioinform. 2013 |
Bioinformatics and computational biology
protein structure prediction |
0.1 | 2 | 2008 | PREDICT-2ND: a tool for generalized protein local structure prediction · Bioinform. 2008 Calibrating E-values for hidden Markov models using reverse-sequence null models · Bioinform. 2005 |
Bioinformatics and computational biology › cancer genomics
somatic mutation analysis |
0.1 | 1 | 2011 | CHASM and SNVBox: toolkit for detecting biologically important single nucleotide mutations in cancer · Bioinform. 2011 |
Bioinformatics and computational biology › genomics
variant analysis |
0.1 | 1 | 2011 | Using bioinformatics to predict the functional impact of SNVs · Bioinform. 2011 |
Bioinformatics and computational biology
genomics |
0.1 | 1 | 2009 | LS-SNP/PDB: annotated non-synonymous SNPs mapped to Protein Data Bank structures · Bioinform. 2009 |
Bioinformatics and computational biology › genomics › variant annotation
SNP annotation |
0.1 | 1 | 2009 | LS-SNP/PDB: annotated non-synonymous SNPs mapped to Protein Data Bank structures · Bioinform. 2009 |
Bioinformatics and computational biology › protein structure prediction › protein structural feature prediction
local structure prediction |
0.1 | 1 | 2008 | PREDICT-2ND: a tool for generalized protein local structure prediction · Bioinform. 2008 |
Bioinformatics and computational biology › protein structure prediction › template-based modeling
fold recognition |
0.1 | 1 | 2005 | Calibrating E-values for hidden Markov models using reverse-sequence null models · Bioinform. 2005 |
Bioinformatics and computational biology › genome annotation
genomic variant annotation |
0.1 | 1 | 2005 | LS-SNP: large-scale annotation of coding non-synonymous SNPs based on multiple information sources · Bioinform. 2005 |
Bioinformatics and computational biology
protein sequence analysis |
0.1 | 1 | 2005 | Calibrating E-values for hidden Markov models using reverse-sequence null models · Bioinform. 2005 |
Bioinformatics and computational biology › sequence analysis
sequence similarity search |
0.1 | 1 | 2005 | Calibrating E-values for hidden Markov models using reverse-sequence null models · Bioinform. 2005 |
Bioinformatics and computational biology › protein function prediction › protein classification
GPCR classification |
0.0 | 1 | 2002 | Classifying G-protein coupled receptors with support vector machines · Bioinform. 2002 |
Bioinformatics and computational biology
protein function prediction |
0.0 | 1 | 2002 | Classifying G-protein coupled receptors with support vector machines · Bioinform. 2002 |
Bioinformatics and computational biology › structural bioinformatics › protein structure representation
protein structure visualization |
0.0 | 1 | 2009 | LS-SNP/PDB: annotated non-synonymous SNPs mapped to Protein Data Bank structures · Bioinform. 2009 |
Bioinformatics and computational biology
structural biology |
0.0 | 1 | 2009 | LS-SNP/PDB: annotated non-synonymous SNPs mapped to Protein Data Bank structures · Bioinform. 2009 |
Bioinformatics and computational biology › sequence analysis
profile hidden markov model |
0.0 | 1 | 1998 | Weighting hidden Markov models for maximum discrimination · Bioinform. 1998 |
Bioinformatics and computational biology
sequence analysis |
0.0 | 1 | 1998 | Weighting hidden Markov models for maximum discrimination · Bioinform. 1998 |
Bioinformatics and computational biology › sequence analysis
sequence weighting |
0.0 | 1 | 1998 | Weighting hidden Markov models for maximum discrimination · Bioinform. 1998 |
Methods — techniques the papers use, named apart from their topics
ensemble visualization · 0.6bayesian hierarchical model · 0.6predictive scoring · 0.2predictive feature database · 0.1bioinformatics prediction tools · 0.1sequence profile · 0.1multi-layer neural networks · 0.1hidden markov model · 0.1protein structure modeling · 0.1pathway mapping · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Estimation of cancer cell fractions and clone trees from multi-region sequencing of tumorsabstractMOTIVATION: Multi-region sequencing of solid tumors can improve our understanding of intratumor subclonal diversity and the evolutionary history of mutational events. Due to uncertainty in clonal composition and the multitude of possible ancestral relationships between clones, elucidating the most probable relationships from bulk tumor sequencing poses statistical and computational challenges. RESULTS: We developed a Bayesian hierarchical model called PICTograph to model uncertainty in assigning mutations to subclones, to enable posterior distributions of cancer cell fractions (CCFs) and to visualize the most probable ancestral relationships between subclones. Compared with available methods, PICTograph provided more consistent and accurate estimates of CCFs and improved tree inference over a range of simulated clonal diversity. Application of PICTograph to multi-region whole-exome sequencing of tumors from individuals with pancreatic cancer precursor lesions confirmed known early-occurring mutations and indicated substantial molecular diversity, including 6-12 distinct subclones and intra-sample mixing of subclones. Using ensemble-based visualizations, we highlight highly probable evolutionary relationships recovered in multiple models. PICTograph provides a useful approximation to evolutionary inference from cross-sectional multi-region sequencing, particularly for complex cases. AVAILABILITY AND IMPLEMENTATION: https://github.com/KarchinLab/pictograph. The data underlying this article will be shared on reasonable request to the corresponding author. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lily Zheng, Noushin Niknafs, Laura D. Wood, Rachel Karchin, Robert B. Scharpf |
Bioinform. | 4 |
| 2017 | A novel approach for selecting combination clinical markers of pathology applied to a large retrospective cohort of surgically resected pancreatic cystsabstractOBJECTIVE: Our objective was to develop an approach for selecting combinatorial markers of pathology from diverse clinical data types. We demonstrate this approach on the problem of pancreatic cyst classification. MATERIALS AND METHODS: We analyzed 1026 patients with surgically resected pancreatic cysts, comprising 584 intraductal papillary mucinous neoplasms, 332 serous cystadenomas, 78 mucinous cystic neoplasms, and 42 solid-pseudopapillary neoplasms. To derive optimal markers for cyst classification from the preoperative clinical and radiological data, we developed a statistical approach for combining any number of categorical, dichotomous, or continuous-valued clinical parameters into individual predictors of pathology. The approach is unbiased and statistically rigorous. Millions of feature combinations were tested using 10-fold cross-validation, and the most informative features were validated in an independent cohort of 130 patients with surgically resected pancreatic cysts. RESULTS: We identified combinatorial clinical markers that classified serous cystadenomas with 95% sensitivity and 83% specificity; solid-pseudopapillary neoplasms with 89% sensitivity and 86% specificity; mucinous cystic neoplasms with 91% sensitivity and 83% specificity; and intraductal papillary mucinous neoplasms with 94% sensitivity and 90% specificity. No individual features were as accurate as the combination markers. We further validated these combinatorial markers on an independent cohort of 130 pancreatic cysts, and achieved high and well-balanced accuracies. Overall sensitivity and specificity for identifying patients requiring surgical resection was 84% and 81%, respectively. CONCLUSIONS: Our approach identified combinatorial markers for pancreatic cyst classification that had improved performance relative to the individual features they comprise. In principle, this approach can be applied to any clinical dataset comprising dichotomous, categorical, and continuous-valued parameters. David L. Masica, Marco Dal Molin, Christopher L. Wolfgang, Tyler M. Tomita, Mohammad R. Ostovaneh, Amanda Blackford, Robert A. Moran, Joanna K. Law, Thomas Barkley, Michael Goggins, Marcia Irene Canto, Meredith E. Pittman, James R. Eshleman, Syed Z. Ali, Elliot K. Fishman, Ihab R. Kamel, Siva P. Raman, Atif K. Zaheer, Nita Ahuja, Martin A. Makary, Matthew J. Weiss, Kenzo Hirose, John L. Cameron, Neda Rezaee, Young Joon Ahn, Yuxuan Wang 0011, Simeon Springer, Luis L. Diaz Jr., Nickolas Papadopoulos, Ralph H. Hruban, Kenneth W. Kinzler, Bert Vogelstein, Rachel Karchin, Anne Marie Lennon |
J. Am. Medical Informatics Assoc. | 35 |
| 2016 | Genome Landscapes of Disease: Strategies to Predict the Phenotypic Consequences of Human Germline and Somatic VariationabstractComputational biology can marry disciplines to help solve some of the most pressing problems in medical research.When built on fundamental evolutionary, biological, and/or physics principles and provided with large quantities of diverse experimental data, computing power, and rigorous statistics tools, efficient and effective computational strategies can help unlock the secrets of the genome to cure disease.Computational biology has undertaken this challenge, spearheading efforts to decode the genetic blueprints encrypted in the human germline and in somatic variations to decipher complex phenotypic consequences.It devises new and creative approaches to sift through the immense-and rapidly growing-assembled human genetic material, its products, the proteome and the metabolome, and their possible clinical ramifications.The cell is the basic unit of life; the nucleus houses the genetic material that is believed to have provided a distinct advantage to the evolving cell.The organization of the genome varies; it depends on cell type, stage of development, differentiation, disease status, and more.The higher-order spatial and temporal organization of genomes-which itself is a function of conditions and environment-is a driver of biological function in differentiation, development, and disease.Studies of genomics, epigenetics, big data analysis, imaging, and clinical cell and molecular biology all benefit from rigorous computational biology analyses and modeling.They profit from the testable hypotheses that computational biology provides.These link the genetic material to the physiological cell state; however, in-depth understanding of-and predicting-the phenotypic relationship to genome landscapes also requires effective algorithms to unravel the outcomes of small-and large-scale alterations on DNA, RNA, and protein molecules.Diseases can be viewed as perturbed states of molecular systems.They can be of different types, including single-gene (monogenic) diseases and multifactorial diseases, such as cancers, immune system diseases, neurodegenerative diseases, cardiovascular diseases, and metabolic diseases.They may involve genetic alterations or be more complex.They can also be infectious diseases where interacting molecular networks of both pathogens and humans are involved, with the pathogen protein subverting the cell's machinery.The rapid development of next-generation methods for whole-genome, whole-exome, and targeted sequencing, complemented by earlier microarray technologies, including comparative Rachel Karchin, Ruth Nussinov |
PLoS Comput. Biol. | 1 |
| 2016 | Towards Increasing the Clinical Relevance of In Silico Methods to Predict Pathogenic Missense VariantsabstractAs genetic sequencing throughput continues to accelerate, so does the accumulation of variants of unknown clinical significance.The great majority of these variants cause amino acid substitutions (cSNVs) in protein sequence.The need to interpret these variants continues to motivate development of better in silico bioinformatic methods.Despite the development of dozens of such methods over the past 15 years, clinically relevant prediction accuracy remains elusive.Here, we present some recent progress and shortcomings in the development of bioinformatics missense variant classifiers, and we argue for the increased use of endophenotypes.Endophenotypes are quantitative measurements that are correlated with phenotypes via shared genetic causes (e.g., enzyme catalytic activity, serum cholesterol or glucose level, volumetric lung capacity).In many cases, endophenotypes are more directly influenced by genetic variation, increasing their power to detect genotype-endophenotype associations relative to genotypephenotype associations.The data required to train and benchmark bioinformatic methods to predict endophenotype from cSNVs and other variant types is increasingly available and could be made widely available by concerted community effort to enhance locus-specific and disease variant databases.We highlight some currently available data and present results from published bioinformatics studies that use endophenotypes. David L. Masica, Rachel Karchin |
PLoS Comput. Biol. | 2 |
| 2015 | SubClonal Hierarchy Inference from Somatic Mutations: Automatic Reconstruction of Cancer Evolutionary Trees from Multi-region Next Generation SequencingabstractRecent improvements in next-generation sequencing of tumor samples and the ability to identify somatic mutations at low allelic fractions have opened the way for new approaches to model the evolution of individual cancers. The power and utility of these models is increased when tumor samples from multiple sites are sequenced. Temporal ordering of the samples may provide insight into the etiology of both primary and metastatic lesions and rationalizations for tumor recurrence and therapeutic failures. Additional insights may be provided by temporal ordering of evolving subclones--cellular subpopulations with unique mutational profiles. Current methods for subclone hierarchy inference tightly couple the problem of temporal ordering with that of estimating the fraction of cancer cells harboring each mutation. We present a new framework that includes a rigorous statistical hypothesis test and a collection of tools that make it possible to decouple these problems, which we believe will enable substantial progress in the field of subclone hierarchy inference. The methods presented here can be flexibly combined with methods developed by others addressing either of these problems. We provide tools to interpret hypothesis test results, which inform phylogenetic tree construction, and we introduce the first genetic algorithm designed for this purpose. The utility of our framework is systematically demonstrated in simulations. For most tested combinations of tumor purity, sequencing coverage, and tree complexity, good power (≥ 0.8) can be achieved and Type 1 error is well controlled when at least three tumor samples are available from a patient. Using data from three published multi-region tumor sequencing studies of (murine) small cell lung cancer, acute myeloid leukemia, and chronic lymphocytic leukemia, in which the authors reconstructed subclonal phylogenetic trees by manual expert curation, we show how different configurations of our tools can identify either a single tree in agreement with the authors, or a small set of trees, which include the authors' preferred tree. Our results have implications for improved modeling of tumor evolution and the importance of multi-region tumor sequencing. Noushin Niknafs, Violeta Beleva Guthrie, Daniel Q. Naiman, Rachel Karchin |
PLoS Comput. Biol. | 4 |
| 2014 | A Probabilistic Model to Predict Clinical Phenotypic Traits from Genome SequencingabstractGenetic screening is becoming possible on an unprecedented scale. However, its utility remains controversial. Although most variant genotypes cannot be easily interpreted, many individuals nevertheless attempt to interpret their genetic information. Initiatives such as the Personal Genome Project (PGP) and Illumina's Understand Your Genome are sequencing thousands of adults, collecting phenotypic information and developing computational pipelines to identify the most important variant genotypes harbored by each individual. These pipelines consider database and allele frequency annotations and bioinformatics classifications. We propose that the next step will be to integrate these different sources of information to estimate the probability that a given individual has specific phenotypes of clinical interest. To this end, we have designed a Bayesian probabilistic model to predict the probability of dichotomous phenotypes. When applied to a cohort from PGP, predictions of Gilbert syndrome, Graves' disease, non-Hodgkin lymphoma, and various blood groups were accurate, as individuals manifesting the phenotype in question exhibited the highest, or among the highest, predicted probabilities. Thirty-eight PGP phenotypes (26%) were predicted with area-under-the-ROC curve (AUC)>0.7, and 23 (15.8%) of these were statistically significant, based on permutation tests. Moreover, in a Critical Assessment of Genome Interpretation (CAGI) blinded prediction experiment, the models were used to match 77 PGP genomes to phenotypic profiles, generating the most accurate prediction of 16 submissions, according to an independent assessor. Although the models are currently insufficiently accurate for diagnostic utility, we expect their performance to improve with growth of publicly available genomics data and model refinement by domain experts. Yun-Ching Chen, Christopher Douville, Noushin Niknafs, Grace H. T. Yeo, Violeta Beleva Guthrie, Hannah Carter, Peter D. Stenson, David N. Cooper, Sean D. Mooney, Rachel Karchin |
PLoS Comput. Biol. | 12 |
| 2013 | CRAVAT: cancer-related analysis of variants toolkitabstractSUMMARY: Advances in sequencing technology have greatly reduced the costs incurred in collecting raw sequencing data. Academic laboratories and researchers therefore now have access to very large datasets of genomic alterations but limited time and computational resources to analyse their potential biological importance. Here, we provide a web-based application, Cancer-Related Analysis of Variants Toolkit, designed with an easy-to-use interface to facilitate the high-throughput assessment and prioritization of genes and missense alterations important for cancer tumorigenesis. Cancer-Related Analysis of Variants Toolkit provides predictive scores for germline variants, somatic mutations and relative gene importance, as well as annotations from published literature and databases. Results are emailed to users as MS Excel spreadsheets and/or tab-separated text files. AVAILABILITY: http://www.cravat.us/ Christopher Douville, Hannah Carter, Rick Kim, Noushin Niknafs, Mark Diekhans, Peter D. Stenson, David N. Cooper, Michael C. Ryan, Rachel Karchin |
Bioinform. | 9 |
| 2011 | Using bioinformatics to predict the functional impact of SNVsabstractMOTIVATION: The past decade has seen the introduction of fast and relatively inexpensive methods to detect genetic variation across the genome and exponential growth in the number of known single nucleotide variants (SNVs). There is increasing interest in bioinformatics approaches to identify variants that are functionally important from millions of candidate variants. Here, we describe the essential components of bioinformatics tools that predict functional SNVs. RESULTS: Bioinformatics tools have great potential to identify functional SNVs, but the black box nature of many tools can be a pitfall for researchers. Understanding the underlying methods, assumptions and biases of these tools is essential to their intelligent application. Melissa S. Cline, Rachel Karchin |
Bioinform. | 2 |
| 2011 | CHASM and SNVBox: toolkit for detecting biologically important single nucleotide mutations in cancerabstractSUMMARY: Thousands of cancer exomes are currently being sequenced, yielding millions of non-synonymous single nucleotide variants (SNVs) of possible relevance to disease etiology. Here, we provide a software toolkit to prioritize SNVs based on their predicted contribution to tumorigenesis. It includes a database of precomputed, predictive features covering all positions in the annotated human exome and can be used either stand-alone or as part of a larger variant discovery pipeline. AVAILABILITY AND IMPLEMENTATION: MySQL database, source code and binaries freely available for academic/government use at http://wiki.chasmsoftware.org, Source in Python and C++. Requires 32 or 64-bit Linux system (tested on Fedora Core 8,10,11 and Ubuntu 10), 2.5*≤ Python <3.0*, MySQL server >5.0, 60 GB available hard disk space (50 MB for software and data files, 40 GB for MySQL database dump when uncompressed), 2 GB of RAM. Wing Chung Wong, Dewey Kim, Hannah Carter, Mark Diekhans, Michael C. Ryan, Rachel Karchin |
Bioinform. | 6 |
| 2011 | Network Models of TEM β-Lactamase Mutations Coevolving under Antibiotic Selection Show Modular Structure and Anticipate Evolutionary TrajectoriesabstractUnderstanding how novel functions evolve (genetic adaptation) is a critical goal of evolutionary biology. Among asexual organisms, genetic adaptation involves multiple mutations that frequently interact in a non-linear fashion (epistasis). Non-linear interactions pose a formidable challenge for the computational prediction of mutation effects. Here we use the recent evolution of β-lactamase under antibiotic selection as a model for genetic adaptation. We build a network of coevolving residues (possible functional interactions), in which nodes are mutant residue positions and links represent two positions found mutated together in the same sequence. Most often these pairs occur in the setting of more complex mutants. Focusing on extended-spectrum resistant sequences, we use network-theoretical tools to identify triple mutant trajectories of likely special significance for adaptation. We extrapolate evolutionary paths (n = 3) that increase resistance and that are longer than the units used to build the network (n = 2). These paths consist of a limited number of residue positions and are enriched for known triple mutant combinations that increase cefotaxime resistance. We find that the pairs of residues used to build the network frequently decrease resistance compared to their corresponding singlets. This is a surprising result, given that their coevolution suggests a selective advantage. Thus, β-lactamase adaptation is highly epistatic. Our method can identify triplets that increase resistance despite the underlying rugged fitness landscape and has the unique ability to make predictions by placing each mutant residue position in its functional context. Our approach requires only sequence information, sufficient genetic diversity, and discrete selective pressures. Thus, it can be used to analyze recent evolutionary events, where coevolution analysis methods that use phylogeny or statistical coupling are not possible. Improving our ability to assess evolutionary trajectories will help predict the evolution of clinically relevant genes and aid in protein design. Violeta Beleva Guthrie, Jennifer Allen, Manel Camps, Rachel Karchin |
PLoS Comput. Biol. | 4 |
| 2009 | Next generation tools for the annotation of human SNPsabstractComputational biology has the opportunity to play an important role in the identification of functional single nucleotide polymorphisms (SNPs) discovered in large-scale genotyping studies, ultimately yielding new drug targets and biomarkers. The medical genetics and molecular biology communities are increasingly turning to computational biology methods to prioritize interesting SNPs found in linkage and association studies. Many such methods are now available through web interfaces, but the interested user is confronted with an array of predictive results that are often in disagreement with each other. Many tools today produce results that are difficult to understand without bioinformatics expertise, are biased towards non-synonymous SNPs, and do not necessarily reflect up-to-date versions of their source bioinformatics resources, such as public SNP repositories. Here, I assess the utility of the current generation of webservers; and suggest improvements for the next generation of webservers to better deliver value to medical geneticists and molecular biologists. Rachel Karchin |
Briefings Bioinform. | 1 |
| 2009 | LS-SNP/PDB: annotated non-synonymous SNPs mapped to Protein Data Bank structuresabstractSUMMARY: LS-SNP/PDB is a new WWW resource for genome-wide annotation of human non-synonymous (amino acid changing) SNPs. It serves high-quality protein graphics rendered with UCSF Chimera molecular visualization software. The system is kept up-to-date by an automated, high-throughput build pipeline that systematically maps human nsSNPs onto Protein Data Bank structures and annotates several biologically relevant features. AVAILABILITY: LS-SNP/PDB is available at (http://ls-snp.icm.jhu.edu/ls-snp-pdb) and via links from protein data bank (PDB) biology and chemistry tabs, UCSC Genome Browser Gene Details and SNP Details pages and PharmGKB Gene Variants Downloads/Cross-References pages. Michael C. Ryan, Mark Diekhans, Stephanie Lien, Yun Liu 0013, Rachel Karchin |
Bioinform. | 5 |
| 2008 | PREDICT-2ND: a tool for generalized protein local structure predictionabstractMOTIVATION: Predictions of protein local structure, derived from sequence alignment information alone, provide visualization tools for biologists to evaluate the importance of amino acid residue positions of interest in the absence of X-ray crystal/NMR structures or homology models. They are also useful as inputs to sequence analysis and modeling tools, such as hidden Markov models (HMMs), which can be used to search for homology in databases of known protein structure. In addition, local structure predictions can be used as a component of cost functions in genetic algorithms that predict protein tertiary structure. We have developed a program (predict-2nd) that trains multilayer neural networks and have applied it to numerous local structure alphabets, tuning network parameters such as the number of layers, the number of units in each layer and the window sizes of each layer. We have had the most success with four-layer networks, with gradually increasing window sizes at each layer. RESULTS: Because the four-layer neural nets occasionally get trapped in poor local optima, our training protocol now uses many different random starts, with short training runs, followed by more training on the best performing networks from the short runs. One recent addition to the program is the option to add a guide sequence to the profile inputs, increasing the number of inputs per position by 20. We find that use of a guide sequence provides a small but consistent improvement in the predictions for several different local-structure alphabets. AVAILABILITY: Local structure prediction with the methods described here is available for use online at http://www.soe.ucsc.edu/compbio/SAM_T08/T08-query.html. The source code and example networks for PREDICT-2ND are available at http://www.soe.ucsc.edu/~karplus/predict-2nd/ A required C++ library is available at http://www.soe.ucsc.edu/~karplus/ultimate/ Sol Katzman, Christian Barrett, Grant Thiltgen, Rachel Karchin, Kevin Karplus |
Bioinform. | 4 |
| 2007 | Functional Impact of Missense Variants in BRCA1 Predicted by Supervised LearningabstractMany individuals tested for inherited cancer susceptibility at the BRCA1 gene locus are discovered to have variants of unknown clinical significance (UCVs). Most UCVs cause a single amino acid residue (missense) change in the BRCA1 protein. They can be biochemically assayed, but such evaluations are time-consuming and labor-intensive. Computational methods that classify and suggest explanations for UCV impact on protein function can complement functional tests. Here we describe a supervised learning approach to classification of BRCA1 UCVs. Using a novel combination of 16 predictive features, the algorithms were applied to retrospectively classify the impact of 36 BRCA1 C-terminal (BRCT) domain UCVs biochemically assayed to measure transactivation function and to blindly classify 54 documented UCVs. Majority vote of three supervised learning algorithms is in agreement with the assay for more than 94% of the UCVs. Two UCVs found deleterious by both the assay and the classifiers reveal a previously uncharacterized putative binding site. Clinicians may soon be able to use computational classifiers such as those described here to better inform patients. These classifiers can be adapted to other cancer susceptibility genes and systematically applied to prioritize the growing number of potential causative loci and variants found by large-scale disease association studies. Rachel Karchin, Alvaro N. A. Monteiro, Sean V. Tavtigian, Marcelo A. Carvalho, Andrej Sali |
PLoS Comput. Biol. | 1 |
| 2005 | LS-SNP: large-scale annotation of coding non-synonymous SNPs based on multiple information sourcesabstractMOTIVATION: The NCBI dbSNP database lists over 9 million single nucleotide polymorphisms (SNPs) in the human genome, but currently contains limited annotation information. SNPs that result in amino acid residue changes (nsSNPs) are of critical importance in variation between individuals, including disease and drug sensitivity. RESULTS: We have developed LS-SNP, a genomic scale software pipeline to annotate nsSNPs. LS-SNP comprehensively maps nsSNPs onto protein sequences, functional pathways and comparative protein structure models, and predicts positions where nsSNPs destabilize proteins, interfere with the formation of domain-domain interfaces, have an effect on protein-ligand binding or severely impact human health. It currently annotates 28,043 validated SNPs that produce amino acid residue substitutions in human proteins from the SwissProt/TrEMBL database. Annotations can be viewed via a web interface either in the context of a genomic region or by selecting sets of SNPs, genes, proteins or pathways. These results are useful for identifying candidate functional SNPs within a gene, haplotype or pathway and in probing molecular mechanisms responsible for functional impacts of nsSNPs. AVAILABILITY: http://www.salilab.org/LS-SNP CONTACT: [email protected] SUPPLEMENTARY INFORMATION: http://salilab.org/LS-SNP/supp-info.pdf. Rachel Karchin, Mark Diekhans, Libusha Kelly, Daryl J. Thomas, Ursula Pieper 0001, Narayanan Eswar, David Haussler, Andrej Sali |
Bioinform. | 1 |
| 2005 | Calibrating E-values for hidden Markov models using reverse-sequence null modelsabstractMOTIVATION: Hidden Markov models (HMMs) calculate the probability that a sequence was generated by a given model. Log-odds scoring provides a context for evaluating this probability, by considering it in relation to a null hypothesis. We have found that using a reverse-sequence null model effectively removes biases owing to sequence length and composition and reduces the number of false positives in a database search. Any scoring system is an arbitrary measure of the quality of database matches. Significance estimates of scores are essential, because they eliminate model- and method-dependent scaling factors, and because they quantify the importance of each match. Accurate computation of the significance of reverse-sequence null model scores presents a problem, because the scores do not fit the extreme-value (Gumbel) distribution commonly used to estimate HMM scores' significance. RESULTS: To get a better estimate of the significance of reverse-sequence null model scores, we derive a theoretical distribution based on the assumption of a Gumbel distribution for raw HMM scores and compare estimates based on this and other distribution families. We derive estimation methods for the parameters of the distributions based on maximum likelihood and on moment matching (least-squares fit for Student's t-distribution). We evaluate the modeled distributions of scores, based on how well they fit the tail of the observed distribution for data not used in the fitting and on the effects of the improved E-values on our HMM-based fold-recognition methods. The theoretical distribution provides some improvement in fitting the tail and in providing fewer false positives in the fold-recognition test. An ad hoc distribution based on assuming a stretched exponential tail does an even better job. The use of Student's t to model the distribution fits well in the middle of the distribution, but provides too heavy a tail. The moment-matching methods fit the tails better than maximum-likelihood methods. AVAILABILITY: Information on obtaining the SAM program suite (free for academic use), as well as a server interface, is available at http://www.soe.ucsc.edu/research/compbio/sam.html and the open-source random sequence generator with varying compositional biases is available at http://www.soe.ucsc.edu/research/compbio/gen_sequence Kevin Karplus, Rachel Karchin, George Shackelford, Richard Hughey |
Bioinform. | 2 |
| 2002 | Classifying G-protein coupled receptors with support vector machinesabstractAbstract Motivation: The enormous amount of protein sequence data uncovered by genome research has increased the demand for computer software that can automate the recognition of new proteins. We discuss the relative merits of various automated methods for recognizing G-Protein Coupled Receptors (GPCRs), a superfamily of cell membrane proteins. GPCRs are found in a wide range of organisms and are central to a cellular signalling network that regulates many basic physiological processes. They are the focus of a significant amount of current pharmaceutical research because they play a key role in many diseases. However, their tertiary structures remain largely unsolved. The methods described in this paper use only primary sequence information to make their predictions. We compare a simple nearest neighbor approach (BLAST), methods based on multiple alignments generated by a statistical profile Hidden Markov Model (HMM), and methods, including Support Vector Machines (SVMs), that transform protein sequences into fixed-length feature vectors. Results: The last is the most computationally expensive method, but our experiments show that, for those interested in annotation-quality classification, the results are worth the effort. In two-fold cross-validation experiments testing recognition of GPCR subfamilies that bind a specific ligand (such as a histamine molecule), the errors per sequence at the Minimum Error Point (MEP) were 13.7% for multi-class SVMs, 17.1% for our SVMtree method of hierarchical multi-class SVM classification, 25.5% for BLAST, 30% for profile HMMs, and 49% for classification based on nearest neighbor feature vector Kernel Nearest Neighbor (kernNN). The percentage of true positives recognized before the first false positive was 65% for both SVM methods, 13% for BLAST, 5% for profile HMMs and 4% for kernNN. Availability: We have set up a web server for GPCR subfamily classification based on hierarchical multi-class SVMs at http://www.soe.ucsc.edu/research/compbio/gpcr-subclass. By scanning predicted peptides found in the human genome with the SVMtree server, we have identified a large number of genes that encode GPCRs. A list of our predictions for human GPCRs is available at http://www.soe.ucsc.edu/research/compbio/gpcr˙hg/class˙results. We also provide suggested subfamily classification for 18 sequences previously identified as unclassified Class A (rhodopsin-like) GPCRs in GPCRDB (Horn et al. , Nucleic Acids Res. , 26, 277–281, 1998), available at http://www.soe.ucsc.edu/research/compbio/gpcr/classA˙unclassified/. Contact: [email protected] * To whom correspondence should be addressed. Rachel Karchin, Kevin Karplus, David Haussler |
Bioinform. | 1 |
| 1998 | Weighting hidden Markov models for maximum discriminationabstractMOTIVATION: Hidden Markov models can efficiently and automatically build statistical representations of related sequences. Unfortunately, training sets are frequently biased toward one subgroup of sequences, leading to an insufficiently general model. This work evaluates sequence weighting methods based on the maximum-discrimination idea. RESULTS: One good method scales sequence weights by an exponential that ranges between 0.1 for the best scoring sequence and 1.0 for the worst. Experiments with a curated data set show that while training with one or two sequences performed worse than single-sequence Probabilistic Smith-Waterman, training with five or ten sequences reduced errors by 20% and 51%, respectively. This new version of the SAM HMM suite outperforms HMMer (17% reduction over PSW for 10 training sequences), Meta-MEME (28% reduction), and unweighted SAM (31% reduction). AVAILABILITY: A WWW server, as well as information on obtaining the Sequence Alignment and Modeling (SAM) software suite and additional data from this work, can be found at http://www.cse.ucse. edu/research/compbio/sam.html Rachel Karchin, Richard Hughey |
Bioinform. | 1 |