VLDB 2026 Research / reviewers in the wild / expert
Hannah Carter
dblp:86/7240
· DBLP profile ↗
5ranked-venue papers
0as first author
1since 2021 · last 2024
0000-0002-1729-2463ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › genomics
genomic variant analysis |
0.8 | 1 | 2024 | GRIEVOUS: your command-line general for resolving cross-dataset genotype inconsistencies · Bioinform. 2024 |
Bioinformatics and computational biology
cancer genomics |
0.3 | 2 | 2013 | CRAVAT: cancer-related analysis of variants toolkit · Bioinform. 2013 CHASM and SNVBox: toolkit for detecting biologically important single nucleotide mutations in cancer · Bioinform. 2011 |
Bioinformatics and computational biology › genomics
variant annotation |
0.2 | 1 | 2013 | CRAVAT: cancer-related analysis of variants toolkit · Bioinform. 2013 |
Bioinformatics and computational biology › statistical genetics
variant prioritization |
0.2 | 1 | 2013 | CRAVAT: cancer-related analysis of variants toolkit · Bioinform. 2013 |
Bioinformatics and computational biology › cancer genomics
somatic mutation analysis |
0.1 | 1 | 2011 | CHASM and SNVBox: toolkit for detecting biologically important single nucleotide mutations in cancer · Bioinform. 2011 |
Methods — techniques the papers use, named apart from their topics
variant normalization · 0.8predictive scoring · 0.2predictive feature database · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | GRIEVOUS: your command-line general for resolving cross-dataset genotype inconsistenciesabstractSUMMARY: Harmonizing variant indexing and allele assignments across datasets is crucial for data integrity in cross-dataset studies such as multi-cohort genome-wide association studies, meta-analyses, and the development, validation, and application of polygenic risk scores. Ensuring this indexing and allele consistency is a laborious, time-consuming, and error-prone process requiring a certain degree of computational proficiency. Here, we introduce GRIEVOUS, a command-line tool for cross-dataset variant homogenization. By means of an internal database and a custom indexing methodology, GRIEVOUS identifies, formats, and aligns all biallelic single nucleotide polymorphisms (SNPs) across all summary statistic and genotype files of interest. Upon completion of dataset harmonization, GRIEVOUS can also be used to extract the maximal set of biallelic SNPs common to all datasets. AVAILABILITY AND IMPLEMENTATION: GRIEVOUS and all supporting documentation and tutorials can be found at https://github.com/jvtalwar/GRIEVOUS. It is freely and publicly available under the MIT license and can be installed via pip. James V. Talwar, Adam R. Klie, Meghana S. Pagadala, Hannah Carter |
Bioinform. | 4 |
| 2019 | Rare variant phasing using paired tumor: normal sequence dataabstractBACKGROUND: In standard high throughput sequencing analysis, genetic variants are not assigned to a homologous chromosome of origin. This process, called haplotype phasing, can reveal information important for understanding the relationship between genetic variants and biological phenotypes. For example, in genes that carry multiple heterozygous missense variants, phasing resolves whether one or both gene copies are altered. Here, we present a novel approach to phasing variants that takes advantage of unique properties of paired tumor:normal sequencing data from cancer studies. RESULTS: VAF phasing uses changes in variant allele frequency (VAF) between tumor and normal samples in regions of somatic chromosomal gain or loss to phase germline variants. We apply VAF phasing to 6180 samples from the Cancer Genome Atlas (TCGA) and demonstrate that our method is highly concordant with other standard phasing methods, and can phase an average of 33% more variants than other read-backed phasing methods. Using variant annotation tools designed to score gene haplotypes, we find a suggestive association between carrying multiple missense variants in a single copy of a cancer predisposition gene and earlier age of cancer diagnosis. CONCLUSIONS: VAF phasing exploits unique properties of tumor genomes to increase the number of germline variants that can be phased over standard read-backed methods in paired tumor:normal samples. Our phase-informed association testing results call attention to the need to develop more tools for assessing the joint effect of multiple genetic variants. Alexandra R. Buckley, Trey Ideker, Hannah Carter, Nicholas J. Schork |
BMC Bioinform. | 3 |
| 2014 | A Probabilistic Model to Predict Clinical Phenotypic Traits from Genome SequencingabstractGenetic screening is becoming possible on an unprecedented scale. However, its utility remains controversial. Although most variant genotypes cannot be easily interpreted, many individuals nevertheless attempt to interpret their genetic information. Initiatives such as the Personal Genome Project (PGP) and Illumina's Understand Your Genome are sequencing thousands of adults, collecting phenotypic information and developing computational pipelines to identify the most important variant genotypes harbored by each individual. These pipelines consider database and allele frequency annotations and bioinformatics classifications. We propose that the next step will be to integrate these different sources of information to estimate the probability that a given individual has specific phenotypes of clinical interest. To this end, we have designed a Bayesian probabilistic model to predict the probability of dichotomous phenotypes. When applied to a cohort from PGP, predictions of Gilbert syndrome, Graves' disease, non-Hodgkin lymphoma, and various blood groups were accurate, as individuals manifesting the phenotype in question exhibited the highest, or among the highest, predicted probabilities. Thirty-eight PGP phenotypes (26%) were predicted with area-under-the-ROC curve (AUC)>0.7, and 23 (15.8%) of these were statistically significant, based on permutation tests. Moreover, in a Critical Assessment of Genome Interpretation (CAGI) blinded prediction experiment, the models were used to match 77 PGP genomes to phenotypic profiles, generating the most accurate prediction of 16 submissions, according to an independent assessor. Although the models are currently insufficiently accurate for diagnostic utility, we expect their performance to improve with growth of publicly available genomics data and model refinement by domain experts. Yun-Ching Chen, Christopher Douville, Noushin Niknafs, Grace H. T. Yeo, Violeta Beleva Guthrie, Hannah Carter, Peter D. Stenson, David N. Cooper, Sean D. Mooney, Rachel Karchin |
PLoS Comput. Biol. | 7 |
| 2013 | CRAVAT: cancer-related analysis of variants toolkitabstractSUMMARY: Advances in sequencing technology have greatly reduced the costs incurred in collecting raw sequencing data. Academic laboratories and researchers therefore now have access to very large datasets of genomic alterations but limited time and computational resources to analyse their potential biological importance. Here, we provide a web-based application, Cancer-Related Analysis of Variants Toolkit, designed with an easy-to-use interface to facilitate the high-throughput assessment and prioritization of genes and missense alterations important for cancer tumorigenesis. Cancer-Related Analysis of Variants Toolkit provides predictive scores for germline variants, somatic mutations and relative gene importance, as well as annotations from published literature and databases. Results are emailed to users as MS Excel spreadsheets and/or tab-separated text files. AVAILABILITY: http://www.cravat.us/ Christopher Douville, Hannah Carter, Rick Kim, Noushin Niknafs, Mark Diekhans, Peter D. Stenson, David N. Cooper, Michael C. Ryan, Rachel Karchin |
Bioinform. | 2 |
| 2011 | CHASM and SNVBox: toolkit for detecting biologically important single nucleotide mutations in cancerabstractSUMMARY: Thousands of cancer exomes are currently being sequenced, yielding millions of non-synonymous single nucleotide variants (SNVs) of possible relevance to disease etiology. Here, we provide a software toolkit to prioritize SNVs based on their predicted contribution to tumorigenesis. It includes a database of precomputed, predictive features covering all positions in the annotated human exome and can be used either stand-alone or as part of a larger variant discovery pipeline. AVAILABILITY AND IMPLEMENTATION: MySQL database, source code and binaries freely available for academic/government use at http://wiki.chasmsoftware.org, Source in Python and C++. Requires 32 or 64-bit Linux system (tested on Fedora Core 8,10,11 and Ubuntu 10), 2.5*≤ Python <3.0*, MySQL server >5.0, 60 GB available hard disk space (50 MB for software and data files, 40 GB for MySQL database dump when uncompressed), 2 GB of RAM. Wing Chung Wong, Dewey Kim, Hannah Carter, Mark Diekhans, Michael C. Ryan, Rachel Karchin |
Bioinform. | 3 |