Keith A. Crandall

dblp:35/5987 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
4since 2021 · last 2022
0000-0002-0836-3389ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 4 since 2021
YearPublicationVenuePosition
2022 Synergies between centralized and federated approaches to data quality: a report from the national COVID cohort collaborative
abstract
OBJECTIVE: In response to COVID-19, the informatics community united to aggregate as much clinical data as possible to characterize this new disease and reduce its impact through collaborative analytics. The National COVID Cohort Collaborative (N3C) is now the largest publicly available HIPAA limited dataset in US history with over 6.4 million patients and is a testament to a partnership of over 100 organizations. MATERIALS AND METHODS: We developed a pipeline for ingesting, harmonizing, and centralizing data from 56 contributing data partners using 4 federated Common Data Models. N3C data quality (DQ) review involves both automated and manual procedures. In the process, several DQ heuristics were discovered in our centralized context, both within the pipeline and during downstream project-based analysis. Feedback to the sites led to many local and centralized DQ improvements. RESULTS: Beyond well-recognized DQ findings, we discovered 15 heuristics relating to source Common Data Model conformance, demographics, COVID tests, conditions, encounters, measurements, observations, coding completeness, and fitness for use. Of 56 sites, 37 sites (66%) demonstrated issues through these heuristics. These 37 sites demonstrated improvement after receiving feedback. DISCUSSION: We encountered site-to-site differences in DQ which would have been challenging to discover using federated checks alone. We have demonstrated that centralized DQ benchmarking reveals unique opportunities for DQ improvement that will support improved research analytics locally and in aggregate. CONCLUSION: By combining rapid, continual assessment of DQ with a large volume of multisite data, it is possible to support more nuanced scientific questions with the scale and rigor that they require.
Emily R. Pfaff, Andrew T. Girvin, Davera Gabriel, Kristin Kostka, Michele Morris, Matvey Palchuk, Harold P. Lehmann, Benjamin R. C. Amor, Mark Bissell, Katie R. Bradwell, Sigfried Gold, Stephanie S. Hong, Johanna Loomba, Amin Manna, Julie A. McMurry, Emily Niehaus, Nabeel Qureshi, Anita Walden, Xiaohan Tanner Zhang, Richard L. Zhu, Richard A. Moffitt, Christopher G. Chute, William G. Adams, Shaymaa Al-Shukri, Alfred Anzalone, Ahmad Baghal, Tellen D. Bennett, Elmer V. Bernstam, Mark M. Bissell, Brian Bush, Thomas R. Campion Jr., Victor Castro, Jack Chang, Deepa D. Chaudhari, Wenjin Chen, San Chu, James J. Cimino, Keith A. Crandall, Mark Crooks, Sara J. Deakyne Davies, John Dipalazzo, David A. Dorr, Daniel Eckrich, Sarah E. Eltinge, Daniel G. Fort, Georgiy Golovko, Snehil Gupta, Melissa A. Haendel, Janos G. Hajagos, David A. Hanauer, Brett M. Harnett, Ronald Horswell, Nancy Huang, Steven G. Johnson, Michael Kahn, Kamil Khanipov, Curtis Kieler, Katherine Ruiz De Luzuriaga, Sarah E. Maidlow, Ashley Martinez, Jomol Mathew, James C. McClay, Gabriel McMahan, Brian Melancon, Stéphane M. Meystre, Lucio Miele, Hiroki Morizono, Ray Pablo, Lav P. Patel, Jimmy Phuong, Daniel J. Popham, Claudia P. Pulgarin, Indra Neil Sarkar, Nancy Sazo, Soko Setoguchi, Selvin Soby, Sirisha Surampalli, Christine Suver, Uma Maheswara Reddy Vangala, Shyam Visweswaran, James von Oehsen, Kellie M. Walters, Laura K. Wiley, David A. Williams, Adrian H. Zai
J. Am. Medical Informatics Assoc.38
2021 COVID-19 biomarkers and their overlap with comorbidities in a disease biomarker data model
abstract
In response to the COVID-19 outbreak, scientists and medical researchers are capturing a wide range of host responses, symptoms and lingering postrecovery problems within the human population. These variable clinical manifestations suggest differences in influential factors, such as innate and adaptive host immunity, existing or underlying health conditions, comorbidities, genetics and other factors-compounding the complexity of COVID-19 pathobiology and potential biomarkers associated with the disease, as they become available. The heterogeneous data pose challenges for efficient extrapolation of information into clinical applications. We have curated 145 COVID-19 biomarkers by developing a novel cross-cutting disease biomarker data model that allows integration and evaluation of biomarkers in patients with comorbidities. Most biomarkers are related to the immune (SAA, TNF-∝ and IP-10) or coagulation (D-dimer, antithrombin and VWF) cascades, suggesting complex vascular pathobiology of the disease. Furthermore, we observe commonality with established cancer biomarkers (ACE2, IL-6, IL-4 and IL-2) as well as biomarkers for metabolic syndrome and diabetes (CRP, NLR and LDL). We explore these trends as we put forth a COVID-19 biomarker resource (https://data.oncomx.org/covid19) that will help researchers and diagnosticians alike.
Nikhita Gogate, Daniel Lyman, Amanda Bell, Edmund Cauley, Keith A. Crandall, Ashia Joseph, Robel Y. Kahsay, Darren A. Natale, Lynn M. Schriml, Sabyasach Sen, Raja Mazumder
Briefings Bioinform.5
2021 Omics community detection using multi-resolution clustering
abstract
MOTIVATION: The discovery of biologically interpretable and clinically actionable communities in heterogeneous omics data is a necessary first step toward deriving mechanistic insights into complex biological phenomena. Here, we present a novel clustering approach, omeClust, for community detection in omics profiles by simultaneously incorporating similarities among measurements and the overall complex structure of the data. RESULTS: We show that omeClust outperforms published methods in inferring the true community structure as measured by both sensitivity and misclassification rate on simulated datasets. We further validated omeClust in diverse, multiple omics datasets, revealing new communities and functionally related groups in microbial strains, cell line gene expression patterns and fetal genomic variation. We also derived enrichment scores attributable to putatively meaningful biological factors in these datasets that can serve as hypothesis generators facilitating new sets of testable hypotheses. AVAILABILITY AND IMPLEMENTATION: omeClust is open-source software, and the implementation is available online at http://github.com/omicsEye/omeClust. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Gholamali Rahnavard, Suvo Chatterjee, Bahar Sayoldin, Keith A. Crandall, Fasil Tekola-Ayele, Himel Mallick
Bioinform.4
2021 Each patient is a research biorepository: informatics-enabled research on surplus clinical specimens via the living BioBank
abstract
The ability to analyze human specimens is the pillar of modern-day translational research. To enhance the research availability of relevant clinical specimens, we developed the Living BioBank (LBB) solution, which allows for just-in-time capture and delivery of phenotyped surplus laboratory medicine specimens. The LBB is a system-of-systems integrating research feasibility databases in i2b2, a real-time clinical data warehouse, and an informatics system for institutional research services management (SPARC). LBB delivers deidentified clinical data and laboratory specimens. We further present an extension to our solution, the Living µBiome Bank, that allows the user to request and receive phenotyped specimen microbiome data. We discuss the details of the implementation of the LBB system and the necessary regulatory oversight for this solution. The conducted institutional focus group of translational investigators indicates an overall positive sentiment towards potential scientific results generated with the use of LBB. Reference implementation of LBB is available at https://LivingBioBank.musc.edu.
Alexander V. Alekseyenko, Bashir Hamidi, Trevor D. Faith, Keith A. Crandall, Jennifer G. Powers, Christopher Metts, James E. Madory, Steven L. Carroll, Jihad S. Obeid, Leslie Lenert
J. Am. Medical Informatics Assoc.4
2020 ReQTL: identifying correlations between expressed SNVs and gene expression using RNA-sequencing data
abstract
MOTIVATION: By testing for associations between DNA genotypes and gene expression levels, expression quantitative trait locus (eQTL) analyses have been instrumental in understanding how thousands of single nucleotide variants (SNVs) may affect gene expression. As compared to DNA genotypes, RNA genetic variation represents a phenotypic trait that reflects the actual allele content of the studied system. RNA genetic variation at expressed SNV loci can be estimated using the proportion of alleles bearing the variant nucleotide (variant allele fraction, VAFRNA). VAFRNA is a continuous measure which allows for precise allele quantitation in loci where the RNA alleles do not scale with the genotype count. We describe a method to correlate VAFRNA with gene expression and assess its ability to identify genetically regulated expression solely from RNA-sequencing (RNA-seq) datasets. RESULTS: We introduce ReQTL, an eQTL modification which substitutes the DNA allele count for the variant allele fraction at expressed SNV loci in the transcriptome (VAFRNA). We exemplify the method on sets of RNA-seq data from human tissues obtained though the Genotype-Tissue Expression (GTEx) project and demonstrate that ReQTL analyses are computationally feasible and can identify a subset of expressed eQTL loci. AVAILABILITY AND IMPLEMENTATION: A toolkit to perform ReQTL analyses is available at https://github.com/HorvathLab/ReQTL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Liam F. Spurr, Nawaf Alomran, Pavlos Bousounis, Dacian Reece-Stremtan, N. M. Prashant, Piotr Slowinski, Muzi Li, Justin Sein, Gabriel Asher, Keith A. Crandall, Krasimira Tsaneva-Atanasova, Anelia Horvath
Bioinform.12
2019 Telescope: Characterization of the retrotranscriptome by accurate estimation of transposable element expression
abstract
Characterization of Human Endogenous Retrovirus (HERV) expression within the transcriptomic landscape using RNA-seq is complicated by uncertainty in fragment assignment because of sequence similarity. We present Telescope, a computational software tool that provides accurate estimation of transposable element expression (retrotranscriptome) resolved to specific genomic locations. Telescope directly addresses uncertainty in fragment assignment by reassigning ambiguously mapped fragments to the most probable source transcript as determined within a Bayesian statistical model. We demonstrate the utility of our approach through single locus analysis of HERV expression in 13 ENCODE cell types. When examined at this resolution, we find that the magnitude and breadth of the retrotranscriptome can be vastly different among cell types. Furthermore, our approach is robust to differences in sequencing technology and demonstrates that the retrotranscriptome has potential to be used for cell type identification. We compared our tool with other approaches for quantifying transposable element (TE) expression, and found that Telescope has the greatest resolution, as it estimates expression at specific TE insertions rather than at the TE subfamily level. Telescope performs highly accurate quantification of the retrotranscriptomic landscape in RNA-seq experiments, revealing a differential complexity in the transposable element biology of complex systems not previously observed. Telescope is available at https://github.com/mlbendall/telescope.
Matthew L. Bendall, Miguel de Mulder, Luis Pedro Iñiguez, Aarón Lecanda-Sánchez, Marcos Pérez-Losada, Mario A. Ostrowski, R. Brad Jones, Lubbertus C. F. Mulder, Gustavo Reyes-Terán, Keith A. Crandall, Christopher E. Ormsby, Douglas F. Nixon
PLoS Comput. Biol.10
2017 Genome Polymorphism Detection Through Relaxed de Bruijn Graph Construction
abstract
Comparing genomes to identify polymorphisms is a difficult task, especially beyond single nucleotide poly-morphisms. Polymorphism detection is important in disease association studies as well as in phylogenetic tree reconstruc-tion. We present a method for identifying polymorphisms in genomes by using a modified version de Bruijn graphs, data structures widely used in genome assembly from Next-Generation Sequencing. Using our method, we are able to identify polymorphisms that exist within a genome as well as well as see graph structures that form in the de Bruijn graph for particular types of polymorphisms (translocations, etc.).
M. Stanley Fujimoto, Cole A. Lyman, Anton Suvorov, Paul M. Bodily, Quinn Snell, Keith A. Crandall, Seth M. Bybee, Mark J. Clement
BIBE6
2017 Whole Genome Phylogenetic Tree Reconstruction Using Colored de Bruijn Graphs
abstract
We present kleuren, a novel assembly-free method to reconstruct phylogenetic trees using the Colored de Bruijn Graph. kleuren works by constructing the Colored de Bruijn Graph and then traversing it, finding bubble structures in the graph that provide phylogenetic signal. The bubbles are then aligned and concatenated to form a supermatrix, from which a phylogenetic tree is inferred. We introduce the algorithms that kleuren uses to accomplish this task, and show its performance on reconstructing the phylogenetic tree of 12 Drosophila species. kleuren reconstructed the established phylogenetic tree accurately, and is a viable tool for phylogenetic tree reconstruction using whole genome sequences. Software package available at: https://github.com/Colelyman/kleuren.
Cole A. Lyman, M. Stanley Fujimoto, Anton Suvorov, Paul M. Bodily, Quinn Snell, Keith A. Crandall, Seth M. Bybee, Mark J. Clement
BIBE6
2014 Clinical PathoScope: rapid alignment and filtration for accurate pathogen identification in clinical samples using unassembled sequencing data
abstract
BACKGROUND: The use of sequencing technologies to investigate the microbiome of a sample can positively impact patient healthcare by providing therapeutic targets for personalized disease treatment. However, these samples contain genomic sequences from various sources that complicate the identification of pathogens. RESULTS: Here we present Clinical PathoScope, a pipeline to rapidly and accurately remove host contamination, isolate microbial reads, and identify potential disease-causing pathogens. We have accomplished three essential tasks in the development of Clinical PathoScope. First, we developed an optimized framework for pathogen identification using a computational subtraction methodology in concordance with read trimming and ambiguous read reassignment. Second, we have demonstrated the ability of our approach to identify multiple pathogens in a single clinical sample, accurately identify pathogens at the subspecies level, and determine the nearest phylogenetic neighbor of novel or highly mutated pathogens using real clinical sequencing data. Finally, we have shown that Clinical PathoScope outperforms previously published pathogen identification methods with regard to computational speed, sensitivity, and specificity. CONCLUSIONS: Clinical PathoScope is the only pathogen identification method currently available that can identify multiple pathogens from mixed samples and distinguish between very closely related species and strains in samples with very few reads per pathogen. Furthermore, Clinical PathoScope does not rely on genome assembly and thus can more rapidly complete the analysis of a clinical sample when compared with current assembly-based methods. Clinical PathoScope is freely available at: http://sourceforge.net/projects/pathoscope/.
Allyson L. Byrd, Joseph F. Perez-Rogers, Solaiappan Manimaran, Eduardo Castro-Nallar, Ian Toma, Tim McCaffrey, Marc Siegel, Gary Benson, Keith A. Crandall
BMC Bioinform.9
2014 Using phylogenetically-informed annotation (PIA) to search for light-interacting genes in transcriptomes from non-model organisms
abstract
BACKGROUND: Tools for high throughput sequencing and de novo assembly make the analysis of transcriptomes (i.e. the suite of genes expressed in a tissue) feasible for almost any organism. Yet a challenge for biologists is that it can be difficult to assign identities to gene sequences, especially from non-model organisms. Phylogenetic analyses are one useful method for assigning identities to these sequences, but such methods tend to be time-consuming because of the need to re-calculate trees for every gene of interest and each time a new data set is analyzed. In response, we employed existing tools for phylogenetic analysis to produce a computationally efficient, tree-based approach for annotating transcriptomes or new genomes that we term Phylogenetically-Informed Annotation (PIA), which places uncharacterized genes into pre-calculated phylogenies of gene families. RESULTS: We generated maximum likelihood trees for 109 genes from a Light Interaction Toolkit (LIT), a collection of genes that underlie the function or development of light-interacting structures in metazoans. To do so, we searched protein sequences predicted from 29 fully-sequenced genomes and built trees using tools for phylogenetic analysis in the Osiris package of Galaxy (an open-source workflow management system). Next, to rapidly annotate transcriptomes from organisms that lack sequenced genomes, we repurposed a maximum likelihood-based Evolutionary Placement Algorithm (implemented in RAxML) to place sequences of potential LIT genes on to our pre-calculated gene trees. Finally, we implemented PIA in Galaxy and used it to search for LIT genes in 28 newly-sequenced transcriptomes from the light-interacting tissues of a range of cephalopod mollusks, arthropods, and cubozoan cnidarians. Our new trees for LIT genes are available on the Bitbucket public repository ( http://bitbucket.org/osiris_phylogenetics/pia/ ) and we demonstrate PIA on a publicly-accessible web server ( http://galaxy-dev.cnsi.ucsb.edu/pia/ ). CONCLUSIONS: Our new trees for LIT genes will be a valuable resource for researchers studying the evolution of eyes or other light-interacting structures. We also introduce PIA, a high throughput method for using phylogenetic relationships to identify LIT genes in transcriptomes from non-model organisms. With simple modifications, our methods may be used to search for different sets of genes or to annotate data sets from taxa outside of Metazoa.
Daniel I. Speiser, M. Sabrina Pankey, Alexander K. Zaharoff, Barbara A. Battelle, Heather D. Bracken-Grissom, Jesse W. Breinholt, Seth M. Bybee, Thomas W. Cronin, Anders Garm, Annie R. Lindgren, Nipam H. Patel, Megan L. Porter, Meredith E. Protas, Ajna S. Rivera, Jeanne M. Serb, Kirk S. Zigler, Keith A. Crandall, Todd H. Oakley
BMC Bioinform.17
2012 Phylogenetic search through partial tree mixing
abstract
BACKGROUND: Recent advances in sequencing technology have created large data sets upon which phylogenetic inference can be performed. Current research is limited by the prohibitive time necessary to perform tree search on a reasonable number of individuals. This research develops new phylogenetic algorithms that can operate on tens of thousands of species in a reasonable amount of time through several innovative search techniques. RESULTS: When compared to popular phylogenetic search algorithms, better trees are found much more quickly for large data sets. These algorithms are incorporated in the PSODA application available at http://dna.cs.byu.edu/psoda CONCLUSIONS: The use of Partial Tree Mixing in a partition based tree space allows the algorithm to quickly converge on near optimal tree regions. These regions can then be searched in a methodical way to determine the overall optimal phylogenetic solution.
Kenneth Sundberg, Mark J. Clement, Quinn Snell, Dan Ventura, Michael Whiting, Keith A. Crandall
BMC Bioinform.6
2003 TreeSAAP: Selection on Amino Acid Properties using phylogenetic trees
abstract
Abstract Summary: The software program TreeSAAP measures the selective influences on 31 structural and biochemical amino acid properties during cladogenesis, and performs goodness-of-fit and categorical statistical tests. Availability: The TreeSAAP package (executables for Windows PC or Macintosh OSX, Java source code, documentation, and instruction manual) is available at http://genome.cs.byu.edu/treesaap.htm. UNIX version is available upon request. Contact: [email protected] * To whom correspondence should be addressed.
Steve Woolley, Justin Johnson 0002, Keith A. Crandall, David A. McClellan
Bioinform.4
1998 MODELTEST: testing the model of DNA substitution
abstract
SUMMARY: The program MODELTEST uses log likelihood scores to establish the model of DNA evolution that best fits the data. AVAILABILITY: The MODELTEST package, including the source code and some documentation is available at http://bioag.byu. edu/zoology/crandall_lab/modeltest.html.
David Posada, Keith A. Crandall
Bioinform.2