EDBT 2026 Demo / reviewers in the wild / expert
Derrick E. Fouts
dblp:154/3959
· DBLP profile ↗
6ranked-venue papers
0as first author
0since 2021 · last 2019
0000-0003-4323-7668ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
comparative genomics |
0.7 | 2 | 2019 | Large-scale comparative analysis of microbial pan-genomes using PanOCT · Bioinform. 2019 GGRaSP: a R-package for selecting representative genomes using Gaussian mixture models · Bioinform. 2018 |
Bioinformatics and computational biology › comparative genomics › pangenomics
pan-genome analysis |
0.4 | 1 | 2019 | Large-scale comparative analysis of microbial pan-genomes using PanOCT · Bioinform. 2019 |
Bioinformatics and computational biology › statistical genetics › genomic prediction
genomic selection |
0.3 | 1 | 2018 | GGRaSP: a R-package for selecting representative genomes using Gaussian mixture models · Bioinform. 2018 |
Bioinformatics and computational biology › genomics
microbial genomics |
0.3 | 1 | 2017 | LOCUST: a custom sequence locus typer for classifying microbial isolates · Bioinform. 2017 |
Bioinformatics and computational biology › metagenomics
antimicrobial resistance gene detection |
0.1 | 1 | 2019 | Large-scale comparative analysis of microbial pan-genomes using PanOCT · Bioinform. 2019 |
Bioinformatics and computational biology › genomics › microbial genomics
bacterial genome analysis |
0.1 | 1 | 2018 | GGRaSP: a R-package for selecting representative genomes using Gaussian mixture models · Bioinform. 2018 |
Bioinformatics and computational biology
genomics |
0.1 | 1 | 2018 | GGRaSP: a R-package for selecting representative genomes using Gaussian mixture models · Bioinform. 2018 |
Methods — techniques the papers use, named apart from their topics
sequence alignment · 0.4hierarchical clustering · 0.4unsupervised clustering · 0.3gaussian mixture model · 0.3locus typing · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Large-scale comparative analysis of microbial pan-genomes using PanOCTabstractSUMMARY: The JCVI pan-genome pipeline is a collection of programs to run PanOCT and tools that support and extend the capabilities of PanOCT. PanOCT (pan-genome ortholog clustering tool) is a tool for pan-genome analysis of closely related prokaryotic species or strains. The JCVI Pan-Genome Pipeline wrapper invokes command-line utilities that prepare input genomes, invoke third-party tools such as NCBI Blast+, run PanOCT, generate a consensus pan-genome, annotate features of the pan-genome, detect sets of genes of interest such as antimicrobial resistance (AMR) genes and generate figures, tables and html pages to visualize the results. The pipeline can run in a hierarchical mode, lowering the RAM and compute resources used. AVAILABILITY AND IMPLEMENTATION: Source code, demo data, and detailed documentation are freely available at https://github.com/JCVenterInstitute/PanGenomePipeline. Jason M. Inman, Granger G. Sutton, Erin Beck, Lauren M. Brinkac, Thomas H. Clarke, Derrick E. Fouts |
Bioinform. | 6 |
| 2019 | OMeta: an ontology-based, data-driven metadata tracking systemabstractBACKGROUND: The development of high-throughput sequencing and analysis has accelerated multi-omics studies of thousands of microbial species, metagenomes, and infectious disease pathogens. Omics studies are enabling genotype-phenotype association studies which identify genetic determinants of pathogen virulence and drug resistance, as well as phylogenetic studies designed to track the origin and spread of disease outbreaks. These omics studies are complex and often employ multiple assay technologies including genomics, metagenomics, transcriptomics, proteomics, and metabolomics. To maximize the impact of omics studies, it is essential that data be accompanied by detailed contextual metadata (e.g., specimen, spatial-temporal, phenotypic characteristics) in clear, organized, and consistent formats. Over the years, many metadata standards developed by various metadata standards initiatives have arisen; the Genomic Standards Consortium's minimal information standards (MIxS), the GSCID/BRC Project and Sample Application Standard. Some tools exist for tracking metadata, but they do not provide event based capabilities to configure, collect, validate, and distribute metadata. To address this gap in the scientific community, an event based data-driven application, OMeta, was created that allows users to quickly configure, collect, validate, distribute, and integrate metadata. RESULTS: A data-driven web application, OMeta, has been developed for use by researchers consisting of a browser-based interface, a command-line interface (CLI), and server-side components that provide an intuitive platform for configuring, capturing, viewing, and sharing metadata. Project and sample metadata can be set based on existing standards or based on projects goals. Recorded information includes details on the biological samples, procedures, protocols, and experimental technologies, etc. This information can be organized based on events, including sample collection, sample quantification, sequencing assay, and analysis results. OMeta enables configuration in various presentation types: checkbox, file, drop-box, ontology, and fields can be configured to use the National Center for Biomedical Ontology (NCBO), a biomedical ontology server. Furthermore, OMeta maintains a complete audit trail of all changes made by users and allows metadata export in comma separated value (CSV) format for convenient deposition of data into public databases. CONCLUSIONS: We present, OMeta, a web-based software application that is built on data-driven principles for configuring and customizing data standards, capturing, curating, and sharing metadata. Indresh Singh, Mehmet Kuscuoglu, Derek M. Harkins, Granger G. Sutton, Derrick E. Fouts, Karen E. Nelson |
BMC Bioinform. | 5 |
| 2018 | GGRaSP: a R-package for selecting representative genomes using Gaussian mixture modelsabstractMotivation: The vast number of available sequenced bacterial genomes occasionally exceeds the facilities of comparative genomic methods or is dominated by a single outbreak strain, and thus a diverse and representative subset is required. Generation of the reduced subset currently requires a priori supervised clustering and sequence-only selection of medoid genomic sequences, independent of any additional genome metrics or strain attributes. Results: The Gaussian Genome Representative Selector with Prioritization (GGRaSP) R-package described below generates a reduced subset of genomes that prioritizes maintaining genomes of interest to the user as well as minimizing the loss of genetic variation. The package also allows for unsupervised clustering by modeling the genomic relationships using a Gaussian mixture model to select an appropriate cluster threshold. We demonstrate the capabilities of GGRaSP by generating a reduced list of 315 genomes from a genomic dataset of 4600 Escherichia coli genomes, prioritizing selection by type strain and by genome completeness. Availability and implementaion: GGRaSP is available at https://github.com/JCVenterInstitute/ggrasp/. Supplementary information: Supplementary data are available at Bioinformatics online. Thomas H. Clarke, Lauren M. Brinkac, Granger G. Sutton, Derrick E. Fouts |
Bioinform. | 4 |
| 2018 | PanACEA: a bioinformatics tool for the exploration and visualization of bacterial pan-chromosomesabstractBACKGROUND: Bacterial pan-genomes, comprised of conserved and variable genes across multiple sequenced bacterial genomes, allow for identification of genomic regions that are phylogenetically discriminating or functionally important. Pan-genomes consist of large amounts of data, which can restrict researchers ability to locate and analyze these regions. Multiple software packages are available to visualize pan-genomes, but currently their ability to address these concerns are limited by using only pre-computed data sets, prioritizing core over variable gene clusters, or by not accounting for pan-chromosome positioning in the viewer. RESULTS: We introduce PanACEA (Pan-genome Atlas with Chromosome Explorer and Analyzer), which utilizes locally-computed interactive web-pages to view ordered pan-genome data. It consists of multi-tiered, hierarchical display pages that extend from pan-chromosomes to both core and variable regions to single genes. Regions and genes are functionally annotated to allow for rapid searching and visual identification of regions of interest with the option that user-supplied genomic phylogenies and metadata can be incorporated. PanACEA's memory and time requirements are within the capacities of standard laptops. The capability of PanACEA as a research tool is demonstrated by highlighting a variable region important in differentiating strains of Enterobacter hormaechei. CONCLUSIONS: PanACEA can rapidly translate the results of pan-chromosome programs into an intuitive and interactive visual representation. It will empower researchers to visually explore and identify regions of the pan-chromosome that are most biologically interesting, and to obtain publication quality images of these regions. Thomas H. Clarke, Lauren M. Brinkac, Jason M. Inman, Granger G. Sutton, Derrick E. Fouts |
BMC Bioinform. | 5 |
| 2017 | LOCUST: a custom sequence locus typer for classifying microbial isolatesabstractSUMMARY: LOCUST is a custom sequence locus typer tool for classifying microbial genomes. It provides a fully automated opportunity to customize the classification of genome-wide nucleotide variant data most relevant to biological research. AVAILABILITY AND IMPLEMENTATION: Source code, demo data, and detailed documentation are freely available at http://sourceforge.net/projects/locustyper . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lauren M. Brinkac, Erin Beck, Jason M. Inman, Pratap Venepally, Derrick E. Fouts, Granger G. Sutton |
Bioinform. | 5 |
| 2014 | NeatFreq: reference-free data reduction and coverage normalization for De Novo sequence assemblyabstractBACKGROUND: Deep shotgun sequencing on next generation sequencing (NGS) platforms has contributed significant amounts of data to enrich our understanding of genomes, transcriptomes, amplified single-cell genomes, and metagenomes. However, deep coverage variations in short-read data sets and high sequencing error rates of modern sequencers present new computational challenges in data interpretation, including mapping and de novo assembly. New lab techniques such as multiple displacement amplification (MDA) of single cells and sequence independent single primer amplification (SISPA) allow for sequencing of organisms that cannot be cultured, but generate highly variable coverage due to amplification biases. RESULTS: Here we introduce NeatFreq, a software tool that reduces a data set to more uniform coverage by clustering and selecting from reads binned by their median kmer frequency (RMKF) and uniqueness. Previous algorithms normalize read coverage based on RMKF, but do not include methods for the preferred selection of (1) extremely low coverage regions produced by extremely variable sequencing of random-primed products and (2) 2-sided paired-end sequences. The algorithm increases the incorporation of the most unique, lowest coverage, segments of a genome using an error-corrected data set. NeatFreq was applied to bacterial, viral plaque, and single-cell sequencing data. The algorithm showed an increase in the rate at which the most unique reads in a genome were included in the assembled consensus while also reducing the count of duplicative and erroneous contigs (strings of high confidence overlaps) in the deliverable consensus. The results obtained from conventional Overlap-Layout-Consensus (OLC) were compared to simulated multi-de Bruijn graph assembly alternatives trained for variable coverage input using sequence before and after normalization of coverage. Coverage reduction was shown to increase processing speed and reduce memory requirements when using conventional bacterial assembly algorithms. CONCLUSIONS: The normalization of deep coverage spikes, which would otherwise inhibit consensus resolution, enables High Throughput Sequencing (HTS) assembly projects to consistently run to completion with existing assembly software. The NeatFreq software package is free, open source and available at https://github.com/bioh4x/NeatFreq . Jamison M. McCorrison, Pratap Venepally, Indresh Singh, Derrick E. Fouts, Roger S. Lasken, Barbara A. Methé |
BMC Bioinform. | 4 |