EDBT 2026 Demo / reviewers in the wild / expert
C. Titus Brown
dblp:11/4436
· DBLP profile ↗
10ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0001-6001-2677ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A comparative review of short genetic variant databases across humans and animal speciesabstractUnderstanding genetic variation is necessary to unravel the complexities of evolution and diverse traits across species. Short genetic variants (<50 bp in length) represent key genomic variations that play a crucial role in shaping the genetic landscape of populations. Human short genetic variant databases are often considered the gold standard for variant repositories, and this review compares them with variant databases created for various animal species. This review examines the methodologies, data integration, and various applications that differ or align between human and animal databases. The goal is to identify challenges in leveraging genetic variation across species, recommend strategies for overcoming these challenges, and suggest future directions for research and database development. Samantha L. Van Buren, C. Titus Brown, Tomasz Szmatola, Carrie J. Finno |
Briefings Bioinform. | 2 |
| 2024 | LINgroups as a Robust Principled Approach to Compare and Integrate Multiple Bacterial TaxonomiesabstractAs a central organizing principle of biology, bacteria and archaea are classified into a hierarchical structure across taxonomic ranks from kingdom to subspecies. Traditionally, this organization was based on observable characteristics of form and chemistry but recently, bacterial taxonomy has been robustly quantified using comparisons of sequenced genomes, as exemplified in the Genome Taxonomy Database (GTDB). Such genome-based taxonomies resolve genomes down to genera and species and are useful in many contexts yet lack the flexibility and resolution of a fine-grained approach. The Life Identification Number (LIN) approach is a common, quantitative framework to tie existing (and future) bacterial taxonomies together, increase the resolution of genome-based discrimination of taxa, and extend taxonomic identification below the species level in a principled way. Utilizing LINgroup as an organizational concept helps resolve some of the confusion and unforeseen negative effects resulting from nomenclature changes of microorganisms that are closely related by overall genomic similarity (often due to genome-based reclassification). Our experimental results demonstrate the value of LINs and LINgroups in mapping between taxonomies, translating between different nomenclatures, and integrating them into a single taxonomic framework. They also reveal the robustness of LIN assignment to hyper-parameter changes when considering within-species taxonomic groups. Reza Mazloom, N. Tessa Pierce-Ward, Parul Sharma, Leighton Pritchard, C. Titus Brown, Boris A. Vinatzer, Lenwood S. Heath |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2022 | Evaluation of taxonomic classification and profiling methods for long-read shotgun metagenomic sequencing datasetsabstractBACKGROUND: Long-read shotgun metagenomic sequencing is gaining in popularity and offers many advantages over short-read sequencing. The higher information content in long reads is useful for a variety of metagenomics analyses, including taxonomic classification and profiling. The development of long-read specific tools for taxonomic classification is accelerating, yet there is a lack of information regarding their relative performance. Here, we perform a critical benchmarking study using 11 methods, including five methods designed specifically for long reads. We applied these tools to several mock community datasets generated using Pacific Biosciences (PacBio) HiFi or Oxford Nanopore Technology sequencing, and evaluated their performance based on read utilization, detection metrics, and relative abundance estimates. RESULTS: Our results show that long-read classifiers generally performed best. Several short-read classification and profiling methods produced many false positives (particularly at lower abundances), required heavy filtering to achieve acceptable precision (at the cost of reduced recall), and produced inaccurate abundance estimates. By contrast, two long-read methods (BugSeq, MEGAN-LR & DIAMOND) and one generalized method (sourmash) displayed high precision and recall without any filtering required. Furthermore, in the PacBio HiFi datasets these methods detected all species down to the 0.1% abundance level with high precision. Some long-read methods, such as MetaMaps and MMseqs2, required moderate filtering to reduce false positives to resemble the precision and recall of the top-performing methods. We found read quality affected performance for methods relying on protein prediction or exact k-mer matching, and these methods performed better with PacBio HiFi datasets. We also found that long-read datasets with a large proportion of shorter reads (< 2 kb length) resulted in lower precision and worse abundance estimates, relative to length-filtered datasets. Finally, for classification methods, we found that the long-read datasets produced significantly better results than short-read datasets, demonstrating clear advantages for long-read metagenomic sequencing. CONCLUSIONS: Our critical assessment of available methods provides best-practice recommendations for current research using long reads and establishes a baseline for future benchmarking studies. Daniel M. Portik, C. Titus Brown, N. Tessa Pierce-Ward |
BMC Bioinform. | 2 |
| 2022 | Ten simple rules and a template for creating workflows-as-applicationsabstractAs bioinformatics analyses increase in size and complexity, workflow managers are becoming more popular for building pipelines [1][2][3].Workflow managers, such as Snakemake [4], Nextflow [5], and Cromwell [6] with WDL or CWL [7], empower researchers to build robust pipelines that call a series of tools and scripts to perform a bespoke analysis.Workflow managers enable non-bioinformaticians to run published pipelines with confidence, and workflow managers with graphical user interfaces such as Galaxy [8] and BioWorkflow [9] have helped nonbioinformaticians create their own simple pipelines.Earlier tools for workflow management have been around for a while, including GNU Make, ruffus [10], doit [11], rake for ruby [12], and Makeflow [13].However, the integration of cluster and cloud computing support in Snakemake, Nextflow, and Cromwell helped drive their current popularity.The use of workflow managers facilitates following the FAIR (Findable, Accessible, Interoperable, Reusable) guiding principles for open scientific research [14].Interestingly, many existing bioinformatics command line tools are wrappers for a series of other software, but since that is the goal of workflow managers, they can be used instead.Examples of command line tools built on a workflow manager include Hecatomb [15,16], ATLAS [17], VirSorter2 [18], spacegraphcats [19], BlobToolKit [20], and PGAP [21].These tools all consist of two key components: a convenience launcher, which provides the command line interface for the tool and compiles the configuration from user command line arguments, and the workflow pipeline and associated files, which performs the actual analysis.Developing bioinformatics software is much quicker and easier when not reinventing the wheel.For instance, if a tool needs to parse a GenBank file, is it better to code that process manually, or simply load a library that is designed to robustly parse and validate these files?The same concept applies to how bioinformatics software runs.It is possible to write functions to compare file timestamps to add reentrancy to the software, catch error codes, and throw meaningful messages when system calls fail, validate and cleanup intermediate files, allow interaction with a job scheduling system, and perform steps in isolated containers, etc.When workflow managers were still in their infancy, the authors of several popular genome Michael J. Roach 0001, N. Tessa Pierce-Ward, Radoslaw Suchecki, Vijini Mallawaarachchi, Bhavya Nalagampalli Papudeshi, Scott A. Handley, C. Titus Brown, Nathan S. Watson-Haigh, Robert A. Edwards |
PLoS Comput. Biol. | 7 |
| 2021 | MQF and buffered MQF: quotient filters for efficient storage of k-mers with their counts and metadataabstractBACKGROUND: Specialized data structures are required for online algorithms to efficiently handle large sequencing datasets. The counting quotient filter (CQF), a compact hashtable, can efficiently store k-mers with a skewed distribution. RESULT: Here, we present the mixed-counters quotient filter (MQF) as a new variant of the CQF with novel counting and labeling systems. The new counting system adapts to a wider range of data distributions for increased space efficiency and is faster than the CQF for insertions and queries in most of the tested scenarios. A buffered version of the MQF can offload storage to disk, trading speed of insertions and queries for a significant memory reduction. The labeling system provides a flexible framework for assigning labels to member items while maintaining good data locality and a concise memory representation. These labels serve as a minimal perfect hash function but are ~ tenfold faster than BBhash, with no need to re-analyze the original data for further insertions or deletions. CONCLUSIONS: The MQF is a flexible and efficient data structure that extends our ability to work with high throughput sequencing data. Moustafa Shokrof, C. Titus Brown, Tamer A. Mansour |
BMC Bioinform. | 2 |
| 2021 | Ten simple rules to cultivate transdisciplinary collaboration in data scienceabstractAuthor(s): Sahneh, Faryad; Balk, Meghan A; Kisley, Marina; Chan, Chi-kwan; Fox, Mercury; Nord, Brian; Lyons, Eric; Swetnam, Tyson; Huppenkothen, Daniela; Sutherland, Will; Walls, Ramona L; Quinn, Daven P; Tarin, Tonantzin; LeBauer, David; Ribes, David; Birnie, Dunbar P; Lushbough, Carol; Carr, Eric; Nearing, Grey; Fischer, Jeremy; Tyle, Kevin; Carrasco, Luis; Lang, Meagan; Rose, Peter W; Rushforth, Richard R; Roy, Samapriya; Matheson, Thomas; Lee, Tina; Brown, C Titus; Teal, Tracy K; Papeș, Monica; Kobourov, Stephen; Merchant, Nirav | Editor(s): Schwartz, Russell Faryad Sahneh, Meghan A. Balk, Marina Kisley, Chi-Kwan Chan, Mercury Fox, Brian Nord, Eric Lyons 0002, Tyson Lee Swetnam, Daniela Huppenkothen, Will Sutherland, Ramona L. Walls, Daven P. Quinn, Tonantzin Tarin, David S. LeBauer, David Ribes, Dunbar P. Birnie III, Carol Lushbough, Eric Carr, Grey Nearing, Jeremy Fischer, Kevin Tyle, Luis Carrasco, Meagan Lang, Peter W. Rose, Richard R. Rushforth, Samapriya Roy, Thomas Matheson, Tina Lee, C. Titus Brown, Tracy K. Teal, Monica Papes, Stephen G. Kobourov, Nirav C. Merchant |
PLoS Comput. Biol. | 29 |
| 2020 | VA-Store: A Virtual Approximate Store Approach to Supporting Repetitive Big Data in Genome Sequence AnalysesabstractIn recent years, we have witnessed an increasing demand to process big data in numerous applications. It is observed that there often exist substantial amounts of repetitive data in different portions of a big data repository/dataset for applications such as genome sequence analyses. In this paper, we present a novel method, called the VA-Store, to reduce the large space requirement for repetitive data in prevailing genome sequence analysis tasks using k-mers (i.e., subsequences of length k) with multiple k values. The VA-Store maintains a physical store for one portion of the input dataset (i.e., k0-mers) and supports multiple virtual stores for other portions of the dataset (i.e., k-mers with k ≠ k0). Utilizing important relationships among repetitive data, the VA-Store transforms a given query on a virtual store into one or more queries on the physical store for execution. Both precise and approximate transformations are considered. Accuracy estimation models for approximate solutions are derived. Query optimization strategies are suggested to improve query performance. Our experiments using real and synthetic datasets demonstrate that the VA-Store is quite promising in providing effective storage and efficient query processing for solving a kernel database problem on repetitive big data for genome sequence analysis applications. Xianying Liu, Qiang Zhu 0001, Sakti Pramanik, C. Titus Brown, Gang Qian |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2019 | Sequencing data discovery with MetaSeekabstractSUMMARY: Sequencing data resources have increased exponentially in recent years, as has interest in large-scale meta-analyses of integrated next-generation sequencing datasets. However, curation of integrated datasets that match a user's particular research priorities is currently a time-intensive and imprecise task. MetaSeek is a sequencing data discovery tool that enables users to flexibly search and filter on any metadata field to quickly find the sequencing datasets that meet their needs. MetaSeek automatically scrapes metadata from all publicly available datasets in the Sequence Read Archive, cleans and parses messy, user-provided metadata into a structured, standard-compliant database and predicts missing fields where possible. MetaSeek provides a web-based graphical user interface and interactive visualization dashboard, as well as a programmatic API to rapidly search, filter, visualize, save, share and download matching sequencing metadata. AVAILABILITY AND IMPLEMENTATION: The MetaSeek online interface is available at https://www.metaseek.cloud/. The MetaSeek database can also be accessed via API to programmatically search, filter and download all metadata. MetaSeek source code, metadata scrapers and documents are available at https://github.com/MetaSeek-Sequencing-Data-Discovery/metaseek/. Adrienne Hoarfrost, C. Titus Brown, Carol Arnosti |
Bioinform. | 3 |
| 2005 | Paircomp, FamilyRelationsII and Cartwheel: tools for interspecific sequence comparisonabstractBACKGROUND: Comparative sequence analysis is an effective and increasingly common way to identify cis-regulatory regions in animal genomes. RESULTS: We describe three tools for comparative analysis of pairs of BAC-sized genomic regions. Paircomp is a tool that does windowed (ungapped) comparisons of two sequences and reports all matches above a set threshold. FamilyRelationsII is a graphical viewer for comparisons that enables interactive exploration of several different kinds of comparisons. Cartwheel is a Web site and compute-cluster management system used to execute and store comparisons for display by FamilyRelationsII. These tools are specialized for the discovery of cis-regulatory regions in animal genomes. All tools and their source code are freely available at http://family.caltech.edu/. CONCLUSION: These tools have been shown to effectively identify regulatory regions in echinoderms, mammals, and nematodes. C. Titus Brown, Yuan Xie 0003, Eric H. Davidson, R. Andrew Cameron |
BMC Bioinform. | 1 |
| 1999 | Visualizing Evolutionary Activity of GenotypesabstractWe introduce a method for visualizing evolutionary activity of genotypes. Following a proposal of Bedau and Packard [11], we define a genotype's evolutionary activity in terms of the history of its concentration in the evolving population. To visualize this evolutionary activity we graph the distribution of evolutionary activity in the population of genotypes as a function of time. Adaptively significant genotypes trace a salient line or "wave" in these graphs. The quality of these waves indicates a variety of neutral variation, and random genetic drift. We apply this method in an evolutionary model of self-replicating assembly language programs competing for room in a two-dimensional space. Comparison with fitness graphs and with a nonadaptive analogue of this model shows how this method highlights adaptively significant events. Mark A. Bedau, C. Titus Brown |
Artif. Life | 2 |