Bryce Kille

dblp:295/3361 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0003-2946-6915ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
4 papers
Bioinformatics and computational biology · 100%
Databases, data mining, and information retrieval
1 paper
Indexing and storage engines · 100%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › sequence analysis
k-mer sampling
1.522025
A near-tight lower bound on the density of forward sampling schemes · Bioinform. 2025
Minmers are a generalization of minimizers that enable unbiased local Jaccard estimation · Bioinform. 2023
Bioinformatics and computational biology
sequence analysis
1.522025
A near-tight lower bound on the density of forward sampling schemes · Bioinform. 2025
Minmers are a generalization of minimizers that enable unbiased local Jaccard estimation · Bioinform. 2023
Bioinformatics and computational biology
comparative genomics
1.022024
Parsnp 2.0: scalable core-genome alignment for massive microbial datasets · Bioinform. 2024
Minmers are a generalization of minimizers that enable unbiased local Jaccard estimation · Bioinform. 2023
Bioinformatics and computational biology › sequence alignment › genome alignment
multiple genome alignment
0.812024
Parsnp 2.0: scalable core-genome alignment for massive microbial datasets · Bioinform. 2024
Bioinformatics and computational biology › sequence analysis › k-mer sampling
minimizers scheme
0.312025
A near-tight lower bound on the density of forward sampling schemes · Bioinform. 2025
Bioinformatics and computational biology › genomics › microbial genomics
bacterial genome analysis
0.212024
Parsnp 2.0: scalable core-genome alignment for massive microbial datasets · Bioinform. 2024
Bioinformatics and computational biology › genomics
microbial genomics
0.212024
Parsnp 2.0: scalable core-genome alignment for massive microbial datasets · Bioinform. 2024

Methods — techniques the papers use, named apart from their topics

count-min sketch · 1.0bloom filter · 1.0combinatorial lower bound proof · 0.9parallel partitioning · 0.8rolling minhash · 0.7
YearPublicationVenuePosition
2026 A Configurable Gomoku Framework for Learning Heuristics
abstract
Learning to think computationally is a key part of learning computer science. This Gomoku (five-in-a-row) assignment is part of Rice University's summer bridge program, where incoming students from under-resourced secondary schools have a chance to develop computational thinking and coding skills. This project is a fun way for students to practice coding, think about data representation, and explore heuristics. We provide a configurable framework that, given a (student written) heuristic function to evaluate a board, enables students to pit their approach against each other in tournaments of varying difficulty on different size game boards.
Bryce Kille, Stian du Preez, Risa B. Myers
ITiCSE (2)1
2025 A near-tight lower bound on the density of forward sampling schemes
abstract
MOTIVATION: Sampling k-mers is a ubiquitous task in sequence analysis algorithms. Sampling schemes such as the often-used random minimizer scheme are particularly appealing as they guarantee at least one k-mer is selected out of every w consecutive k-mers. Sampling fewer k-mers often leads to an increase in efficiency of downstream methods. Thus, developing schemes that have low density, i.e. have a small proportion of sampled k-mers, is an active area of research. After over a decade of consistent efforts in both decreasing the density of practical schemes and increasing the lower bound on the best possible density, there is still a large gap between the two. RESULTS: We prove a near-tight lower bound on the density of forward sampling schemes, a class of schemes that generalizes minimizer schemes. For small w and k, we observe that our bound is tight when k≡1(mod w). For large w and k, the bound can be approximated by 1w+k⌈w+kw⌉. Importantly, our lower bound implies that existing schemes are much closer to achieving optimal density than previously known. For example, with the current default minimap2 HiFi settings w = 19 and k = 19, we show that the best known scheme for these parameters, the double decycling-set-based minimizer of Pellow et al. is at most 3% denser than optimal, compared to the previous gap of at most 50%. Furthermore, when k≡1(mod w) and the alphabet size σ goes to ∞, we show that mod-minimizers introduced by Groot Koerkamp and Pibiri achieve optimal density matching our lower bound. AVAILABILITY AND IMPLEMENTATION: Minimizer implementations: github.com/RagnarGrootKoerkamp/minimizers ILP and analysis: github.com/treangenlab/sampling-scheme-analysis.
Bryce Kille, Ragnar Groot Koerkamp, Drake McAdams, Alan Liu, Todd J. Treangen
Bioinform.1
2024 Parsnp 2.0: scalable core-genome alignment for massive microbial datasets
abstract
MOTIVATION: Since 2016, the number of microbial species with available reference genomes in NCBI has more than tripled. Multiple genome alignment, the process of identifying nucleotides across multiple genomes which share a common ancestor, is used as the input to numerous downstream comparative analysis methods. Parsnp is one of the few multiple genome alignment methods able to scale to the current era of genomic data; however, there has been no major release since its initial release in 2014. RESULTS: To address this gap, we developed Parsnp v2, which significantly improves on its original release. Parsnp v2 provides users with more control over executions of the program, allowing Parsnp to be better tailored for different use-cases. We introduce a partitioning option to Parsnp, which allows the input to be broken up into multiple parallel alignment processes which are then combined into a final alignment. The partitioning option can reduce memory usage by over 4× and reduce runtime by over 2×, all while maintaining a precise core-genome alignment. The partitioning workflow is also less susceptible to complications caused by assembly artifacts and minor variation, as alignment anchors only need to be conserved within their partition and not across the entire input set. We highlight the performance on datasets involving thousands of bacterial and viral genomes. AVAILABILITY AND IMPLEMENTATION: Parsnp v2 is available at https://github.com/marbl/parsnp.
Bryce Kille, Michael G. Nute, Victor Huang, Eddie Kim, Adam M. Phillippy, Todd J. Treangen
Bioinform.1
2023 Minmers are a generalization of minimizers that enable unbiased local Jaccard estimation
abstract
MOTIVATION: The Jaccard similarity on k-mer sets has shown to be a convenient proxy for sequence identity. By avoiding expensive base-level alignments and comparing reduced sequence representations, tools such as MashMap can scale to massive numbers of pairwise comparisons while still providing useful similarity estimates. However, due to their reliance on minimizer winnowing, previous versions of MashMap were shown to be biased and inconsistent estimators of Jaccard similarity. This directly impacts downstream tools that rely on the accuracy of these estimates. RESULTS: To address this, we propose the minmer winnowing scheme, which generalizes the minimizer scheme by use of a rolling minhash with multiple sampled k-mers per window. We show both theoretically and empirically that minmers yield an unbiased estimator of local Jaccard similarity, and we implement this scheme in an updated version of MashMap. The minmer-based implementation is over 10 times faster than the minimizer-based version under the default ANI threshold, making it well-suited for large-scale comparative genomics applications. AVAILABILITY AND IMPLEMENTATION: MashMap3 is available at https://github.com/marbl/MashMap.
Bryce Kille, Erik Garrison, Todd J. Treangen, Adam M. Phillippy
Bioinform.1
2021 Fast Processing and Querying of 170TB of Genomics Data via a Repeated And Merged BloOm Filter (RAMBO)
abstract
DNA sequencing, especially of microbial genomes and metagenomes, has been at the core of recent research advances in large-scale comparative genomics. The data deluge has resulted in exponential growth in genomic datasets over the past years and has shown no sign of slowing down. Several recent attempts have been made to tame the computational burden of sequence search on these terabyte and petabyte-scale datasets, including raw reads and assembled genomes. However, no known implementation provides both fast query and construction time, keeps the low false-positive requirement, and offers cheap storage of the data structure. We propose a data structure for search called RAMBO (Repeated And Merged BloOm Filter) which is significantly faster in query time than state-of-the-art genome indexing methods- COBS (Compact bit-sliced signature index), Sequence Bloom Trees, HowDeSBT, and SSBT. Furthermore, it supports insertion and query process parallelism, cheap updates for streaming inputs, has a zero false-negative rate, a low false-positive rate, and a small index size. RAMBO converts the search problem into set membership testing among K documents. Interestingly, it is a count-min sketch type arrangement of a membership testing utility (Bloom Filter in our case). The simplicity of the algorithm and embarrassingly parallel architecture allows us to stream and index a 170TB whole-genome sequence dataset in a mere 9 hours on a cluster of 100 nodes while competing methods require weeks.
Minghao Yan, Benjamin Coleman, Bryce Kille, Ryan A. Leo Elworth, Tharun Medini, Todd J. Treangen, Anshumali Shrivastava
SIGMOD Conference4