Michael B. Hall

dblp:77/4501 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2025
0000-0003-3683-6208ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › genomics › genome analysis
genome size estimation
0.912025
Genome size estimation from long read overlaps · Bioinform. 2025
Bioinformatics and computational biology › sequence analysis
genomic sequence analysis
0.912025
Genome size estimation from long read overlaps · Bioinform. 2025

Methods — techniques the papers use, named apart from their topics

read-to-read overlap analysis · 0.9median estimation · 0.9
YearPublicationVenuePosition
2025 Genome size estimation from long read overlaps
abstract
MOTIVATION: Accurate genome size estimation is an important component of genomic analyses such as assembly and coverage calculation, though existing tools are primarily optimized for short-read data. RESULTS: We present LRGE, a novel tool that uses read-to-read overlap information to estimate genome size in a reference-free manner. LRGE calculates per-read genome size estimates by analysing the expected number of overlaps for each read, considering read lengths and a minimum overlap threshold. The final size is taken as the median of these estimates, ensuring robustness to outliers such as reads with no overlaps. Additionally, LRGE provides an expected confidence range for the estimate. We validate LRGE on a large, diverse bacterial dataset and confirm it generalizes to eukaryotic datasets. On bacterial genomes, LRGE outperforms k-mer-based methods in both accuracy and computational efficiency and produces genome size estimates comparable to those from assembly-based approaches, like Raven, while using significantly less computational resources. AVAILABILITY AND IMPLEMENTATION: Our method, LRGE (Long Read-based Genome size Estimation from overlaps), is implemented in Rust and is available as a precompiled binary for most architectures, a Bioconda package, a prebuilt container image, and a crates.io package as a binary (lrge) or library (liblrge). The source code is available at https://github.com/mbhall88/lrge and an archive at https://doi.org/10.5281/zenodo.17183812 under an MIT license.
Michael B. Hall, Lachlan James M. Coin
Bioinform.1