Abigail L. Manson

dblp:344/6943 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-3800-0714ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021
YearPublicationVenuePosition
2026 Response to 'On using clustering statistics for assessing plasmid binning tools accuracy'
abstract
This response addresses the comments raised by Dr. Epain and colleagues in their Letter to the Editor titled 'On using clustering statistics for assessing plasmid binning tools accuracy' in response to our paper 'Circling in on plasmids: benchmarking plasmid detection and reconstruction tools for short-read data from diverse species'. In their letter, the authors caution against using homogeneity and completeness measures to evaluate the accuracy of plasmid binning tools due to issues related to the length and content of contigs. They also refer to PlasEval, a plasmid binning evaluation tool that they recently developed. In response, we evaluated the impact of contig size and content on the results of our study, and also repeated the benchmarking of plasmid reconstruction tools using PlasEval. Though the absolute values of metrics differed across tests, we observed nearly identical rankings among plasmid reconstruction tools as originally reported in our study, with gplas2 outperforming all other tools. Nevertheless, approaches specifically designed for plasmid data may be better suited than clustering metrics to evaluate plasmid reconstruction.
Colin J. Worby, Thomas Abeel, Ashlee M. Earl, Abigail L. Manson
Briefings Bioinform.5
2025 Circling in on plasmids: benchmarking plasmid detection and reconstruction tools for short-read data from diverse species
abstract
The ability to detect and reconstruct plasmids from genome assemblies is crucial for studying the evolution and spread of antimicrobial resistance and virulence in bacteria. Though long-read sequencing technologies have made reconstructing plasmids easier, most (97%) of the bacterial genome assemblies in the public domain are generated from short-read data. Work to compare plasmid reconstruction tools has focused primarily on Escherichia coli, leaving gaps in our understanding of how well these tools perform on other, less well-characterized, taxa. Using high-quality assemblies as ground truth, we benchmarked 12 plasmid detection tools (which identify plasmid contigs in assemblies) and four plasmid reconstruction tools (which group contigs from the same plasmid together). We tested their ability to characterize diverse plasmids from short-read assemblies representing a wide range of Enterobacterales and Enterococcus species, including newly discovered and poorly characterized species collected from nonhuman hosts. Plasmer, PlasmidEC, PlaScope, and gplas2 were the highest-scoring plasmid detection tools, performing well for both Enterobacterales and enterococci. The two major determinants of accurate plasmid detection were representation in plasmid databases-with Enterobacterales plasmids being more easily detected than those from enterococci-and assembly contiguity, which was also key for successful plasmid reconstruction. Gplas2 performed best for plasmid reconstruction; however, less than half of plasmids were perfectly reconstructed, suggesting that substantial room for improvement remains in this class of tools.
Célia Souque, Colin J. Worby, Terrance Shea, Nicoletta Commins, Joshua T. Smith, Arjun M. Miklos, Thomas Abeel, Ashlee M. Earl, Abigail L. Manson
Briefings Bioinform.10
2025 Fast and exact gap-affine partial order alignment with POASTA
abstract
MOTIVATION: Partial order alignment is a widely used method for computing multiple sequence alignments, with applications in genome assembly and pangenomics, among many others. Current algorithms to compute the optimal, gap-affine partial order alignment do not scale well to larger graphs and sequences. While heuristic approaches exist, they do not guarantee optimal alignment and sacrifice alignment accuracy. RESULTS: We present POASTA, a new optimal algorithm for partial order alignment that exploits long stretches of matching sequence between the graph and a query. We benchmarked POASTA against the state-of-the-art on several diverse bacterial gene datasets and demonstrated an average speed-up of 4.1× and up to 9.8×, using less memory. POASTA's memory scaling characteristics enabled the construction of much larger POA graphs than previously possible, as demonstrated by megabase-length alignments of 342 Mycobacterium tuberculosis sequences. AVAILABILITY AND IMPLEMENTATION: POASTA is available on Github at https://github.com/broadinstitute/poasta.
Lucas R. van Dijk, Abigail L. Manson, Ashlee M. Earl, Kiran V. Garimella, Thomas Abeel
Bioinform.2
2024 SAFPred: synteny-aware gene function prediction for bacteria using protein embeddings
abstract
MOTIVATION: Today, we know the function of only a small fraction of the protein sequences predicted from genomic data. This problem is even more salient for bacteria, which represent some of the most phylogenetically and metabolically diverse taxa on Earth. This low rate of bacterial gene annotation is compounded by the fact that most function prediction algorithms have focused on eukaryotes, and conventional annotation approaches rely on the presence of similar sequences in existing databases. However, often there are no such sequences for novel bacterial proteins. Thus, we need improved gene function prediction methods tailored for bacteria. Recently, transformer-based language models-adopted from the natural language processing field-have been used to obtain new representations of proteins, to replace amino acid sequences. These representations, referred to as protein embeddings, have shown promise for improving annotation of eukaryotes, but there have been only limited applications on bacterial genomes. RESULTS: To predict gene functions in bacteria, we developed SAFPred, a novel synteny-aware gene function prediction tool based on protein embeddings from state-of-the-art protein language models. SAFpred also leverages the unique operon structure of bacteria through conserved synteny. SAFPred outperformed both conventional sequence-based annotation methods and state-of-the-art methods on multiple bacterial species, including for distant homolog detection, where the sequence similarity to the proteins in the training set was as low as 40%. Using SAFPred to identify gene functions across diverse enterococci, of which some species are major clinical threats, we identified 11 previously unrecognized putative novel toxins, with potential significance to human and animal health. AVAILABILITY AND IMPLEMENTATION: https://github.com/AbeelLab/safpred.
Aysun Urhan, Bianca-Maria Cosma, Ashlee M. Earl, Abigail L. Manson, Thomas Abeel
Bioinform.4
2023 Comparison of long- and short-read metagenomic assembly for low-abundance species and resistance genes
abstract
Recent technological and computational advances have made metagenomic assembly a viable approach to achieving high-resolution views of complex microbial communities. In previous benchmarking, short-read (SR) metagenomic assemblers had the highest accuracy, long-read (LR) assemblers generated the most contiguous sequences and hybrid (HY) assemblers balanced length and accuracy. However, no assessments have specifically compared the performance of these assemblers on low-abundance species, which include clinically relevant organisms in the gut. We generated semi-synthetic LR and SR datasets by spiking small and increasing amounts of Escherichia coli isolate reads into fecal metagenomes and, using different assemblers, examined E. coli contigs and the presence of antibiotic resistance genes (ARGs). For ARG assembly, although SR assemblers recovered more ARGs with high accuracy, even at low coverages, LR assemblies allowed for the placement of ARGs within longer, E. coli-specific contigs, thus pinpointing their taxonomic origin. HY assemblies identified resistance genes with high accuracy and had lower contiguity than LR assemblies. Each assembler type's strengths were maintained even when our isolate was spiked in with a competing strain, which fragmented and reduced the accuracy of all assemblies. For strain characterization and determining gene context, LR assembly is optimal, while for base-accurate gene identification, SR assemblers outperform other options. HY assembly offers contiguity and base accuracy, but requires generating data on multiple platforms, and may suffer high misassembly rates when strain diversity exists. Our results highlight the trade-offs associated with each approach for recovering low-abundance taxa, and that the optimal approach is goal-dependent.
Sosie Yorki, Terrance Shea, Christina Cuomo, Bruce J. Walker, Regina C. Larocque, Abigail L. Manson, Ashlee M. Earl, Colin J. Worby
Briefings Bioinform.6