Robert S. Harris

dblp:48/3842 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
1since 2021 · last 2022
0000-0001-5464-6892ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
6 papers
Bioinformatics and computational biology · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
genomics
0.412019
Noise-cancelling repeat finder: uncovering tandem repeats in error-prone long-read sequencing data · Bioinform. 2019
Bioinformatics and computational biology › genomics › next-generation sequencing data analysis
long-read sequencing analysis
0.412019
Noise-cancelling repeat finder: uncovering tandem repeats in error-prone long-read sequencing data · Bioinform. 2019
Bioinformatics and computational biology › sequence analysis › repeat detection
tandem repeat detection
0.412019
Noise-cancelling repeat finder: uncovering tandem repeats in error-prone long-read sequencing data · Bioinform. 2019
Bioinformatics and computational biology › sequence analysis › sequence assembly
genome assembly
0.312018
RecoverY: k-mer-based read classification for Y-chromosome-specific sequencing and assembly · Bioinform. 2018
Bioinformatics and computational biology
sequence analysis
0.312017
AllSome Sequence Bloom Trees · RECOMB 2017
Bioinformatics and computational biology › sequence analysis
sequence indexing
0.312017
AllSome Sequence Bloom Trees · RECOMB 2017
Bioinformatics and computational biology › sequence analysis
read mapping
0.212022
The minimizer Jaccard estimator is biased and inconsistent · Bioinform. 2022
Bioinformatics and computational biology
comparative genomics
0.112007
HomologMiner: looking for homologous genomic groups in whole genomes · Bioinform. 2007
Bioinformatics and computational biology › genomics
genome analysis
0.112007
HomologMiner: looking for homologous genomic groups in whole genomes · Bioinform. 2007
Bioinformatics and computational biology › sequence analysis
repeat detection
0.112007
HomologMiner: looking for homologous genomic groups in whole genomes · Bioinform. 2007
Bioinformatics and computational biology › sequence analysis
sequence similarity search
0.112007
HomologMiner: looking for homologous genomic groups in whole genomes · Bioinform. 2007

Methods — techniques the papers use, named apart from their topics

minimizer sketch · 0.6partitioning · 0.4bloom filter · 0.4noise-cancelling algorithm · 0.4k-mer abundance analysis · 0.3automatic parameter selection · 0.3tandem repeat detection · 0.1sequence clustering · 0.1
YearPublicationVenuePosition
2022 The minimizer Jaccard estimator is biased and inconsistent
abstract
MOTIVATION: Sketching is now widely used in bioinformatics to reduce data size and increase data processing speed. Sketching approaches entice with improved scalability but also carry the danger of decreased accuracy and added bias. In this article, we investigate the minimizer sketch and its use to estimate the Jaccard similarity between two sequences. RESULTS: We show that the minimizer Jaccard estimator is biased and inconsistent, which means that the expected difference (i.e. the bias) between the estimator and the true value is not zero, even in the limit as the lengths of the sequences grow. We derive an analytical formula for the bias as a function of how the shared k-mers are laid out along the sequences. We show both theoretically and empirically that there are families of sequences where the bias can be substantial (e.g. the true Jaccard can be more than double the estimate). Finally, we demonstrate that this bias affects the accuracy of the widely used mashmap read mapping tool. AVAILABILITY AND IMPLEMENTATION: Scripts to reproduce our experiments are available at https://github.com/medvedevgroup/minimizer-jaccard-estimator/tree/main/reproduce. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mahdi Belbasi, Antonio Blanca, Robert S. Harris, David Koslicki, Paul Medvedev
Bioinform.3
2020 Improved representation of sequence bloom trees
abstract
MOTIVATION: Algorithmic solutions to index and search biological databases are a fundamental part of bioinformatics, providing underlying components to many end-user tools. Inexpensive next generation sequencing has filled publicly available databases such as the Sequence Read Archive beyond the capacity of traditional indexing methods. Recently, the Sequence Bloom Tree (SBT) and its derivatives were proposed as a way to efficiently index such data for queries about transcript presence. RESULTS: We build on the SBT framework to construct the HowDe-SBT data structure, which uses a novel partitioning of information to reduce the construction and query time as well as the size of the index. Compared to previous SBT methods, on real RNA-seq data, HowDe-SBT can construct the index in less than 36% of the time and with 39% less space and can answer small-batch queries at least five times faster. We also develop a theoretical framework in which we can analyze and bound the space and query performance of HowDe-SBT compared to other SBT methods. AVAILABILITY AND IMPLEMENTATION: HowDe-SBT is available as a free open source program on https://github.com/medvedevgroup/HowDeSBT. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Robert S. Harris, Paul Medvedev
Bioinform.1
2019 Noise-cancelling repeat finder: uncovering tandem repeats in error-prone long-read sequencing data
abstract
SUMMARY: Tandem DNA repeats can be sequenced with long-read technologies, but cannot be accurately deciphered due to the lack of computational tools taking high error rates of these technologies into account. Here we introduce Noise-Cancelling Repeat Finder (NCRF) to uncover putative tandem repeats of specified motifs in noisy long reads produced by Pacific Biosciences and Oxford Nanopore sequencers. Using simulations, we validated the use of NCRF to locate tandem repeats with motifs of various lengths and demonstrated its superior performance as compared to two alternative tools. Using real human whole-genome sequencing data, NCRF identified long arrays of the (AATGG)n repeat involved in heat shock stress response. AVAILABILITY AND IMPLEMENTATION: NCRF is implemented in C, supported by several python scripts, and is available in bioconda and at https://github.com/makovalab-psu/NoiseCancellingRepeatFinder. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Robert S. Harris, Monika Cechova, Kateryna D. Makova
Bioinform.1
2018 RecoverY: k-mer-based read classification for Y-chromosome-specific sequencing and assembly
abstract
Motivation: The haploid mammalian Y chromosome is usually under-represented in genome assemblies due to high repeat content and low depth due to its haploid nature. One strategy to ameliorate the low coverage of Y sequences is to experimentally enrich Y-specific material before assembly. As the enrichment process is imperfect, algorithms are needed to identify putative Y-specific reads prior to downstream assembly. A strategy that uses k-mer abundances to identify such reads was used to assemble the gorilla Y. However, the strategy required the manual setting of key parameters, a time-consuming process leading to sub-optimal assemblies. Results: We develop a method, RecoverY, that selects Y-specific reads by automatically choosing the abundance level at which a k-mer is deemed to originate from the Y. This algorithm uses prior knowledge about the Y chromosome of a related species or known Y transcript sequences. We evaluate RecoverY on both simulated and real data, for human and gorilla, and investigate its robustness to important parameters. We show that RecoverY leads to a vastly superior assembly compared to alternate strategies of filtering the reads or contigs. Compared to the preliminary strategy used by Tomaszkiewicz et al., we achieve a 33% improvement in assembly size and a 20% improvement in the NG50, demonstrating the power of automatic parameter selection. Availability and implementation: Our tool RecoverY is freely available at https://github.com/makovalab-psu/RecoverY. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Samarth Rangavittal, Robert S. Harris, Monika Cechova, Marta Tomaszkiewicz, Rayan Chikhi, Kateryna D. Makova, Paul Medvedev
Bioinform.2
2017 AllSome Sequence Bloom Trees
Chen Sun 0014, Robert S. Harris, Rayan Chikhi, Paul Medvedev
RECOMB2
2014 HECTOR: a parallel multistage homopolymer spectrum based error corrector for 454 sequencing data
abstract
BACKGROUND: Current-generation sequencing technologies are able to produce low-cost, high-throughput reads. However, the produced reads are imperfect and may contain various sequencing errors. Although many error correction methods have been developed in recent years, none explicitly targets homopolymer-length errors in the 454 sequencing reads. RESULTS: We present HECTOR, a parallel multistage homopolymer spectrum based error corrector for 454 sequencing data. In this algorithm, for the first time we have investigated a novel homopolymer spectrum based approach to handle homopolymer insertions or deletions, which are the dominant sequencing errors in 454 pyrosequencing reads. We have evaluated the performance of HECTOR, in terms of correction quality, runtime and parallel scalability, using both simulated and real pyrosequencing datasets. This performance has been further compared to that of Coral, a state-of-the-art error corrector which is based on multiple sequence alignment and Acacia, a recently published error corrector for amplicon pyrosequences. Our evaluations reveal that HECTOR demonstrates comparable correction quality to Coral, but runs 3.7× faster on average. In addition, HECTOR performs well even when the coverage of the dataset is low. CONCLUSION: Our homopolymer spectrum based approach is theoretically capable of processing arbitrary-length homopolymer-length errors, with a linear time complexity. HECTOR employs a multi-threaded design based on a master-slave computing model. Our experimental results show that HECTOR is a practical 454 pyrosequencing read error corrector which is competitive in terms of both correction quality and speed. The source code and all simulated data are available at: http://hector454.sourceforge.net.
Adrianto Wirawan, Robert S. Harris, Yongchao Liu 0004, Bertil Schmidt, Jan Schröder
BMC Bioinform.2
2007 HomologMiner: looking for homologous genomic groups in whole genomes
abstract
MOTIVATION: Complex genomes contain numerous repeated sequences, and genomic duplication is believed to be a main evolutionary mechanism to obtain new functions. Several tools are available for de novo repeat sequence identification, and many approaches exist for clustering homologous protein sequences. We present an efficient new approach to identify and cluster homologous DNA sequences with high accuracy at the level of whole genomes, excluding low-complexity repeats, tandem repeats and annotated interspersed repeats. We also determine the boundaries of each group member so that it closely represents a biological unit, e.g. a complete gene, or a partial gene coding a protein domain. RESULTS: We developed a program called HomologMiner to identify homologous groups applicable to genome sequences that have been properly marked for low-complexity repeats and annotated interspersed repeats. We applied it to the whole genomes of human (hg17), macaque (rheMac2) and mouse (mm8). Groups obtained include gene families (e.g. olfactory receptor gene family, zinc finger families), unannotated interspersed repeats and additional homologous groups that resulted from recent segmental duplications. Our program incorporates several new methods: a new abstract definition of consistent duplicate units, a new criterion to remove moderately frequent tandem repeats, and new algorithmic techniques. We also provide preliminary analysis of the output on the three genomes mentioned above, and show several applications including identifying boundaries of tandem gene clusters and novel interspersed repeat families. AVAILABILITY: All programs and datasets are downloadable from www.bx.psu.edu/miller_lab.
Minmei Hou, Piotr Berman, Chih-Hao Hsu, Robert S. Harris
Bioinform.4
1995 String matching with left-to-right networks
Jens Gregor, Robert S. Harris
Pattern Recognit. Lett.2