Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Richard Wilton

dblp:73/5395 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
3since 2021 · last 2023
0000-0003-1263-5532ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 5 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
4 papers
Bioinformatics and computational biology · 100%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 64% GPUs and heterogeneous computing · 36%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
sequence alignment
1.222023
Short-read aligner performance in germline variant identification · Bioinform. 2023
Performance optimization in DNA short-read alignment · Bioinform. 2022
Bioinformatics and computational biology › sequence analysis › read mapping
short read alignment
1.222023
Short-read aligner performance in germline variant identification · Bioinform. 2023
Performance optimization in DNA short-read alignment · Bioinform. 2022
Bioinformatics and computational biology › genomics
variant calling
0.712023
Short-read aligner performance in germline variant identification · Bioinform. 2023
Bioinformatics and computational biology › genomics
genomic data management
0.412019
The Terabase Search Engine: a large-scale relational database of short-read sequences · Bioinform. 2019
Bioinformatics and computational biology › biological database
sequence database
0.412019
The Terabase Search Engine: a large-scale relational database of short-read sequences · Bioinform. 2019
Bioinformatics and computational biology › epigenomics › DNA methylation › DNA methylation analysis
bisulfite sequencing
0.312018
Arioc: GPU-accelerated alignment of short bisulfite-treated reads · Bioinform. 2018
Bioinformatics and computational biology › sequence analysis
read mapping
0.312018
Arioc: GPU-accelerated alignment of short bisulfite-treated reads · Bioinform. 2018
Bioinformatics and computational biology › genomics › variant calling
germline variant calling
0.212023
Short-read aligner performance in germline variant identification · Bioinform. 2023
High-performance computing
performance optimization at scale
0.212022
Performance optimization in DNA short-read alignment · Bioinform. 2022
GPUs and heterogeneous computing
GPU computing
0.112018
Arioc: GPU-accelerated alignment of short bisulfite-treated reads · Bioinform. 2018

Methods — techniques the papers use, named apart from their topics

benchmarking · 1.8performance profiling · 1.1relational indexing · 0.8GPU acceleration · 0.7
YearPublicationVenuePosition
2023 Short-read aligner performance in germline variant identification
abstract
MOTIVATION: Read alignment is an essential first step in the characterization of DNA sequence variation. The accuracy of variant-calling results depends not only on the quality of read alignment and variant-calling software but also on the interaction between these complex software tools. RESULTS: In this review, we evaluate short-read aligner performance with the goal of optimizing germline variant-calling accuracy. We examine the performance of three general-purpose short-read aligners-BWA-MEM, Bowtie 2, and Arioc-in conjunction with three germline variant callers: DeepVariant, FreeBayes, and GATK HaplotypeCaller. We discuss the behavior of the read aligners with regard to the data elements on which the variant callers rely, and illustrate how the runtime configurations of these software tools combine to affect variant-calling performance. AVAILABILITY AND IMPLEMENTATION: The quick brown fox jumps over the lazy dog.
Richard Wilton, Alex Szalay
Bioinform.1
2023 BiocMAP: a Bioconductor-friendly, GPU-accelerated pipeline for bisulfite-sequencing data
abstract
BACKGROUND: Bisulfite sequencing is a powerful tool for profiling genomic methylation, an epigenetic modification critical in the understanding of cancer, psychiatric disorders, and many other conditions. Raw data generated by whole genome bisulfite sequencing (WGBS) requires several computational steps before it is ready for statistical analysis, and particular care is required to process data in a timely and memory-efficient manner. Alignment to a reference genome is one of the most computationally demanding steps in a WGBS workflow, taking several hours or even days with commonly used WGBS-specific alignment software. This naturally motivates the creation of computational workflows that can utilize GPU-based alignment software to greatly speed up the bottleneck step. In addition, WGBS produces raw data that is large and often unwieldy; a lack of memory-efficient representation of data by existing pipelines renders WGBS impractical or impossible to many researchers. RESULTS: We present BiocMAP, a Bioconductor-friendly methylation analysis pipeline consisting of two modules, to address the above concerns. The first module performs computationally-intensive read alignment using Arioc, a GPU-accelerated short-read aligner. Since GPUs are not always available on the same computing environments where traditional CPU-based analyses are convenient, the second module may be run in a GPU-free environment. This module extracts and merges DNA methylation proportions-the fractions of methylated cytosines across all cells in a sample at a given genomic site. Bioconductor-based output objects in R utilize an on-disk data representation to drastically reduce required main memory and make WGBS projects computationally feasible to more researchers. CONCLUSIONS: BiocMAP is implemented using Nextflow and available at http://research.libd.org/BiocMAP/ . To enable reproducible analysis across a variety of typical computing environments, BiocMAP can be containerized with Docker or Singularity, and executed locally or with the SLURM or SGE scheduling engines. By providing Bioconductor objects, BiocMAP's output can be integrated with powerful analytical open source software for analyzing methylation data.
Nicholas J. Eagles, Richard Wilton, Andrew E. Jaffe, Leonardo Collado-Torres
BMC Bioinform.2
2022 Performance optimization in DNA short-read alignment
abstract
SUMMARY: Over the past decade, short-read sequence alignment has become a mature technology. Optimized algorithms, careful software engineering and high-speed hardware have contributed to greatly increased throughput and accuracy. With these improvements, many opportunities for performance optimization have emerged. In this review, we examine three general-purpose short-read alignment tools-BWA-MEM, Bowtie 2 and Arioc-with a focus on performance optimization. We analyze the performance-related behavior of the algorithms and heuristics each tool implements, with the goal of arriving at practical methods of improving processing speed and accuracy. We indicate where an aligner's default behavior may result in suboptimal performance, explore the effects of computational constraints such as end-to-end mapping and alignment scoring threshold, and discuss sources of imprecision in the computation of alignment scores and mapping quality. With this perspective, we describe an approach to tuning short-read aligner performance to meet specific data-analysis and throughput requirements while avoiding potential inaccuracies in subsequent analysis of alignment results. Finally, we illustrate how this approach avoids easily overlooked pitfalls and leads to verifiable improvements in alignment speed and accuracy. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Appendices referenced in this article are available at Bioinformatics online.
Richard Wilton, Alex Szalay
Bioinform.1
2020 Arioc: High-concurrency short-read alignment on multiple GPUs
abstract
In large DNA sequence repositories, archival data storage is often coupled with computers that provide 40 or more CPU threads and multiple GPU (general-purpose graphics processing unit) devices. This presents an opportunity for DNA sequence alignment software to exploit high-concurrency hardware to generate short-read alignments at high speed. Arioc, a GPU-accelerated short-read aligner, can compute WGS (whole-genome sequencing) alignments ten times faster than comparable CPU-only alignment software. When two or more GPUs are available, Arioc's speed increases proportionately because the software executes concurrently on each available GPU device. We have adapted Arioc to recent multi-GPU hardware architectures that support high-bandwidth peer-to-peer memory accesses among multiple GPUs. By modifying Arioc's implementation to exploit this GPU memory architecture we obtained a further 1.8x-2.9x increase in overall alignment speeds. With this additional acceleration, Arioc computes two million short-read alignments per second in a four-GPU system; it can align the reads from a human WGS sequencer run-over 500 million 150nt paired-end reads-in less than 15 minutes. As WGS data accumulates exponentially and high-concurrency computational resources become widespread, Arioc addresses a growing need for timely computation in the short-read data analysis toolchain.
Richard Wilton, Alex Szalay
PLoS Comput. Biol.1
2019 The Terabase Search Engine: a large-scale relational database of short-read sequences
abstract
MOTIVATION: DNA sequencing archives have grown to enormous scales in recent years, and thousands of human genomes have already been sequenced. The size of these data sets has made searching the raw read data infeasible without high-performance data-query technology. Additionally, it is challenging to search a repository of short-read data using relational logic and to apply that logic across samples from multiple whole-genome sequencing samples. RESULTS: We have built a compact, efficiently-indexed database that contains the raw read data for over 250 human genomes, encompassing trillions of bases of DNA, and that allows users to search these data in real-time. The Terabase Search Engine enables retrieval from this database of all the reads for any genomic location in a matter of seconds. Users can search using a range of positions or a specific sequence that is aligned to the genome on the fly. AVAILABILITY AND IMPLEMENTATION: Public access to the Terabase Search Engine database is available at http://tse.idies.jhu.edu. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Richard Wilton, Sarah J. Wheelan, Alex Szalay, Steven Salzberg
Bioinform.1
2018 Arioc: GPU-accelerated alignment of short bisulfite-treated reads
abstract
Motivation: The alignment of bisulfite-treated DNA sequences (BS-seq reads) to a large genome involves a significant computational burden beyond that required to align non-bisulfite-treated reads. In the analysis of BS-seq data, this can present an important performance bottleneck that can be mitigated by appropriate algorithmic and software-engineering improvements. One strategy is to modify the read-alignment algorithms by integrating the logic related to BS-seq alignment, with the goal of making the software implementation amenable to optimizations that lead to higher speed and greater sensitivity than might otherwise be attainable. Results: We evaluated this strategy using Arioc, a short-read aligner that uses GPU (general-purpose graphics processing unit) hardware to accelerate computationally-expensive programming logic. We integrated the BS-seq computational logic into both GPU and CPU code throughout the Arioc implementation. We then carried out a read-by-read comparison of Arioc's reported alignments with the alignments reported by well-known CPU-based BS-seq read aligners. With simulated reads, Arioc's accuracy is equal to or better than the other read aligners we evaluated. With human sequencing reads, Arioc's throughput is at least 10 times faster than existing BS-seq aligners across a wide range of sensitivity settings. Availability and implementation: The Arioc software is available for download at https://github.com/RWilton/Arioc. It is released under a BSD open-source license. Supplementary information: Supplementary data are available at Bioinformatics online.
Richard Wilton, Xin Li 0131, Andrew P. Feinberg, Alex Szalay
Bioinform.1