EDBT 2026 Demo / reviewers in the wild / expert
Richard Wilton
dblp:73/5395
· DBLP profile ↗
6ranked-venue papers
5as first author
3since 2021 · last 2023
0000-0003-1263-5532ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 5 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Bioinformatics and computational biology · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
High-performance computing · 64% GPUs and heterogeneous computing · 36% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
sequence alignment |
1.2 | 2 | 2023 | Short-read aligner performance in germline variant identification · Bioinform. 2023 Performance optimization in DNA short-read alignment · Bioinform. 2022 |
Bioinformatics and computational biology › sequence analysis › read mapping
short read alignment |
1.2 | 2 | 2023 | Short-read aligner performance in germline variant identification · Bioinform. 2023 Performance optimization in DNA short-read alignment · Bioinform. 2022 |
Bioinformatics and computational biology › genomics
variant calling |
0.7 | 1 | 2023 | Short-read aligner performance in germline variant identification · Bioinform. 2023 |
Bioinformatics and computational biology › genomics
genomic data management |
0.4 | 1 | 2019 | The Terabase Search Engine: a large-scale relational database of short-read sequences · Bioinform. 2019 |
Bioinformatics and computational biology › biological database
sequence database |
0.4 | 1 | 2019 | The Terabase Search Engine: a large-scale relational database of short-read sequences · Bioinform. 2019 |
Bioinformatics and computational biology › epigenomics › DNA methylation › DNA methylation analysis
bisulfite sequencing |
0.3 | 1 | 2018 | Arioc: GPU-accelerated alignment of short bisulfite-treated reads · Bioinform. 2018 |
Bioinformatics and computational biology › sequence analysis
read mapping |
0.3 | 1 | 2018 | Arioc: GPU-accelerated alignment of short bisulfite-treated reads · Bioinform. 2018 |
Bioinformatics and computational biology › genomics › variant calling
germline variant calling |
0.2 | 1 | 2023 | Short-read aligner performance in germline variant identification · Bioinform. 2023 |
High-performance computing
performance optimization at scale |
0.2 | 1 | 2022 | Performance optimization in DNA short-read alignment · Bioinform. 2022 |
GPUs and heterogeneous computing
GPU computing |
0.1 | 1 | 2018 | Arioc: GPU-accelerated alignment of short bisulfite-treated reads · Bioinform. 2018 |
Methods — techniques the papers use, named apart from their topics
benchmarking · 1.8performance profiling · 1.1relational indexing · 0.8GPU acceleration · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Short-read aligner performance in germline variant identificationabstractMOTIVATION: Read alignment is an essential first step in the characterization of DNA sequence variation. The accuracy of variant-calling results depends not only on the quality of read alignment and variant-calling software but also on the interaction between these complex software tools. RESULTS: In this review, we evaluate short-read aligner performance with the goal of optimizing germline variant-calling accuracy. We examine the performance of three general-purpose short-read aligners-BWA-MEM, Bowtie 2, and Arioc-in conjunction with three germline variant callers: DeepVariant, FreeBayes, and GATK HaplotypeCaller. We discuss the behavior of the read aligners with regard to the data elements on which the variant callers rely, and illustrate how the runtime configurations of these software tools combine to affect variant-calling performance. AVAILABILITY AND IMPLEMENTATION: The quick brown fox jumps over the lazy dog. Richard Wilton, Alex Szalay |
Bioinform. | 1 |
| 2023 | BiocMAP: a Bioconductor-friendly, GPU-accelerated pipeline for bisulfite-sequencing dataabstractBACKGROUND: Bisulfite sequencing is a powerful tool for profiling genomic methylation, an epigenetic modification critical in the understanding of cancer, psychiatric disorders, and many other conditions. Raw data generated by whole genome bisulfite sequencing (WGBS) requires several computational steps before it is ready for statistical analysis, and particular care is required to process data in a timely and memory-efficient manner. Alignment to a reference genome is one of the most computationally demanding steps in a WGBS workflow, taking several hours or even days with commonly used WGBS-specific alignment software. This naturally motivates the creation of computational workflows that can utilize GPU-based alignment software to greatly speed up the bottleneck step. In addition, WGBS produces raw data that is large and often unwieldy; a lack of memory-efficient representation of data by existing pipelines renders WGBS impractical or impossible to many researchers. RESULTS: We present BiocMAP, a Bioconductor-friendly methylation analysis pipeline consisting of two modules, to address the above concerns. The first module performs computationally-intensive read alignment using Arioc, a GPU-accelerated short-read aligner. Since GPUs are not always available on the same computing environments where traditional CPU-based analyses are convenient, the second module may be run in a GPU-free environment. This module extracts and merges DNA methylation proportions-the fractions of methylated cytosines across all cells in a sample at a given genomic site. Bioconductor-based output objects in R utilize an on-disk data representation to drastically reduce required main memory and make WGBS projects computationally feasible to more researchers. CONCLUSIONS: BiocMAP is implemented using Nextflow and available at http://research.libd.org/BiocMAP/ . To enable reproducible analysis across a variety of typical computing environments, BiocMAP can be containerized with Docker or Singularity, and executed locally or with the SLURM or SGE scheduling engines. By providing Bioconductor objects, BiocMAP's output can be integrated with powerful analytical open source software for analyzing methylation data. Nicholas J. Eagles, Richard Wilton, Andrew E. Jaffe, Leonardo Collado-Torres |
BMC Bioinform. | 2 |
| 2022 | Performance optimization in DNA short-read alignmentabstractSUMMARY: Over the past decade, short-read sequence alignment has become a mature technology. Optimized algorithms, careful software engineering and high-speed hardware have contributed to greatly increased throughput and accuracy. With these improvements, many opportunities for performance optimization have emerged. In this review, we examine three general-purpose short-read alignment tools-BWA-MEM, Bowtie 2 and Arioc-with a focus on performance optimization. We analyze the performance-related behavior of the algorithms and heuristics each tool implements, with the goal of arriving at practical methods of improving processing speed and accuracy. We indicate where an aligner's default behavior may result in suboptimal performance, explore the effects of computational constraints such as end-to-end mapping and alignment scoring threshold, and discuss sources of imprecision in the computation of alignment scores and mapping quality. With this perspective, we describe an approach to tuning short-read aligner performance to meet specific data-analysis and throughput requirements while avoiding potential inaccuracies in subsequent analysis of alignment results. Finally, we illustrate how this approach avoids easily overlooked pitfalls and leads to verifiable improvements in alignment speed and accuracy. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Appendices referenced in this article are available at Bioinformatics online. Richard Wilton, Alex Szalay |
Bioinform. | 1 |
| 2020 | Arioc: High-concurrency short-read alignment on multiple GPUsabstractIn large DNA sequence repositories, archival data storage is often coupled with computers that provide 40 or more CPU threads and multiple GPU (general-purpose graphics processing unit) devices. This presents an opportunity for DNA sequence alignment software to exploit high-concurrency hardware to generate short-read alignments at high speed. Arioc, a GPU-accelerated short-read aligner, can compute WGS (whole-genome sequencing) alignments ten times faster than comparable CPU-only alignment software. When two or more GPUs are available, Arioc's speed increases proportionately because the software executes concurrently on each available GPU device. We have adapted Arioc to recent multi-GPU hardware architectures that support high-bandwidth peer-to-peer memory accesses among multiple GPUs. By modifying Arioc's implementation to exploit this GPU memory architecture we obtained a further 1.8x-2.9x increase in overall alignment speeds. With this additional acceleration, Arioc computes two million short-read alignments per second in a four-GPU system; it can align the reads from a human WGS sequencer run-over 500 million 150nt paired-end reads-in less than 15 minutes. As WGS data accumulates exponentially and high-concurrency computational resources become widespread, Arioc addresses a growing need for timely computation in the short-read data analysis toolchain. Richard Wilton, Alex Szalay |
PLoS Comput. Biol. | 1 |
| 2019 | The Terabase Search Engine: a large-scale relational database of short-read sequencesabstractMOTIVATION: DNA sequencing archives have grown to enormous scales in recent years, and thousands of human genomes have already been sequenced. The size of these data sets has made searching the raw read data infeasible without high-performance data-query technology. Additionally, it is challenging to search a repository of short-read data using relational logic and to apply that logic across samples from multiple whole-genome sequencing samples. RESULTS: We have built a compact, efficiently-indexed database that contains the raw read data for over 250 human genomes, encompassing trillions of bases of DNA, and that allows users to search these data in real-time. The Terabase Search Engine enables retrieval from this database of all the reads for any genomic location in a matter of seconds. Users can search using a range of positions or a specific sequence that is aligned to the genome on the fly. AVAILABILITY AND IMPLEMENTATION: Public access to the Terabase Search Engine database is available at http://tse.idies.jhu.edu. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Richard Wilton, Sarah J. Wheelan, Alex Szalay, Steven Salzberg |
Bioinform. | 1 |
| 2018 | Arioc: GPU-accelerated alignment of short bisulfite-treated readsabstractMotivation: The alignment of bisulfite-treated DNA sequences (BS-seq reads) to a large genome involves a significant computational burden beyond that required to align non-bisulfite-treated reads. In the analysis of BS-seq data, this can present an important performance bottleneck that can be mitigated by appropriate algorithmic and software-engineering improvements. One strategy is to modify the read-alignment algorithms by integrating the logic related to BS-seq alignment, with the goal of making the software implementation amenable to optimizations that lead to higher speed and greater sensitivity than might otherwise be attainable. Results: We evaluated this strategy using Arioc, a short-read aligner that uses GPU (general-purpose graphics processing unit) hardware to accelerate computationally-expensive programming logic. We integrated the BS-seq computational logic into both GPU and CPU code throughout the Arioc implementation. We then carried out a read-by-read comparison of Arioc's reported alignments with the alignments reported by well-known CPU-based BS-seq read aligners. With simulated reads, Arioc's accuracy is equal to or better than the other read aligners we evaluated. With human sequencing reads, Arioc's throughput is at least 10 times faster than existing BS-seq aligners across a wide range of sensitivity settings. Availability and implementation: The Arioc software is available for download at https://github.com/RWilton/Arioc. It is released under a BSD open-source license. Supplementary information: Supplementary data are available at Bioinformatics online. Richard Wilton, Xin Li 0131, Andrew P. Feinberg, Alex Szalay |
Bioinform. | 1 |