VLDB 2026 Research / reviewers in the wild / expert
Jack Wadden
dblp:148/9820
· DBLP profile ↗
10ranked-venue papers
4as first author
2since 2021 · last 2021
0000-0002-3055-3656ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Hardware accelerators and domain-specific architectures · 42% High-performance computing · 18% Reconfigurable computing and FPGAs · 17% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
bioinformatics accelerator |
1.0 | 2 | 2021 | SquiggleFilter: An Accelerator for Portable Virus Detection · MICRO 2021 Accelerated Seeding for Genome Sequence Alignment with Enumerated Radix Trees · ISCA 2021 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.5 | 1 | 2021 | Accelerated Seeding for Genome Sequence Alignment with Enumerated Radix Trees · ISCA 2021 |
High-performance computing › scientific computing
genomic sequence comparison |
0.5 | 1 | 2021 | Accelerated Seeding for Genome Sequence Alignment with Enumerated Radix Trees · ISCA 2021 |
Hardware accelerators and domain-specific architectures
pattern matching accelerator |
0.4 | 1 | 2019 | Portable Programming with RAPID · IEEE Trans. Parallel Distributed Syst. 2019 |
High-performance computing › performance engineering
performance portability |
0.4 | 1 | 2019 | Portable Programming with RAPID · IEEE Trans. Parallel Distributed Syst. 2019 |
Reconfigurable computing and FPGAs
automata processing |
0.3 | 1 | 2018 | Characterizing and Mitigating Output Reporting Bottlenecks in Spatial Automata Processing Architectures · HPCA 2018 |
Hardware accelerators and domain-specific architectures
spatial architecture |
0.3 | 1 | 2018 | Characterizing and Mitigating Output Reporting Bottlenecks in Spatial Automata Processing Architectures · HPCA 2018 |
Performance modeling and evaluation
workload characterization |
0.3 | 1 | 2018 | Characterizing and Mitigating Output Reporting Bottlenecks in Spatial Automata Processing Architectures · HPCA 2018 |
GPUs and heterogeneous computing
GPU reliability |
0.2 | 1 | 2014 | Real-world design and evaluation of compiler-managed GPU redundant multithreading · ISCA 2014 |
Hardware reliability and fault tolerance › redundancy
redundant multithreading |
0.2 | 1 | 2014 | Real-world design and evaluation of compiler-managed GPU redundant multithreading · ISCA 2014 |
Hardware reliability and fault tolerance
soft errors |
0.2 | 1 | 2014 | Real-world design and evaluation of compiler-managed GPU redundant multithreading · ISCA 2014 |
Hardware reliability and fault tolerance
software fault tolerance |
0.2 | 1 | 2014 | Real-world design and evaluation of compiler-managed GPU redundant multithreading · ISCA 2014 |
Bioinformatics and computational biology › genomics
genome sequencing |
0.1 | 1 | 2021 | SquiggleFilter: An Accelerator for Portable Virus Detection · MICRO 2021 |
Bioinformatics and computational biology › sequence analysis
nanopore sequencing |
0.1 | 1 | 2021 | SquiggleFilter: An Accelerator for Portable Virus Detection · MICRO 2021 |
Bioinformatics and computational biology › sequence analysis
read mapping |
0.1 | 1 | 2021 | Accelerated Seeding for Genome Sequence Alignment with Enumerated Radix Trees · ISCA 2021 |
Bioinformatics and computational biology › sequence alignment
seed-and-extend |
0.1 | 1 | 2021 | Accelerated Seeding for Genome Sequence Alignment with Enumerated Radix Trees · ISCA 2021 |
Data mining
pattern mining |
0.1 | 1 | 2018 | Characterizing and Mitigating Output Reporting Bottlenecks in Spatial Automata Processing Architectures · HPCA 2018 |
Methods — techniques the papers use, named apart from their topics
hardware-software co-design · 1.0enumerated radix tree · 1.0FMD-Index · 1.0imperative-declarative programming model · 0.8automata conversion · 0.8output compression · 0.7architectural simulation · 0.7register-level thread communication · 0.4redundant multithreading · 0.4compiler pass · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Accelerated Seeding for Genome Sequence Alignment with Enumerated Radix TreesabstractRead alignment is a time-consuming step in genome sequencing analysis. The most widely used software for read alignment, BWA-MEM, and the recently published faster version BWA-MEM2 are based on the seed-and-extend paradigm for read alignment. The seeding step of read alignment is a major bottleneck contributing ~40% to the overall execution time of BWA-MEM2 when aligning whole human genome reads from the Platinum Genomes dataset. This is because both BWA-MEM and BWA-MEM2 use a compressed index structure called the FMD-Index, which results in high bandwidth requirements, primarily due to its character-by-character processing of reads. For instance, to seed each read (101 DNA base-pairs stored in 37.8 bytes), the FMD-Index solution in BWA-MEM2 requires ~68.5 KB of index data.We propose a novel indexing data structure named Enumerated Radix Tree (ERT) and design a custom seeding accelerator based on it. ERT improves bandwidth efficiency of BWA-MEM2 by 4.5× while guaranteeing 100% identical output to the original software, and still fitting in 64 GB DRAM. Overall, the proposed seeding accelerator implemented on AWS F1 FPGA (f1.4xlarge) improves seeding throughput of BWA-MEM2 by 3.3×. When combined with seed-extension accelerators, we observe a 2.1× improvement in overall read alignment throughput over BWA-MEM2. The software implementation of ERT is integrated into BWA-MEM2 (ert branch: https://github.com/bwa-mem2/bwa-mem2/tree/ert) and is open sourced for the benefit of the research community. Arun Subramaniyan 0001, Jack Wadden, Kush Goliya, Nathan Ozog, Xiao Wu 0002, Satish Narayanasamy, David T. Blaauw, Reetuparna Das |
ISCA | 2 |
| 2021 | SquiggleFilter: An Accelerator for Portable Virus DetectionabstractThe MinION is a recent-to-market handheld nanopore sequencer. It can be used to determine the whole genome of a target virus in a biological sample. Its Read Until feature allows us to skip sequencing a majority of non-target reads (DNA/RNA fragments), which constitutes more than 99% of all reads in a typical sample. However, it does not have any on-board computing, which significantly limits its portability. Timothy Dunn, Harisankar Sadasivan, Jack Wadden, Kush Goliya, Kuan-Yu Chen 0001, David T. Blaauw, Reetuparna Das, Satish Narayanasamy |
MICRO | 3 |
| 2019 | Portable Programming with RAPIDabstractAs the hardware found within data centers becomes more heterogeneous, it is important to allow for efficient execution of algorithms across architectures. We present RAPID, a high-level programming language and combined imperative and declarative model for functionally- and performance-portable execution of sequential pattern-matching applications across CPUs, GPUs, Field-Programmable Gate Arrays (FPGAs), and Micron's D480 AP. RAPID is clear, maintainable, concise, and efficient both at compile and run time. Language features, such as code abstraction and parallel control structures, map well to pattern-matching problems, providing clarity and maintainability. For generation of efficient runtime code, we present algorithms to convert RAPID programs into finite automata. Our empirical evaluation of applications in the ANMLZoo benchmark suite demonstrates that the automata processing paradigm provides an abstraction that is portable across architectures. We evaluate RAPID programs against custom, baseline implementations previously demonstrated to be significantly accelerated. We also find that RAPID programs are much shorter in length, are expressible at a higher level of abstraction than their handcrafted counterparts, and yield generated code that is often more compact. Kevin Angstadt, Jack Wadden, Westley Weimer, Kevin Skadron |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | Automata Processing in Reconfigurable Architectures: In-the-Cloud Deployment, Cross-Platform Evaluation, and Fast Symbol-Only ReconfigurationabstractWe present a general automata processing framework on FPGAs, which generates an RTL kernel for automata processing together with an AXI and PCIe based I/O circuitry. We implement the framework on both local nodes and cloud platforms (Amazon AWS and Nimbix) with novel features. A full performance comparison of the proposed framework is conducted against state-of-the-art automata processing engines on CPUs, GPUs, and Micron’s Automata Processor using the ANMLZoo benchmark suite and some real-world datasets. Results show that FPGAs enable extremely high-throughput automata processing compared to von Neumann architectures. We also collect the resource utilization and power consumption on the two cloud platforms, and find that the I/O circuitry consumes most of the hardware resources and power. Furthermore, we propose a fast, symbol-only reconfiguration mechanism based on the framework for large pattern sets that cannot fit on a single device and need to be partitioned. The proposed method supports multiple passes of the input stream and reduces the re-compilation cost from hours to seconds. Chunkun Bo, Vinh Dang, Ted Xie, Jack Wadden, Mircea R. Stan, Kevin Skadron |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2018 | Characterizing and Mitigating Output Reporting Bottlenecks in Spatial Automata Processing ArchitecturesabstractAutomata processing has seen a resurgence in importance due to its usefulness for pattern matching and pattern mining of "big data." While large-scale automata processing is known to bottleneck von Neumann processors due to unpredictable memory accesses, spatial architectures excel at automata processing. Spatial architectures can implement automata graphs by wiring together automata states in reconfigurable arrays, allowing parallel automata state computation, and point-to-point state transitions on-chip. However, spatial automata processing architectures can suffer from output constraints (up to 255x in commercial systems!) due to the physical placement of states, output processing architecture design, I/O resources, and the massively parallel nature of the architecture. To understand this bottleneck, we conduct the first known characterization of output requirements of a realistic set of automata processing benchmarks. We find that most benchmarks report fairly frequently, but that few states report at any one time. This observation motivates new output compression schemes and reporting architectures. We evaluate the benefit of one purely software automata transformation and show that output reporting costs can be greatly reduced (improving performance by up to 40% without hardware modification. We then explore bottlenecks in the reporting architecture of a commercial spatial automata processor and propose a new architecture that improves performance by up to 5.1x. Jack Wadden, Kevin Angstadt, Kevin Skadron |
HPCA | 1 |
| 2017 | Automata-to-Routing: An Open-Source Toolchain for Design-Space Exploration of Spatial Automata Processing ArchitecturesabstractNewly-available spatial architectures to accelerate finite-automata processing have spurred research and development on novel automata-based applications. However, spatial automata processing architecture research is lacking, because of a lack of automata optimization and place-and-route tools. To solve this issue, we propose a new, open-source toolchain-Automata-to-Routing (ATR)-that enables design-space exploration of spatial automata architectures. ATR leverages existing open-source tools for both automata processing and FPGA architecture research. To demonstrate the usefulness of this new toolchain, we use it to analyze design choices of spatial automata processing architectures. We first show that ATR is capable of modeling the logic tiles of a commercially-available spatial automata processing architecture. We then use ATR to compare and contrast two different routing architecture methodologies-hierarchical and 2D-mesh-over a set of diverse automata benchmarks. We show that shallower 2D-mesh-style routing fabrics can route complex automata with equal channel width, while using up to 4.2x fewer logic tile resources. Jack Wadden, Samira Manabi Khan, Kevin Skadron |
FCCM | 1 |
| 2017 | REAPR: Reconfigurable engine for automata processingabstractFinite automata have proven their usefulness in high-profile domains ranging from network security to machine learning. While prior work focused on their applicability for purely regular expression workloads such as antivirus and network security rulesets, recent research has shown that automata can optimize the performance for algorithms in other areas such as machine learning and even particle physics. Unfortunately, their emulation on traditional CPU architectures is fundamentally slow and further bottlenecked by memory. In this paper, we present REAPR: Reconfigurable Engine for Automata PRocessing, a flexible framework that synthesizes RTL for automata processing applications as well as I/O to handle data transfer to and from the kernel. We show that even with memory and control flow overheads, FPGAs still enable extremely high-throughput computation of automata workloads compared to other architectures. Ted Xie, Vinh Dang, Jack Wadden, Kevin Skadron, Mircea R. Stan |
FPL | 3 |
| 2016 | Generating efficient and high-quality pseudo-random behavior on Automata ProcessorsabstractMicron's Automata Processor (AP) efficiently emulates non-deterministic finite automata and has been shown to provide large speedups over traditional von Neumann execution for massively parallel, rule-based, data-mining and pattern matching applications. We demonstrate the AP's ability to generate high-quality and energy efficient pseudo-random behavior for use in pseudo-random number generation or in chip simulation. By recognizing that transition rules become probabilistic when input characters are randomized, the AP is also capable of simulating Markov chains. Combining hundreds of parallel Markov chains creates high-quality, high-throughput pseudo-random number sequences with greater power efficiency than state-of-the-art CPU and GPU algorithms. This indicates that the AP could potentially accelerate other Markov Chain-based applications such as agent-based simulation. We explore how to achieve throughputs upwards of 40GB/s per AP chip, with power efficiency 6.8x greater than state-of-the-art pseudo-random number generation on GPUs. Jack Wadden, Nathan Brunelle, Ke Wang 0011, Mohamed El-Hadedy 0001, Gabriel Robins, Mircea R. Stan, Kevin Skadron |
ICCD | 1 |
| 2015 | Regular expression acceleration on the micron automata processor: Brill tagging as a case studyabstractBrill tagging is a classic rule-based algorithm for part-of-speech (POS) tagging that assigns tags, such as nouns, verbs, adjectives, etc., to input tokens. Due to the the intense memory requirements of rule matching, CPU implementations of the Brill tagging algorithm have been found to be slow. We show that Micron's Automata Processor (AP) - a new computing architecture that can perform massively parallel pattern matching - can greatly accelerate the second stage of Brill tagging via rule template matching. The 218 contextual rules are first converted into regular expressions (regex). Regex is used widely in natural language processing (NLP) tasks, thus, this case study involving Brill Tagging also shows how the AP might accelerate other applications that are able to be framed as regexes. We compare single-threaded, and multithreaded versions of Regex matching on an Intel i7 CPU, an Intel XeonPhi co-processor, and the AP. The results show a 63.90X speed-up using the AP as a regex accelerator over the fastest multi-threaded CPU version. We also investigate how performance of regex matching on both CPU architectures varies depending on the complexity of the regex. Taken together, these results demonstrate the potential for significant performance improvements by using accelerators for various NLP computational tasks, particularly those that involve rule-based or pattern-matching approaches. Keira Zhou, Jack Wadden, Jeffrey J. Fox, Ke Wang 0011, Donald E. Brown, Kevin Skadron |
IEEE BigData | 2 |
| 2014 | Real-world design and evaluation of compiler-managed GPU redundant multithreadingabstractReliability for general purpose processing on the GPU (GPGPU) is becoming a weak link in the construction of reliable supercomputer systems. Because hardware protection is expensive to develop, requires dedicated on-chip resources, and is not portable across different architectures, the efficiency of software solutions such as redundant multithreading (RMT) must be explored. This paper presents a real-world design and evaluation of automatic software RMT on GPU hardware. We first describe a compiler pass that automatically converts GPGPU kernels into redundantly threaded versions. We then perform detailed power and performance evaluations of three RMT algorithms, each of which provides fault coverage to a set of structures in the GPU. Using real hardware, we show that compiler-managed software RMT has highly variable costs. We further analyze the individual costs of redundant work scheduling, redundant computation, and inter-thread communication, showing that no single component in general is responsible for high overheads across all applications; instead, certain workload properties tend to cause RMT to perform well or poorly. Finally, we demonstrate the benefit of architectural support for RMT with a specific example of fast, register-level thread communication. Jack Wadden, Alexander Lyashevsky, Sudhanva Gurumurthi, Vilas Sridharan, Kevin Skadron |
ISCA | 1 |