VLDB 2026 Research / reviewers in the wild / expert
Aaron Severance
dblp:15/9240
· DBLP profile ↗
9ranked-venue papers
6as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 6 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Reconfigurable computing and FPGAs · 56% Processor architecture and microarchitecture · 27% Memory systems · 9% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Reconfigurable computing and FPGAs › reconfigurable architecture › reconfigurable processor
soft vector processor |
0.7 | 4 | 2015 | Wavefront Skipping using BRAMs for Conditional Algorithms on Vector Processors · FPGA 2015 Soft vector processors with streaming pipelines · FPGA 2014 Accelerator compiler for the VENICE vector processor · FPGA 2012 |
Processor architecture and microarchitecture
vector processor |
0.2 | 3 | 2015 | VEGAS: soft vector processor with scratchpad memory · FPGA 2011 Wavefront Skipping using BRAMs for Conditional Algorithms on Vector Processors · FPGA 2015 Soft vector processors with streaming pipelines · FPGA 2014 |
Compilers and program optimization
compiler back end |
0.1 | 1 | 2012 | Accelerator compiler for the VENICE vector processor · FPGA 2012 |
Memory systems › on-chip memory
scratchpad memory |
0.1 | 1 | 2011 | VEGAS: soft vector processor with scratchpad memory · FPGA 2011 |
Processor architecture and microarchitecture › instruction-level parallelism
predicated execution |
0.1 | 1 | 2015 | Wavefront Skipping using BRAMs for Conditional Algorithms on Vector Processors · FPGA 2015 |
Processor architecture and microarchitecture › microprocessor design › processor core design
datapath design |
0.1 | 1 | 2014 | Soft vector processors with streaming pipelines · FPGA 2014 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.1 | 1 | 2014 | Soft vector processors with streaming pipelines · FPGA 2014 |
Parallel and multicore computing
data-parallel programming |
0.0 | 1 | 2012 | Accelerator compiler for the VENICE vector processor · FPGA 2012 |
Reconfigurable computing and FPGAs › FPGA-based processor implementation
soft-core processor |
0.0 | 1 | 2011 | VEGAS: soft vector processor with scratchpad memory · FPGA 2011 |
Methods — techniques the papers use, named apart from their topics
high-level data-parallel compilation · 0.3wavefront skipping · 0.2BRAM-based offset storage · 0.2streaming pipeline · 0.2c-based programming · 0.2C macro API · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | Wavefront Skipping using BRAMs for Conditional Algorithms on Vector ProcessorsabstractSoft vector processors can accelerate data parallel algorithms on FPGAs while retaining software programmability. To handle divergent control flow, vector processors typically use mask registers and predicated instructions. These work by executing all branches and finally selecting the correct one. Our work improves FPGA based vector processors by adding wavefront skipping, where wavefronts that are completely masked off are skipped. This accelerates conditional algorithms, particularly useful where elements terminate early if simple tests fail but require extensive processing in the worst case. The difference in logic speed and RAM area for FPGA based circuits versus ASICs led us to a different implementation than used in fixed vector processors, storing wavefront offsets in on-chip BRAM rather than computing wavefronts skipped dynamically. Additionally, we allow for partitioning the wavefronts so that partial wavefronts can skip independently of one another. We show that <5% extra area can give up to 3.2× better performance on conditional algorithms. Partial wavefront skipping may not be generally useful enough to be added to a fixed vector processor; it provides up to 65% more performance for up to 27% more area. In an FGPA, however, the designer can use it to make application specific tradeoffs between area and performance. Aaron Severance, Joe Edwards, Guy Lemieux |
FPGA | 1 |
| 2014 | Soft vector processors with streaming pipelinesabstractSoft vector processors (SVPs) achieve significant performance gains through the use of parallel ALUs. However, since ALUs are used in a time-multiplexed fashion, this does not exploit a key strength of FPGA performance: pipeline parallelism. This paper shows how streaming pipelines can be integrated into the datapath of a SVP to achieve dramatic speedups. The SVP plays an important role in supplying the pipeline with high-bandwidth input data and storing its results using on-chip memory. However, the SVP must also perform the housekeeping tasks necessary to keep the pipeline busy. In particular, it orchestrates data movement between on-chip memory and external DRAM, it pre- or post-processes the data using its own ALUs, and it controls the overall sequence of execution. Since the SVP is programmed in C, these tasks are easier to develop and debug than using a traditional HDL approach. Using the N-body problem as a case study, this paper illustrates how custom streaming pipelines are integrated into the SVP datapath and multiple techniques for generating them. Using a custom pipeline, we demonstrate speedups over 7,000 times and performance-per-ALM over 100 times better than Nios II/f. The custom pipeline is also 50 times faster than a naive Intel Core i7 processor implementation. Aaron Severance, Joe Edwards, Hossein Omidian, Guy Lemieux |
FPGA | 1 |
| 2013 | TputCache: High-frequency, multi-way cache for high-throughput FPGA applicationsabstractThroughput processing involves using many different contexts or threads to solve multiple problems or subproblems in parallel, where the size of the problem is large enough that latency can be tolerated. Bandwidth is required to support multiple concurrent executions, however, and utilizing multiple external memory channels is costly. For small working sets, FPGA designers can use on-chip BRAMs achieve the necessary bandwidth without increasing the system cost. Designing algorithms around fixed-size local memories is difficult, however, as there is no graceful fallback if the problem size exceeds the amount of local memory. This paper introduces TputCache, a cache designed to meet the needs of throughput processing on FPGAs, giving the throughput performance of on-chip BRAMs when the problem size fits in local memory. The design utilizes a replay based architecture to achieve high frequency with very low resource overheads. Aaron Severance, Guy Lemieux |
FPL | 1 |
| 2012 | VENICE: A Compact Vector Processor for FPGA ApplicationsabstractVENICE is a new soft vector processor (SVP) for FPGA applications that is designed for maximum through-put with a small number (1 to 4) of ALUs. By increasing clock speed and eliminating bottlenecks in ALU utilization, VENICE achieves over 2x better performance-per-logic block than VEGAS, the previous best SVP. VENICE is also simpler to program, as its instructions use standard C pointers into a scratchpad memory rather than vector registers. Aaron Severance, Guy Lemieux |
FCCM | 1 |
| 2012 | Accelerator compiler for the VENICE vector processorabstractThis paper describes the compiler design for VENICE, a new soft vector processor (SVP). The compiler is a new back-end target for Microsoft Accelerator, a high-level data parallel library for C++ and C#. This allows us to automatically compile high-level programs into VENICE assembly code, thus avoiding the process of writing assembly code used by previous SVPs. Experimental results show the compiler can generate scalable parallel code with execution times that are comparable to hand-written VENICE assembly code. On data-parallel applications, VENICE at 100MHz on an Altera DE3 platform runs at speeds comparable to one core of a 3.5GHz Intel Xeon W3690 processor, beating it in performance on four of six benchmarks by up to 3.2x. Zhiduo Liu, Aaron Severance, Satnam Singh, Guy Lemieux |
FPGA | 2 |
| 2012 | Pipeline frequency boosting: Hiding dual-ported block RAM latency using intentional clock skewabstractFPGAs are increasingly being used to implement many new applications, including pipelined processor designs. Designers often employ memories to communicate and pass data between these pipeline stages. However, one-cycle communication between sender and receiver is often required. To implement this read-immediately-after-write functionality, bypass registers are needed by most FPGA memory blocks. Read and write latencies to these memories and the bypass can limit clock frequencies, or require extra resources to further pipeline the bypass. Instead of further pipelining the bypass, this paper applies clock skew scheduling to memory write and read ports of a simple bypass circuit. We show that the clock skew provides an improved Fmaxwithout requiring the area overhead of the pipelined bypass. Many configurations of pipelined memory systems are implemented, and their speed and area compared to our design. Memory clock skew scheduling yields the best Fmaxof all techniques which preserve functionality, an improvement of 56% over the baseline clock speed, and 14% over the best conventional design. Furthermore, the suggested technique consumes 46% fewer resources than the next best performing technique. Alexander Brant, Ameer Abdelhadi, Aaron Severance, Guy Lemieux |
FPT | 3 |
| 2012 | VENICE: A compact vector processor for FPGA applicationsabstractThis paper presents VENICE, a new soft vector processor (SVP) for FPGA applications. VENICE differs from previous SVPs in that it was designed for maximum throughput with a small number (1 to 4) of ALUs. By increasing clockspeed and eliminating bottlenecks in ALU utilization, VENICE can achieve over 2x better performance-per-logic block than VEGAS, the previous best SVP. While VENICE can scale to a large number of ALUs, a multiprocessor system of smaller VENICE SVPs is shown to scale better for benchmarks with limited innerloop parallelism. VENICE is also simpler to program, as its instructions use standard C pointers into a scratchpad memory rather than vector registers. Aaron Severance, Guy Lemieux |
FPT | 1 |
| 2011 | VEGAS: soft vector processor with scratchpad memoryabstractThis paper presents VEGAS, a new soft vector architecture, in which the vector processor reads and writes directly to a scratchpad memory instead of a vector register file. The scratchpad memory is a more efficient storage medium than a vector register file, allowing up to 9x more data elements to fit into on-chip memory. In addition, the use of fracturable ALUs in VEGAS allow efficient processing of bytes, halfwords and words in the same processor instance, providing up to 4x the operations compared to existing fixed-width soft vector ALUs. Benchmarks show the new VEGAS architecture is 10x to 208x faster than Nios II and has 1.7x to 3.1x better area-delay product than previous vector work, achieving much higher throughput per unit area. To put this performance in perspective, VEGAS is faster than a leading-edge Intel processor at integer matrix multiply. To ease programming effort and provide full debug support, VEGAS uses a C macro API that outputs vector instructions as standard NIOS II/f custom instructions. Christopher Han-Yu Chou, Aaron Severance, Alex D. Brant, Zhiduo Liu, Saurabh Sant, Guy Lemieux |
FPGA | 2 |
| 2011 | VENICE: A compact vector processor for FPGA applications
Aaron Severance, Guy Lemieux |
Hot Chips Symposium | 1 |