Aaron Severance

dblp:15/9240 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 6 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Reconfigurable computing and FPGAs · 56% Processor architecture and microarchitecture · 27% Memory systems · 9%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs › reconfigurable architecture › reconfigurable processor
soft vector processor
0.742015
Wavefront Skipping using BRAMs for Conditional Algorithms on Vector Processors · FPGA 2015
Soft vector processors with streaming pipelines · FPGA 2014
Accelerator compiler for the VENICE vector processor · FPGA 2012
Processor architecture and microarchitecture
vector processor
0.232015
VEGAS: soft vector processor with scratchpad memory · FPGA 2011
Wavefront Skipping using BRAMs for Conditional Algorithms on Vector Processors · FPGA 2015
Soft vector processors with streaming pipelines · FPGA 2014
Compilers and program optimization
compiler back end
0.112012
Accelerator compiler for the VENICE vector processor · FPGA 2012
Memory systems › on-chip memory
scratchpad memory
0.112011
VEGAS: soft vector processor with scratchpad memory · FPGA 2011
Processor architecture and microarchitecture › instruction-level parallelism
predicated execution
0.112015
Wavefront Skipping using BRAMs for Conditional Algorithms on Vector Processors · FPGA 2015
Processor architecture and microarchitecture › microprocessor design › processor core design
datapath design
0.112014
Soft vector processors with streaming pipelines · FPGA 2014
Reconfigurable computing and FPGAs
FPGA accelerator
0.112014
Soft vector processors with streaming pipelines · FPGA 2014
Parallel and multicore computing
data-parallel programming
0.012012
Accelerator compiler for the VENICE vector processor · FPGA 2012
Reconfigurable computing and FPGAs › FPGA-based processor implementation
soft-core processor
0.012011
VEGAS: soft vector processor with scratchpad memory · FPGA 2011

Methods — techniques the papers use, named apart from their topics

high-level data-parallel compilation · 0.3wavefront skipping · 0.2BRAM-based offset storage · 0.2streaming pipeline · 0.2c-based programming · 0.2C macro API · 0.1
YearPublicationVenuePosition
2015 Wavefront Skipping using BRAMs for Conditional Algorithms on Vector Processors
abstract
Soft vector processors can accelerate data parallel algorithms on FPGAs while retaining software programmability. To handle divergent control flow, vector processors typically use mask registers and predicated instructions. These work by executing all branches and finally selecting the correct one. Our work improves FPGA based vector processors by adding wavefront skipping, where wavefronts that are completely masked off are skipped. This accelerates conditional algorithms, particularly useful where elements terminate early if simple tests fail but require extensive processing in the worst case. The difference in logic speed and RAM area for FPGA based circuits versus ASICs led us to a different implementation than used in fixed vector processors, storing wavefront offsets in on-chip BRAM rather than computing wavefronts skipped dynamically. Additionally, we allow for partitioning the wavefronts so that partial wavefronts can skip independently of one another. We show that <5% extra area can give up to 3.2× better performance on conditional algorithms. Partial wavefront skipping may not be generally useful enough to be added to a fixed vector processor; it provides up to 65% more performance for up to 27% more area. In an FGPA, however, the designer can use it to make application specific tradeoffs between area and performance.
Aaron Severance, Joe Edwards, Guy Lemieux
FPGA1
2014 Soft vector processors with streaming pipelines
abstract
Soft vector processors (SVPs) achieve significant performance gains through the use of parallel ALUs. However, since ALUs are used in a time-multiplexed fashion, this does not exploit a key strength of FPGA performance: pipeline parallelism. This paper shows how streaming pipelines can be integrated into the datapath of a SVP to achieve dramatic speedups. The SVP plays an important role in supplying the pipeline with high-bandwidth input data and storing its results using on-chip memory. However, the SVP must also perform the housekeeping tasks necessary to keep the pipeline busy. In particular, it orchestrates data movement between on-chip memory and external DRAM, it pre- or post-processes the data using its own ALUs, and it controls the overall sequence of execution. Since the SVP is programmed in C, these tasks are easier to develop and debug than using a traditional HDL approach. Using the N-body problem as a case study, this paper illustrates how custom streaming pipelines are integrated into the SVP datapath and multiple techniques for generating them. Using a custom pipeline, we demonstrate speedups over 7,000 times and performance-per-ALM over 100 times better than Nios II/f. The custom pipeline is also 50 times faster than a naive Intel Core i7 processor implementation.
Aaron Severance, Joe Edwards, Hossein Omidian, Guy Lemieux
FPGA1
2013 TputCache: High-frequency, multi-way cache for high-throughput FPGA applications
abstract
Throughput processing involves using many different contexts or threads to solve multiple problems or subproblems in parallel, where the size of the problem is large enough that latency can be tolerated. Bandwidth is required to support multiple concurrent executions, however, and utilizing multiple external memory channels is costly. For small working sets, FPGA designers can use on-chip BRAMs achieve the necessary bandwidth without increasing the system cost. Designing algorithms around fixed-size local memories is difficult, however, as there is no graceful fallback if the problem size exceeds the amount of local memory. This paper introduces TputCache, a cache designed to meet the needs of throughput processing on FPGAs, giving the throughput performance of on-chip BRAMs when the problem size fits in local memory. The design utilizes a replay based architecture to achieve high frequency with very low resource overheads.
Aaron Severance, Guy Lemieux
FPL1
2012 VENICE: A Compact Vector Processor for FPGA Applications
abstract
VENICE is a new soft vector processor (SVP) for FPGA applications that is designed for maximum through-put with a small number (1 to 4) of ALUs. By increasing clock speed and eliminating bottlenecks in ALU utilization, VENICE achieves over 2x better performance-per-logic block than VEGAS, the previous best SVP. VENICE is also simpler to program, as its instructions use standard C pointers into a scratchpad memory rather than vector registers.
Aaron Severance, Guy Lemieux
FCCM1
2012 Accelerator compiler for the VENICE vector processor
abstract
This paper describes the compiler design for VENICE, a new soft vector processor (SVP). The compiler is a new back-end target for Microsoft Accelerator, a high-level data parallel library for C++ and C#. This allows us to automatically compile high-level programs into VENICE assembly code, thus avoiding the process of writing assembly code used by previous SVPs. Experimental results show the compiler can generate scalable parallel code with execution times that are comparable to hand-written VENICE assembly code. On data-parallel applications, VENICE at 100MHz on an Altera DE3 platform runs at speeds comparable to one core of a 3.5GHz Intel Xeon W3690 processor, beating it in performance on four of six benchmarks by up to 3.2x.
Zhiduo Liu, Aaron Severance, Satnam Singh, Guy Lemieux
FPGA2
2012 Pipeline frequency boosting: Hiding dual-ported block RAM latency using intentional clock skew
abstract
FPGAs are increasingly being used to implement many new applications, including pipelined processor designs. Designers often employ memories to communicate and pass data between these pipeline stages. However, one-cycle communication between sender and receiver is often required. To implement this read-immediately-after-write functionality, bypass registers are needed by most FPGA memory blocks. Read and write latencies to these memories and the bypass can limit clock frequencies, or require extra resources to further pipeline the bypass. Instead of further pipelining the bypass, this paper applies clock skew scheduling to memory write and read ports of a simple bypass circuit. We show that the clock skew provides an improved Fmaxwithout requiring the area overhead of the pipelined bypass. Many configurations of pipelined memory systems are implemented, and their speed and area compared to our design. Memory clock skew scheduling yields the best Fmaxof all techniques which preserve functionality, an improvement of 56% over the baseline clock speed, and 14% over the best conventional design. Furthermore, the suggested technique consumes 46% fewer resources than the next best performing technique.
Alexander Brant, Ameer Abdelhadi, Aaron Severance, Guy Lemieux
FPT3
2012 VENICE: A compact vector processor for FPGA applications
abstract
This paper presents VENICE, a new soft vector processor (SVP) for FPGA applications. VENICE differs from previous SVPs in that it was designed for maximum throughput with a small number (1 to 4) of ALUs. By increasing clockspeed and eliminating bottlenecks in ALU utilization, VENICE can achieve over 2x better performance-per-logic block than VEGAS, the previous best SVP. While VENICE can scale to a large number of ALUs, a multiprocessor system of smaller VENICE SVPs is shown to scale better for benchmarks with limited innerloop parallelism. VENICE is also simpler to program, as its instructions use standard C pointers into a scratchpad memory rather than vector registers.
Aaron Severance, Guy Lemieux
FPT1
2011 VEGAS: soft vector processor with scratchpad memory
abstract
This paper presents VEGAS, a new soft vector architecture, in which the vector processor reads and writes directly to a scratchpad memory instead of a vector register file. The scratchpad memory is a more efficient storage medium than a vector register file, allowing up to 9x more data elements to fit into on-chip memory. In addition, the use of fracturable ALUs in VEGAS allow efficient processing of bytes, halfwords and words in the same processor instance, providing up to 4x the operations compared to existing fixed-width soft vector ALUs. Benchmarks show the new VEGAS architecture is 10x to 208x faster than Nios II and has 1.7x to 3.1x better area-delay product than previous vector work, achieving much higher throughput per unit area. To put this performance in perspective, VEGAS is faster than a leading-edge Intel processor at integer matrix multiply. To ease programming effort and provide full debug support, VEGAS uses a C macro API that outputs vector instructions as standard NIOS II/f custom instructions.
Christopher Han-Yu Chou, Aaron Severance, Alex D. Brant, Zhiduo Liu, Saurabh Sant, Guy Lemieux
FPGA2
2011 VENICE: A compact vector processor for FPGA applications
Aaron Severance, Guy Lemieux
Hot Chips Symposium1