Joe Edwards

dblp:141/9220 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Reconfigurable computing and FPGAs · 60% Processor architecture and microarchitecture · 32% Embedded and real-time systems · 8%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs › reconfigurable architecture › reconfigurable processor
soft vector processor
0.422015
Wavefront Skipping using BRAMs for Conditional Algorithms on Vector Processors · FPGA 2015
Soft vector processors with streaming pipelines · FPGA 2014
Processor architecture and microarchitecture
vector processor
0.122015
Wavefront Skipping using BRAMs for Conditional Algorithms on Vector Processors · FPGA 2015
Soft vector processors with streaming pipelines · FPGA 2014
Processor architecture and microarchitecture › instruction-level parallelism
predicated execution
0.112015
Wavefront Skipping using BRAMs for Conditional Algorithms on Vector Processors · FPGA 2015
Processor architecture and microarchitecture › microprocessor design › processor core design
datapath design
0.112014
Soft vector processors with streaming pipelines · FPGA 2014
Reconfigurable computing and FPGAs
FPGA accelerator
0.112014
Soft vector processors with streaming pipelines · FPGA 2014

Methods — techniques the papers use, named apart from their topics

wavefront skipping · 0.2BRAM-based offset storage · 0.2streaming pipeline · 0.2c-based programming · 0.2
YearPublicationVenuePosition
2017 Real-time object detection in software with custom vector instructions and algorithm changes
abstract
Real-time vision applications place stringent performance requirements on embedded systems. To meet performance requirements, embedded systems often require hardware implementations. This approach is unfavorable as hardware development can be difficult to debug, time-consuming, and require extensive skill. This paper presents a case study of accelerating face detection, often part of a complex image processing pipeline, using a software/hardware hybrid approach. As a baseline, the algorithm is initially run on a scalar ARM Cortex-A9 application processor found on a Xilinx Zynq device. Next, using a previously designed vector engine implemented in the FPGA fabric, the algorithm is vectorized, using only standard vector instructions, to achieve a 25× speedup. Then, we accelerate the critical inner loops by adding two hardware-assisted custom vector instructions for an additional 10× speedup, yielding 248× speedup over the initial Cortex-A9 baseline. Collectively, the custom instructions require fewer than 800 lines of VHDL code, including comments and blank lines. Compared to previous hardware-only face detection systems, our work is 1.5 to 6.8 times faster. This approach demonstrates that good performance can be obtained from software-only vectorization, and a small amount of custom hardware can provide a significant acceleration boost.
Joe Edwards, Guy Lemieux
ASAP1
2015 Wavefront Skipping using BRAMs for Conditional Algorithms on Vector Processors
abstract
Soft vector processors can accelerate data parallel algorithms on FPGAs while retaining software programmability. To handle divergent control flow, vector processors typically use mask registers and predicated instructions. These work by executing all branches and finally selecting the correct one. Our work improves FPGA based vector processors by adding wavefront skipping, where wavefronts that are completely masked off are skipped. This accelerates conditional algorithms, particularly useful where elements terminate early if simple tests fail but require extensive processing in the worst case. The difference in logic speed and RAM area for FPGA based circuits versus ASICs led us to a different implementation than used in fixed vector processors, storing wavefront offsets in on-chip BRAM rather than computing wavefronts skipped dynamically. Additionally, we allow for partitioning the wavefronts so that partial wavefronts can skip independently of one another. We show that <5% extra area can give up to 3.2× better performance on conditional algorithms. Partial wavefront skipping may not be generally useful enough to be added to a fixed vector processor; it provides up to 65% more performance for up to 27% more area. In an FGPA, however, the designer can use it to make application specific tradeoffs between area and performance.
Aaron Severance, Joe Edwards, Guy Lemieux
FPGA2
2014 Soft vector processors with streaming pipelines
abstract
Soft vector processors (SVPs) achieve significant performance gains through the use of parallel ALUs. However, since ALUs are used in a time-multiplexed fashion, this does not exploit a key strength of FPGA performance: pipeline parallelism. This paper shows how streaming pipelines can be integrated into the datapath of a SVP to achieve dramatic speedups. The SVP plays an important role in supplying the pipeline with high-bandwidth input data and storing its results using on-chip memory. However, the SVP must also perform the housekeeping tasks necessary to keep the pipeline busy. In particular, it orchestrates data movement between on-chip memory and external DRAM, it pre- or post-processes the data using its own ALUs, and it controls the overall sequence of execution. Since the SVP is programmed in C, these tasks are easier to develop and debug than using a traditional HDL approach. Using the N-body problem as a case study, this paper illustrates how custom streaming pipelines are integrated into the SVP datapath and multiple techniques for generating them. Using a custom pipeline, we demonstrate speedups over 7,000 times and performance-per-ALM over 100 times better than Nios II/f. The custom pipeline is also 50 times faster than a naive Intel Core i7 processor implementation.
Aaron Severance, Joe Edwards, Hossein Omidian, Guy Lemieux
FPGA2