Lesley Shannon

dblp:42/1149 · DBLP profile ↗
← Back
59ranked-venue papers
11as first author
13since 2021 · last 2024
0000-0002-7050-6184ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 58 · 11 first-author · 13 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2024 A Semi Black-Box Adversarial Bit- Flip Attack with Limited DNN Model Information
abstract
Despite the rising prevalence of deep neural networks (DNNs) in cyber-physical systems, their vulnerability to adversarial bit-flip attacks (BFAs) is a noteworthy concern. This paper proposes B3FA, a semi-black-box BFA-based parameter attack on DNNs, assuming the adversary has limited knowledge about the model. We consider practical scenarios often feature a more restricted threat model for real-world systems, contrasting with the typical BFA models that presuppose the adversary's full access to a network's inputs and parameters. The introduced bit-flip approach utilizes a magnitude-based ranking method and a statistical reconstruction technique to identify the vulnerable bits. We demonstrate the effectiveness of B3FA on several DNN models in a semi-black-box setting. For example, B3FA could drop the accuracy of a MobileNetV2 from 69.84% to 9% with only 20 bit-flips in a real-world setting.
Behnam Ghavami, Mani Sadati, Mohammad Shahidzadeh, Lesley Shannon, Steve Wilton
ICCD4
2024 ZOBNN: Zero-Overhead Dependable Design of Binary Neural Networks with Deliberately Quantized Parameters
abstract
Low-precision weights and activations in deep neural networks (DNNs) outperform their full-precision counterparts in terms of hardware efficiency. When implemented with low-precision operations, specifically in the extreme case where network parameters are binarized (i.e. BNNs), the two most frequently mentioned benefits of quantization are reduced memory consumption and a faster inference process. In this paper, we introduce a third advantage of very low-precision neural networks: improved fault-tolerance attribute. We investigate the impact of memory faults on state-of-the-art binary neural networks (BNNs) through comprehensive analysis. Despite the inclusion of floating-point parameters in BNN architectures to improve accuracy, our findings reveal that BNNs are highly sensitive to deviations in these parameters caused by memory faults. In light of this crucial finding, we propose a technique to improve BNN dependability by restricting the range of float parameters through a novel deliberately uniform quantization. The introduced quantization technique results in a reduction in the proportion of floating-point parameters utilized in the BNN, without incurring any additional computational overheads during the inference stage. The extensive experimental fault simulation on the proposed BNN architecture (i.e. ZOBNN) reveal a remarkable 5X enhancement in robustness compared to conventional floating-point DNN. Notably, this improvement is achieved without incurring any computational overhead. Crucially, this enhancement comes without computational overhead. ZOBNN excels in critical edge applications characterized by limited computational resources, prioritizing both dependability and real-time performance.
Behnam Ghavami, Mohammad Shahidzadeh, Lesley Shannon, Steve Wilton
IOLTS3
2024 Designing an IEEE-Compliant FPU that Supports Configurable Precision for Soft Processors
abstract
Field Programmable Gate Arrays (FPGAs) are commonly used to accelerate floating-point (FP) applications. Although researchers have extensively studied FPGA FP implementations, existing work has largely focused on standalone operators and frequency-optimized designs. These works are not suitable for FPGA soft processors which are more sensitive to latency, impose a lower frequency ceiling, and require IEEE FP standard compliance. We present an open-source floating-point unit (FPU) for FPGA RISC-V soft processors that is fully IEEE compliant with configurable levels of FP precision. Our design emphasizes runtime performance with 25% lower latency in the most common instructions compared to previous works while maintaining efficient resource utilization. Our FPU also allows users to explore various mantissa widths without having to rewrite or recompile their algorithms. We use this to investigate the scalability of our reduced-precision FPU across numerous microbenchmark functions as well as more complex case studies. Our experiments show that applications like the discrete cosine transformation and the Black-Scholes model can realize a speedup of more than 1.35x in conjunction with a 43% and 35% reduction in lookup table and flip-flop resources while experiencing less than a 0.025% average loss in numerical accuracy with a 16-bit mantissa width.
Chris Keilbart, Yuhui Gao, Martin Chua, Eric Matthews, Steve Wilton, Lesley Shannon
ACM Trans. Reconfigurable Technol. Syst.6
2023 Designing a configurable IEEE-compliant FPU that supports variable precision for soft processors
abstract
FPGAs are an increasingly popular medium for many high-performance data center workloads and the rapidly-expanding artificial intelligence domain. These applications often make extensive use of floating-point (FP) numbers defined by the IEEE 754 standard [1]. Although researchers have extensively studied FPGA-based hardware FP implementations, existing work has largely focused on standalone and throughput-optimized data-path designs. Such designs optimize performance by increasing throughput with long pipelines and high frequencies. This approach is not suitable for soft processors, which are more sensitive to latency in order to reduce stalls due to data hazards. Additionally, the frequency ceiling imposed by other internal components of the soft processor necessarily limits the maximum operating frequency of the Floating-Point Unit (FPU).
Chris Keilbart, Yuhui Gao, Martin Chua, Eric Matthews, Steve Wilton, Lesley Shannon
FCCM6
2022 FitAct: Error Resilient Deep Neural Networks via Fine-Grained Post-Trainable Activation Functions
abstract
Deep neural networks (DNNs) are increasingly being deployed in safety-critical systems such as personal healthcare devices and self-driving cars. In such DNN-based systems, error resilience is a top priority since faults in DNN inference could lead to mispredictions and safety hazards. For latency-critical DNN inference on resource-constrained edge devices, it is nontrivial to apply conventional redundancy-based fault tolerance techniques. In this paper, we propose FitAct, a low-cost approach to enhance the error resilience of DNNs by deploying fine-grained post-trainable activation functions. The main idea is to precisely bound the activation value of each individual neuron via neuron-wise bounded activation functions, so that it could prevent the fault propagation in the network. To avoid complex DNN model re-training, we propose to decouple the accuracy training and resilience training, and develop a lightweight post-training phase to learn these activation functions with precise bound values. Experimental results on widely used DNN models such as AlexNet, VGG16, and ResNet50 demonstrate that FitAct outperform state-of-the-art studies such as Clip-Act and Ranger in enhancing the DNN error resilience for a wide range of fault rates, while adding manageable runtime and memory space overheads.
Behnam Ghavami, Mani Sadati, Zhenman Fang, Lesley Shannon
DATE4
2022 A Majority-based Approximate Adder for FPGAs
abstract
The most advanced ASIC-based approximate adders are focused on gate or transistor level approximating structures. However, due to architectural differences between ASIC and FPGA, comparable performance gains for FPGA-based approximate adders cannot be obtained using ASIC-based approximation ones. In this paper, we propose a method for designing a low-error approximate adder that effectively deploys the modern FPGA structure. We introduce an FPGA-based approximate adder, named as Majority Approximate Adder (MAA), with less error than the advanced approximate adders. MAA is constructed using an approximate part and an accurate one; i.e. the accurate part is based on a smaller carry-chain compared with the carry-chain of the corresponding accurate adder. In addition, approximate part is designed to use FPGA resources efficiently with a low mean error distance (MED). Experimental results based on Monte-Carlo simulation demonstrates that a 16-bit MAA has a 49.92% lower MED than the state of the art FPGA-based approximate adder. MAA also takes up less area and consumes less power than other FPGA-based approximate adders in the literature.
Behnam Ghavami, Mahdi Sajedi, Mohsen Raji, Zhenman Fang, Lesley Shannon
DSD5
2022 Blind Data Adversarial Bit-flip Attack against Deep Neural Networks
abstract
Because of their high accuracy, deep neural net-works (DNNs) have achieved amazing success in security-critical systems such as medical devices. It has recently been demon-strated that Adversarial Bit Flip Attacks (BFAs) against DNN hardware by flipping a very small number of bits can result in catastrophic accuracy loss. The reliance on test data, however, is a significant drawback of previous state-of-the-art bit-flip attack methods. This is frequently not possible with applications containing sensitive or proprietary data. In this paper, we propose Blind Data Adversarial Bit-flip Attack (BDFA), a novel technique to enable BFA against DNN hardware without any access to the training or testing data. This is achieved by optimizing for a synthetic dataset, which is engineered to match the statistics of batch normalization across different layers of the network and the targeted label. Experimental results show that BDFA could decrease the accuracy of ResNet50 significantly from 75.96% to 13.94% with only 4 bits flips.
Behnam Ghavami, Mani Sadati, Mohammad Shahidzadeh, Zhenman Fang, Lesley Shannon
DSD5
2022 Demystifying the Soft and Hardened Memory Systems of Modern FPGAs for Software Programmers through Microbenchmarking
abstract
Both modern datacenter and embedded Field Programmable Gate Arrays (FPGAs) provide great opportunities for high-performance and high-energy-efficiency computing. With the growing public availability of FPGAs from major cloud service providers such as AWS, Alibaba, and Nimbix, as well as uniform hardware accelerator development tools (such as Xilinx Vitis and Intel oneAPI) for software programmers, hardware and software developers can now easily access FPGA platforms. However, it is nontrivial to develop efficient FPGA accelerators, especially for software programmers who use high-level synthesis (HLS). The major goal of this article is to figure out how to efficiently access the memory system of modern datacenter and embedded FPGAs in HLS-based accelerator designs. This is especially important for memory-bound applications; for example, a naive accelerator design only utilizes less than 5% of the available off-chip memory bandwidth. To achieve our goal, we first identify a comprehensive set of factors that affect the memory bandwidth, including (1) the clock frequency of the accelerator design, (2) the number of concurrent memory access ports, (3) the data width of each port, (4) the maximum burst access length for each port, and (5) the size of consecutive data accesses. Then, we carefully design a set of HLS-based microbenchmarks to quantitatively evaluate the performance of the memory systems of datacenter FPGAs (Xilinx Alveo U200 and U280) and embedded FPGA (Xilinx ZCU104) when changing those affecting factors, and we provide insights into efficient memory access in HLS-based accelerator designs. Comparing between the typically used soft and hardened memory systems, respectively, found on datacenter and embedded FPGAs, we further summarize their unique features and discuss the effective approaches to leverage these systems. To demonstrate the usefulness of our insights, we also conduct two case studies to accelerate the widely used K-nearest neighbors (KNN) and sparse matrix-vector multiplication (SpMV) algorithms on datacenter FPGAs with a soft (and thus more flexible) memory system. Compared to the baseline designs, optimized designs leveraging our insights achieve about \( 3.5\times \) and \( 8.5\times \) speedups for the KNN and SpMV accelerators. Our final optimized KNN and SpMV designs on a Xilinx Alveo U200 FPGA fully utilize its off-chip memory bandwidth, and achieve about \( 5.6\times \) and \( 3.4\times \) speedups over the 24-core CPU implementations.
Alec Lu, Zhenman Fang, Lesley Shannon
ACM Trans. Reconfigurable Technol. Syst.3
2022 Quick-Div: Rethinking Integer Divider Design for FPGA-based Soft-processors
abstract
In today’s FPGA-based soft-processors, one of the slowest instructions is integer division. Compared to the low single-digit latency of other arithmetic operations, the fixed 32-cycle latency of radix-2 division is substantially longer. Given that today’s soft-processors typically only implement radix-2 division—if they support hardware division at all—there is significant potential to improve the performance of integer dividers. In this work, we present a set of high-performance, data-dependent, variable-latency integer dividers for FPGA-based soft-processors that we call Quick-Div . We compare them to various radix-N dividers and provide a thorough analysis in terms of latency and resource usage. In addition, we analyze the frequency scaling for such divider designs when (1) treated as a stand-alone unit and (2) integrated as part of a high-performance soft-processor. Moreover, we provide additional theoretical analysis of different dividers’ behaviour and develop a new better-performing Quick-Div variant, called Quick-radix-4 . Experimental results show that our Quick-radix-4 design can achieve up to 6.8× better performance and 6.1× better performance-per-LUT over the radix-2 divider for applications such as random number generation. Even in cases where division operations constitute as little as 1% of all executed instructions, Quick-radix-4 provides a performance uplift of 16% compared to the radix-2 divider.
Eric Matthews, Alec Lu, Zhenman Fang, Lesley Shannon
ACM Trans. Reconfigurable Technol. Syst.4
2022 Introduction to Special Section on FPGA 2020
abstract
No abstract available.
Lesley Shannon
ACM Trans. Reconfigurable Technol. Syst.1
2021 LEAP: A Deep Learning based Aging-Aware Architecture Exploration Framework for FPGAs
abstract
Transistor aging raises a vital lifetime reliability challenge for FPGA devices in advanced technology nodes. In this paper, we design a tool called LEAP to enable the aging-aware FPGA architecture exploration. The core idea of LEAP is to efficiently model the aging-induced delay degradation at the coarse-grained FPGA basic block level using deep neural networks (DNNs), while achieving almost the same accuracy as the transistor-level simulation. For each type of the FPGA basic block such as LUT and DSP, we first characterize its accurate delay degradation via transistor-level SPICE simulation under a versatile set of aging factors from the FPGA fabric and in-field operation. Then we train one DNN model for each block type to learn the relation between its delay degradation and aging factors. Moreover, we integrate our DNN models into the widely used Verilog-to-Routing (VTR 8) toolflow and generate the aging-aware FPGA architecture file. Experimental results demonstrate that our proposed flow can predict the delay degradation of FPGA blocks more than 104x to 107x faster than transistor-level SPICE simulation, with the maximum prediction error of less than 0.7%. Therefore, FPGA architects can leverage LEAP to explore better aging-aware FPGA architectures.
Behnam Ghavami, Seyed Milad Ebrahimipour, Zhenman Fang, Lesley Shannon
FPGA4
2021 Demystifying the Memory System of Modern Datacenter FPGAs for Software Programmers through Microbenchmarking
abstract
With the public availability of FPGAs from major cloud service providers like AWS, Alibaba, and Nimbix, hardware and software developers can now easily access FPGA platforms. However, it is nontrivial to develop efficient FPGA accelerators, especially for software programmers who use high-level synthesis (HLS).
Alec Lu, Zhenman Fang, Lesley Shannon
FPGA4
2021 MAPLE: A Machine Learning based Aging-Aware FPGA Architecture Exploration Framework
abstract
In this paper, we develop a framework called MAPLE to enable the aging-aware FPGA architecture exploration. The core idea is to efficiently model the aging-induced delay degradation at the coarse-grained FPGA basic block level using deep neural networks (DNNs). For each type of the FPGA basic block such as LUT and DSP, we first characterize its accurate delay degradation via transistor-level SPICE simulation under a versatile set of aging factors from the FPGA fabric and in-field operation. Then we train one DNN model for each block type to quickly and accurately predict the complex relation between its delay degradation and comprehensive aging factors. Moreover, we integrate our DNN models into the widely used Verilog-to-Routing toolflow (VTR 8) to support analyzing the impact of aging-induced delay degradation on the entire large-scale FPGA architecture. Experimental results demonstrate that MAPLE can predict the delay degradation of FPGA blocks 104to 107times faster than transistor-level SPICE simulation, with a prediction error less than 0.7%. Our case study demonstrates that FPGA architects can effectively leverage MAPLE to explore better aging-aware FPGA architectures.
Behnam Ghavami, Milad Ibrahimipour, Zhenman Fang, Lesley Shannon
FPL4
2020 Exploring Writeback Designs for Efficiently Leveraging Parallel-Execution Units in FPGA-Based Soft-Processors
abstract
Maximizing processor performance depends on maximizing the product of instruction throughput and clock frequency. Writeback mechanisms and forwarding networks heavily impact both of these properties along with the resource usage and scalability of the processor design. Furthermore, these mechanisms are typically multiplexer heavy which can make their implementation resource inefficient on FPGAs. In this paper, we explore multiple different writeback and result storage mechanisms using an FPGA-based RISC-V soft-processor (Taiga), exploring both exception-safe and non-exception-safe designs. Writeback mechanisms based on per-unit result storage and centralized storage are explored while leveraging FPGA specific resources such as LUTRAMs. We evaluate the designs based on their impact on instruction throughput, processor frequency, and scalability of both simultaneous instructions in-flight and the number of execution units. As each design has different characteristics, we focus on comparing and contrasting the designs. We find that across all designs, average IPC can vary by up to 11%, with a few designs reaching the maximum IPC of one for some benchmarks. Clock frequency is found to vary by up to 20% across the designs, but is not significantly impacted when increasing the number of execution units. Scaling up the instructions in-flight is found to have the greatest variability, with LUT usage increasing by 3% to 93% across the different designs. Overall, we find that under current constraints, a commit-buffer design provides the highest combination of performance and performance per LUT.
Eric Matthews, Yuhui Gao, Lesley Shannon
FCCM3
2020 Aadam: A Fast, Accurate, and Versatile Aging-Aware Cell Library Delay Model using Feed-Forward Neural Network
abstract
With the CMOS technology scaling, transistor aging has become one major issue affecting circuit reliability and lifetime. There are two major classes of existing studies that model the aging effects in the circuit delay. One is at transistor-level, which is highly accurate but very slow. The other is at gate-level, which is faster but less accurate. Moreover, most prior studies only consider a limited subset or limited value ranges of aging factors.
Seyed Milad Ebrahimipour, Behnam Ghavami, Mohsen Raji, Zhenman Fang, Lesley Shannon
ICCAD6
2019 Rethinking Integer Divider Design for FPGA-Based Soft-Processors
abstract
Most existing soft-processors on FPGAs today support a fixed-latency instruction pipeline. Therefore, for integer division, a simple fixed-latency radix-2 integer divider is typically used, or algorithm-level changes are made to avoid integer divisions. However, for certain important application domains the simple radix-2 integer divider becomes the performance bottleneck, as every 32-bit division operation takes 32 cycles. In this paper, we explore integer divider designs for FPGA-based soft-processors, by leveraging the recent support of variable-latency execution units in their instruction pipeline. We implement a high-performance, data-dependent, variable-latency integer divider called Quick-Div, optimize its performance on FPGAs, and integrate it into a RISC-V soft-processor called Taiga that supports a variable-latency instruction pipeline. We perform a comprehensive analysis and comparison-in terms of cycles, clock frequency, and resource usage-for both the fixed-latency radix-2/4/8/16 dividers and our variable-latency Quick-Div divider with various optimizations. Experimental results on a Xilinx Virtex UltraScale+ VCU118 FPGA board show that our Quick-Div divider can provide over 5x better performance and over 4x better performance/LUT compared to a radix-2 divider for certain applications like random number generation. Finally, through a case study of integer square root, we demonstrate that our Quick-Div divider provides opportunities for reconsidering simpler and faster algorithmic choices.
Eric Matthews, Alec Lu, Zhenman Fang, Lesley Shannon
FCCM4
2019 Efficient PUF-Based Key Generation in FPGAs Using Per-Device Configuration
abstract
Reconfigurable systems often require secret keys to encrypt and decrypt data. Applications requiring high security commonly generate keys based on physical unclonable functions (PUFs), circuits that use random manufacturing variations to produce secret keys that are unique to each device. Implementing PUFs on field-programmable gate arrays (FPGAs) is usually difficult, because the designer has limited control over layout, and each PUF system requires a large area overhead to correct errors in the PUF response bits. In this paper, we extend the state of the art for FPGA-based weak PUFs using a novel methodology of per-device configuration and a new PUF variant derived from the popular FPGA-specific Anderson PUF. The PUF is evaluated using Xilinx XC7Z020 programmable systemon-chips from the Virtex-7 family on Zynq ZedBoard platforms. The design we propose has several advantages over existing work including the Anderson PUF on which it is based. Our design is tunable to minimize the response bias and can be implemented using the common SLICEL components on Xilinx FPGAs. Moreover, the proposed PUF design enables an efficient per-device configuration that reduces bit error rate by over 10× at room temperature and improves response stability by over 2× across all temperatures. We demonstrate that the proposed per-device PUF configuration step leads to roughly 2× savings in area resources for PUFs and error correction as used in key generation.
Mohammad A. Usmani, Shahrzad Keshavarz, Eric Matthews, Lesley Shannon, Russell Tessier, Daniel E. Holcomb
IEEE Trans. Very Large Scale Integr. Syst.4
2018 Evaluating the Performance Efficiency of a Soft-Processor, Variable-Length, Parallel-Execution-Unit Architecture for FPGAs Using the RISC-V ISA
abstract
FPGA-based soft-processors have traditionally focused on fixed-pipeline designs. These designs have limited Instruction Level Parallelism (ILP) and constrain the integration of tightly-coupled accelerators, potentially limiting the speedup they can provide. Recently, it has been proposed that replacing the fixed-pipeline datapath in these soft processors with variable-latency parallel-execution functional units could facilitate the integration of custom instructions. In this paper, we discuss and analyze the architectural impact and requirements for decoupling the pipeline stages and supporting parallel execution units. We find that, relative to a fixed pipeline architecture, our variable-latency, parallel-execution architecture: increases resource usage by 8% LUTs and 9% FlipFlops but results in up to a 42% increase in Instruction Per Cycle (IPC), with an overall improvement of 28% MIPS/LUT. Finally, we analyze the performance tradeoffs of tightly integrating custom instructions into a fixed pipeline versus parallel execution units architecture.
Eric Matthews, Zavier Aguila, Lesley Shannon
FCCM3
2018 Modular Block-RAM-Based Longest-Prefix Match Ternary Content-Addressable Memories
abstract
Ternary Content Addressable Memories (TCAMs) are massively parallel search engines enabling the usage of "don't care" wildcards when searching for data. TCAMs are used in a wide variety of applications, such as routing tables for IP forwarding, which have been recently implemented using FPGAs. However, traditional "brute force" CAM architectures that use FPGA SRAM blocks (BRAMs) involve swapping address and data lines and are very inefficient. In this paper, a novel, efficient and modular technique for Longest-Prefix Match (LPM) TCAMs using FPGA BRAMs is proposed. Hierarchical search is exploited to achieve a linear storage growth and high storage efficiency. Compared to other methods, our LPM-TCAM design accommodates 5.5x more data for the same SRAM area without degrading the performance. A fully parameterized Verilog implementation is being released as an open source library. The library has been extensively tested using Altera's Quartus and ModelSim.
Ameer Abdelhadi, Guy Lemieux, Lesley Shannon
FPL3
2017 Performance impacts and limitations of hardware memory access trace collection
abstract
In today's multicore architectures, complex interactions between applications in the memory system can have a significant and highly variable impact on application execution time. System designers typically use hardware counters to profile execution behaviours and diagnose performance problems. However, hardware counters are not always sufficient and some problems are best identified with full memory access traces. Collecting these traces in software is very expensive; our work explores using dedicated hardware for memory-access trace collection. We analyze the limitations of this approach and its impacts on application performance. Our study is performed on actual hardware using two very different CPU platforms: 1) the PolyBlaze multicore soft processor and 2) the ARM Cortex-A9. In both cases, the data collection is implemented on an FPGA. Using micro-benchmarks designed to test the bounds of memory access behaviour, we illustrate the operational regions of data collection and the impact on system performance. By examining the bandwidth bottlenecks that limit the rate of data collection, as well as hardware architecture choices that can aggravate the impact on application performance, we provide guidelines that can be used to extrapolate our analysis to other systems and processor architectures.
Nicholas C. Doyle, Eric Matthews, Graham M. Holland, Alexandra Fedorova, Lesley Shannon
DATE5
2017 TAIGA: A new RISC-V soft-processor framework enabling high performance CPU architectural features
abstract
Recently, there has been an increased focus on integration of reconfigurable fabric with modern processors. However, existing soft-processors are optimized to leverage older FPGA fabrics, focus primarily on resource minimization and have fixed-pipeline designs that limit the scope for tightly integrated hardware accelerators. In this work, we present Taiga: a RISC-V, 32-bit, soft-processor architecture supporting the RISC-V Multiply/Divide and Atomic operations extensions (RV32IMA) designed to support Linux-based shared-memory systems. The processor design is highly configurable and features a standardized interface for functional units allowing for ease of integration of new functional units. Despite a more complex pipeline, our design uses approximately 33% fewer slices while clocking 39% faster than a LEON3 based system built on a Xilinx Zynq X7CZ020.
Eric Matthews, Lesley Shannon
FPL2
2016 A multi-beam Scan Mode Synthetic Aperture Radar processor suitable for satellite operation
abstract
As FPGA device sizes increase, they offer greater opportunities for on-site data processing, which is potentially useful for reducing the transmission requirements for applications with large data sets. For satellite applications, unlike ASICs, designers also benefit from an FPGA's ability to be reprogrammed to update functionality over the lifetime of a satellite (15+ years) and a mission (often 5+ years), while having significantly lower power costs than GPGPUs or high performance processors. This paper presents the first custom, fully pipelined, adaptable framework for multi-beam Scan Mode Synthetic Aperture Radar (SAR), the only 24/7 remote sensing imaging system that is capable of producing high-resolution global images in any weather conditions. As high resolution SAR or even low-resolution global-coverage generates on the order of hundreds of Megabytes of raw data per second, onboard SAR processing would reduce this transmitted data by orders of magnitude. Our Scan-mode SAR processor is scalable to different bit-widths and frame-sizes. We are able to process 81×730 frames of 8-bit I-Q channels from each scan (1.4 MB) in less than 1.5 ms, approximately 102 times faster than a corresponding estimated C solution leveraging the Intel Integrated Performance Primitives, and 150 times faster than the corresponding fully vectorized software solution run in MATLAB (6-core, 3.5 Ghz CPU).
Mohammad Reza Mohammadnia, Lesley Shannon
ASAP2
2016 Shared Memory Multicore MicroBlaze System with SMP Linux Support
abstract
In this work, we present PolyBlaze, a scalable and configurable multicore platform for FPGA-based embedded systems and systems research. PolyBlaze is an extension of the MicroBlaze soft processor, leveraging the configurability of the MicroBlaze and bringing it into the multicore era with Linux Symmetric Multi-Processor (SMP) support. This work details the hardware modifications required for the MicroBlaze processor and its software stack to enable fully validated SMP operations, including atomic operation support, shared interrupts and timers, and exception handling. New in this work, we present a scalable and flexible memory hierarchy optimized for Field Programmable Gate Arrays (FPGAs), which manages atomic operations and provides support for future flexible memory hierarchies and heterogeneous systems. Also new is an in-depth analysis of key performance characteristics, including memory bandwidth, latency, and resource usage. For all system configurations, bandwidth is found to scale linearly with the addition of processor cores until the memory interface is saturated. Additionally, average memory latency remains constant until the memory interface is saturated; after which, it scales linearly with each additional processor core.
Eric Matthews, Lesley Shannon, Alexandra Fedorova
ACM Trans. Reconfigurable Technol. Syst.2
2015 Technology Scaling in FPGAs: Trends in Applications and Architectures
abstract
Since the release of the first commercial field programmable gate array (FPGA) in 1985, devices have enjoyed continuous improvements in all metrics due to technology scaling, architectural advances and the addition of features. In this paper, we explore performance and utilization trends associated with research designs as a function of FPGA technology progression. The data used is a subset of designs presented at the IEEE International Symposium on Field-Programmable Custom Computing Machines (FCCM) over the past 20 years. These are compared to trends from theoretical and vendor sources, models generated and comparisons made. Finally, we compare operating frequency trends from our analysis to the trends exhibited by a set of vendor IP cores mapped to four generations of devices. The results of this investigation suggest that design implementations are generally following the theoretical trends and that the inclusion of embedded hard IP blocks has provided designers with additional performance benefits.
Lesley Shannon, Veronica Cojocaru, Cong Nguyen Dao, Philip H. W. Leong
FCCM1
2015 Design Space Exploration of L1 Data Caches for FPGA-Based Multiprocessor Systems
abstract
Combining multi-processing with the high level of configurability possible with FPGA-based soft-processors, this paper presents a multiprocessing framework based on the MicroBlaze soft-processor that provides multicore support and fully coherent, independently configurable Level 1 Caches with Linux multicore support. This architecture allows for fine-grain configurability of the system, allowing for FPGA resources to be better optimized for a specific embedded application. We use our framework to explore the L1 Data Cache configuration, developing a metric for efficiency based on resource usage and static application runtime. We find that a Pseudo-Random replacement policy is consistently the more efficient choice for FPGA systems.
Eric Matthews, Nicholas C. Doyle, Lesley Shannon
FPGA3
2014 A methodology for identifying and placing heterogeneous cluster groups based on placement proximity data (abstract only)
abstract
Due to the rapid growth in the size of designs and Field Programmable Gate Arrays (FPGAs), CAD run-time has increased dramatically. Reducing FPGA design compilation times without degrading circuit performance is crucial. In this work, we describe a novel approach for incremental design flows that both identifies tightly grouped FPGA logic blocks and then uses this information during circuit placement. Our approach reduces placement run-time on average by more than 17% while typically maintaining the design's critical path delay and marginally increasing its minimum channel width and wire length on average. Instead of following the traditional approach of evaluating a circuit's pre-placement netlist, this new algorithm analyzes designs post-placement to detect proximity data. It uses this information to non-aggressively extract heterogeneous cluster groupings from the design, which we call "gems," that consist of two to seventeen clusters. We modified VPR's simulated annealing placement algorithm to use our Singularity Placer, which first crushes each cluster grouping into a "singularity," to be treated as a single cluster. We then run the annealer over this condensed circuit, followed by an expansion of the singularities, and a second annealing phase for the entire expanded circuit.
Farnaz Gharibian, Lesley Shannon, Peter Jamieson
FPGA2
2014 Identifying and placing heterogeneously-sized cluster groupings based on FPGA placement data
abstract
Field Programmable Gate Arrays (FPGAs) CAD flow run-time has increased due to the rapid growth in size of designs and FPGAs. Researchers are trying to find new ways to improve compilation time without degrading design performance. In this paper, we present a novel approach that identifies tightly grouped FPGA logic blocks and then uses this information during circuit placement. Our approach is an orthogonal optimization applicable in incremental design and physical optimization, and reduces placement run-time. Specifically, we present a new algorithm that analyzes designs post-placement to extract medium-grained super-clusters that consist of two to seventeen clusters, which we call “gems”. We modified VPR's simulated annealing placement algorithm to place our mixture of gems and clusters. Our new “Singularity Annealing” algorithm first crushes each cluster grouping into a “singularity” (treated as a single cluster). Then, the Singularity Annealer is run over this condensed circuit to obtain an initial placement, followed by an expansion of the singularities. Finally, we run a second low-temperature annealing phase on the entire expanded circuit. Our results show that our system reduces placement run-time on average by 17% while maintains the designs critical path delay, and increases designs channel width, and wirelength by 2% and 6.3%, respectively. We have also presented a test case to show the re-usability of gems in an incremental design example.
Farnaz Gharibian, Lesley Shannon, Peter Jamieson
FPL2
2013 Supergenes in a genetic algorithm for heterogeneous FPGA placement
abstract
Supergenes are an addition to a genetic algorithm's genome that duplicate genes in the genome, represent local optimizations, and have the potential to be expressed overriding the duplicated gene. We introduce supergenes in a genetic algorithm for FPGA placement where a placement algorithm places a mix of fine-grain components and medium-grain components (where a medium-grain component is 2 to 10 times the size of a finegrain component). This is the first placement algorithm, to our knowledge, that can deal with such a mix of components. Our results show that supergenes improve a placement metric (clock speed of the FPGA) by approximately 10%. We also show and explore mutation operators on supergenes, and we experimentally demonstrate that the expression of a supergene can be effectively controlled via a binary function for our placement problem.
Peter Jamieson, Farnaz Gharibian, Lesley Shannon
IEEE Congress on Evolutionary Computation3
2013 Analyzing System-Level Information's Correlation to FPGA Placement
abstract
One popular placement algorithms for Field-Programmable Gate Arrays (FPGAs) is called Simulated Annealing (SA). This algorithm tries to create a good quality placement from a flattened design that no longer contains any high-level information related to the original design hierarchy. Placement is an NP-hard problem, and as the size and complexity of designs implemented on FPGAs increases, SA does not scale well to find good solutions in a timely fashion. In this article, we investigate if system-level information can be reconstructed from a flattened netlist and evaluate how that information is realized in terms of its locality in the final placement. If there is a strong relationship between good quality placements and system-level information, then it may be possible to divide a large design into smaller components and improve the time needed to create a good quality placement. Our preliminary results suggest that the locality property of the information embedded in the system-level HDL structure (i.e. “module”, “always”, and “if” statements) is greatly affected by designer HDL coding style. Therefore, a reconstructive algorithm, called Affinity Propagation, is also considered as a possible method of generating a meaningful coarse-grain picture of the design.
Farnaz Gharibian, Lesley Shannon, Peter Jamieson, Kevin Chung
ACM Trans. Reconfigurable Technol. Syst.2
2012 Bio-inspired walking: A FPGA multicore system for a legged robot
abstract
Previous legged robots use single or multi-microcontroller systems to control their motions. This work is a complete robot control system, implemented as a Multi-Processor System-on-Chip (MPSoC), on a Spartan 3A Field Programmable Gate Array (FPGA). Novel features of this system include encapsulation of the various levels of control, low communication latency between processors (4 clock cycles at 50 MHz), and ease of use for the control system researchers. The MPSoC implementation combines the performance benefits of processing control loops in parallel, with the size and mass advantages of a single IC solution. The system comprises one soft processor that is used for high-level decisions regarding the robot's overall movements, and six soft processors that run independent, low-level control loops for each of the six legs. The low-level control loop frequency can reach up to 2 kHz, and is only limited by the Analog to Digital Converter (ADC) sample rate. Coordination between legs occurs at 100 Hz. This design uses 90% of the user I/Os, 57% of the flip flops, 70% of the LUTs, 18% of the DSPs and 89% of the block RAMs on the FPGA, with a system operating frequency of 50 MHz. A legged robot, Abigaille-III, uses this control system to walk on flat and uneven surfaces.
Michael Henrey, Sean Edmond, Lesley Shannon, Carlo Menon
FPL3
2012 Polyblaze: From one to many bringing the microblaze into the multicore era with Linux SMP support
abstract
Modern computing systems increasingly consist of multiple processor cores. From cell phones to datacenters, multicore computing has become the standard. At the same time, our understanding of the performance impact resource sharing has on these platforms is limited, and therefore, prevents these systems from being fully utilized. As the capacity of FPGAs has grown, they have become a viable method for emulating architecture designs as they offer increased performance and visibility into runtime behaviour compared to simulation. With future systems trending towards asymmetric and heterogeneous systems, and thus further increasing complexity, a framework that enables research in this area is highly desirable. In this work, we present PolyBlaze: a multicore Micro- Blaze based system with Linux Symmetric Multi-Processor (SMP) support on an FPGA. Starting with a single-core, Linux supported, MicroBlaze we detail the changes to the platform, both in hardware and software, required to bring Linux SMP support to the MicroBlaze. We then outline the series of tests performed on our platform to demonstrate both its stability (e.g. more than two weeks of up time) and scalability (up to eight cores on an FPGA, with resource usage increasing linearly with the number of cores).
Eric Matthews, Lesley Shannon, Alexandra Fedorova
FPL2
2012 Minimizing the error: A study of the implementation of an Integer Split-Radix FFT on an FPGA for medical imaging
abstract
Fixed-point arithmetic is used to provide faster and smaller implementations in many digital signal processing applications, including medical imaging, at the expense of decreased accuracy. In particular, when a Fast Fourier Transform (FFT)-Inverse Fast Fourier Transform (IFFT) pair are required as part of the calculation, the error introduced into the calculations can be significant. For some applications, such as Fourier Domain Optical Coherence Tomography (FD-OCT), this degradation is unacceptable. Our study shows that using a conventional fixed-point FFT-IFFT pair, such as Xilinx's FFT core, can produce an average 6-bit error for a 1024-point FFT using 12-bit input data in a 32-bit arithmetic system. The majority of the error is caused by quantization effects, particularly on the phase information of input signal. For this reason, in phase sensitive applications such as FD-OCT, the error dominates the fixed-point calculation: 78% in 16-bit and 51% in 32-bit systems. This work presents a parameterized 32 to 4096-point integer FFT implementation for FPGAs that uses a Split-Radix algorithm to reduce the number of multiplies and improve latency. Integer FFTs are perfectly reconstructible, with zero reconstruction error. Here, we specifically analyze a 1024-point Integer Split-Radix FFT (Int-SRFFT) and IFFT pair that perfectly reconstructs the original 12-bit input data using 22-bit arithmetic, compared to the average 6-bit error for a 1024-point FFT using 32-bit fixed-point arithmetic. The pipelined architecture of this design has a latency of 29.06us for a 1024-point FFT, and a throughput of more than 34 thousand 1024-point FFTs/second for a 22-bit datapath at an operating frequency of 274MHz. Although our Int-SRFFT is perfectly reconstructible, compared to Xilinx's fixed-point FFT, it has ~6% more flipflops, ~63% more LUTs, 5.3x more BRAMs, and a ~44% increase in latency. However, compared to a previous fixed-point SRFFT design on an FPGA, our throughput is 15.5x greater.
Mohammad Reza Mohammadnia, Lesley Shannon
FPT2
2012 Amoeba-Cache: Adaptive Blocks for Eliminating Waste in the Memory Hierarchy
abstract
The fixed geometries of current cache designs do not adapt to the working set requirements of modern applications, causing significant inefficiency. The short block lifetimes and moderate spatial locality exhibited by many applications result in only a few words in the block being touched prior to eviction. Unused words occupy between 17-80% of a 64K L1 cache and between 1%-79% of a 1MB private LLC. This effectively shrinks the cache size, increases miss rate, and wastes on-chip bandwidth. Scaling limitations of wires mean that unused-word transfers comprise a large fraction (11%) of on-chip cache hierarchy energy consumption. We propose Amoeba-Cache, a design that supports a variable number of cache blocks, each of a different granularity. Amoeba-Cache employs a novel organization that completely eliminates the tag array, treating the storage array as uniform and morph able between tags and data. This enables the cache to harvest space from unused words in blocks for additional tag storage, thereby supporting a variable number of tags (and correspondingly, blocks). Amoeba-Cache adjusts individual cache line granularities according to the spatial locality in the application. It adapts to the appropriate granularity both for different data objects in an application as well as for different phases of access to the same data. Overall, compared to a fixed granularity cache, the Amoeba-Cache reduces miss rate on average (geometric mean) by 18% at the L1 level and by 18% at the L2 level and reduces L1-L2 miss bandwidth by ≃46%. Correspondingly, Amoeba-Cache reduces on-chip memory hierarchy energy by as much as 36% (mcf) and improves performance by as much as 50% (art).
Snehasish Kumar, Hongzhou Zhao, Arrvindh Shriraman, Eric Matthews, Sandhya Dwarkadas, Lesley Shannon
MICRO6
2012 Hierarchical Benchmark Circuit Generation for FPGA Architecture Evaluation
abstract
We describe a stochastic circuit generator that can be used to automatically create benchmark circuits for use in FPGA architecture studies. The circuits consist of a hierarchy of interconnected modules, reflecting the structure of circuits designed using a system-on-chip design flow. Within each level of hierarchy, modules can be connected in a bus, star, or dataflow configuration. Our circuit generator is calibrated based on a careful study of existing system-on-chip circuits. We show that our benchmark circuits lead to more realistic architectural conclusions than circuits generated using previous generators.
Cindy Mark, Scott Y. L. Chin, Lesley Shannon, Steve Wilton
ACM Trans. Embed. Comput. Syst.3
2011 FUSE: Front-End User Framework for O/S Abstraction of Hardware Accelerators
abstract
SoCs can be implemented on a single FPGA, offering designers a unique opportunity for Embedded Systems. Instead of defining a fixed architecture early in the design process, the reconfigurable platform allows architectural redesign to meet the system's specific needs. However, the ability to instantiate new modules in the reconfigurable hardware provides a unique set of challenges for integration, particularly to the software (SW) designer. Specifically, the Operating System (OS) cannot automatically abstract these platform changes without redesign. In this paper, we present FUSE, a framework for HW accelerator abstraction that provides: 1) transparency to the SW designer at the application level, and 2) OS support for easy HW accelerator integration. We illustrate FUSE as an API for an embedded Linux OS with POSIX threads on Xilinx's Micro Blaze on a Virtex5. For three different applications and HW accelerators, we achieve performance speedups ranging from 6.4-37×.
Aws Ismail, Lesley Shannon
FCCM2
2011 Scalable, High Performance Fourier Domain Optical Coherence Tomography: Why FPGAs and Not GPGPUs
abstract
Fourier Domain Optical Coherence Tomography (FD-OCT) is an emerging biomedical imaging technology featuring ultra-high resolution and fast imaging speed. Due to the complexity of the FD-OCT algorithm, real time FD-OCT imaging demands high performance computing platforms. However, the scaling of real-time FD-OCT processing for increasing data acquisition rates and 3-dimensional (3D) imaging is quickly outpacing the performance of general purpose processors. Our research analyzes the scalability of accelerating FD-OCT processing on two potential implementation platforms: General Purpose Graphical Processing Units (GPGPUs) and Field Programmable Gate Arrays (FPGAs). We implemented a complete FD-OCT system using a NVIDIA GPGPU as co-processor, with a speed up of 6.9x over general purpose processors (GPPs). We also created a hardware processing engine using FPGAs with a speed up of 15.5x over GPPs for a single pipeline, which can be replicated to further increase performance. Our analysis of the performance and scalability for both platforms shows that, while GPGPUs offer an easy and low cost solution for accelerating FD-OCT, FPGAs are more likely to match the long term demands for real-time, 3D, FD-OCT imaging.
Jian Li 0004, Marinko Sarunic, Lesley Shannon
FCCM3
2011 Exploring FPGA technology mapping for fracturable LUT minimization
abstract
Modern commercial Field-Programmable Gate Array (FPGA) architectures contain look-up-tables (LUTs) that can be “fractured” into two smaller LUTs. The potential of packing two LUTs into a space that could accommodate only one in traditional architectures complicates technology mapping's LUT minimization objective. Previous works introduced edge-recovery techniques and the concept of LUT balancing, both of which produce mappings that pack into fewer fracturable LUTs. We combine these two ideas and evaluate their effectiveness for one commercial and four academic FPGA architectures, all of which contain fracturable LUTs. When used in conjunction, edge-recovery and LUT balancing yield a 8.9% to 16.2% reduction in fracturable LUT use, depending upon architectural constraints.
David Dickin, Lesley Shannon
FPT2
2011 Leveraging reconfigurability in the hardware/software codesign process
abstract
Current technology allows designers to implement complete embedded computing systems on a single FPGA. Using an FPGA as the implementation platform introduces greater flexibility into the design process and allows a new approach to embedded system design. Since there is no cost to reprogramming an FPGA, system performance can be measured on-chip in the runtime environment and the system's architecture can be altered based on an evaluation of the data to meet design requirements. In this article, we discuss a new hardware/software codesign methodology tailored to reconfigurable platforms and a design infrastructure created to incorporate on-chip design tools. This methodology utilizes the FPGA's reconfigurability during the design process to profile and verify system performance, thereby reducing system design time. Our current design infrastructure includes: a system specification tool, two on-chip profiling tools, and an on-chip system verification tool.
Lesley Shannon, Paul Chow
ACM Trans. Reconfigurable Technol. Syst.1
2010 Customizing controller instruction sets for application-specific architectures
abstract
Previous work has proposed the "Systems Integrating Modules with Predefined Physical Links" (SIMPPL) architectural framework as one possible method to shorten the design cycle by utilizing a light weight programmable controller (SIMPPL Controller) as the system-level interface. This paper presents a study of how much improvement in area, power, and performance can be achieved through the customization of the SIMPPL Controller's instruction set. Furthermore, we have created a tool to automatically generate the HDL for SIMPPL Controllers with a user specified instruction set. Our study on an FPGA platform has shown that using a customized SIMPPL Controller with a minimal instruction set results in: an area reduction of 42%, a performance increase of 16%, and a power reduction of 10%.
Jian Li 0004, David Dickin, Lesley Shannon
ASAP3
2010 Odin II - An Open-Source Verilog HDL Synthesis Tool for CAD Research
abstract
In this work, we present Odin II, a framework for Verilog Hardware Description Language (HDL) synthesis that allows researchers to investigate approaches/improvements to different phases of HDL elaboration that have not been previously possible. Odin II's output can be fed into traditional back-end flows for both FPGAs and ASICs so that these improvements can be better quantified. Whereas the original Odin [1] provided an open source synthesis tool, Odin II's synthesis framework offers significant improvements such as a unified environment for both front-end parsing and netlist flattening. Odin II also interfaces directly with VPR [2], a common academic FPGA CAD flow, allowing an architectural description of a target FPGA as an input to enable identification and mapping of design features to custom features. Furthermore, Odin II can also read the netlists from downstream CAD stages into its netlist data-structure to facilitate analysis. Odin II can be used for a wide range of experiments; in this paper, we show three specific instances of how Odin II can be used by ASIC and FPGA researchers for more than basic synthesis. Odin II is open source and released under the MIT License.
Peter Jamieson, Kenneth B. Kent, Farnaz Gharibian, Lesley Shannon
FCCM4
2010 Predicting the performance of application-specific NoCs implemented on FPGAs
abstract
Modern FPGAs are able to implement complex systems such as Systems-on-Chips (SoCs) and Networks-on-Chips (NoCs). Appropriate NoC topology choices for ASICs have been investigated and typically topologies that can be easily mapped to a two-dimensional fabric are used to reduce chip area and ensure electrical characteristics. However, for FPGAs, each device's size and routing fabric are fixed. Since these resources exist independent of use, the choice of topology is only limited by the performance of the NoC itself. In this work, we investigate how topology characteristics impact a NoC's performance on an FPGA. From this analysis, we have created an analytical model that describes the maximum operating frequency of a NoC as a function of the topology's network parameters. This model is in the form of a simple equation that is accurate to within 4.68% across a range of topologies, chip sizes, and device families. It demonstrates how an FPGA's prefabricated routing interconnect provides increased freedom in the selection of application-specific topologies. Furthermore, it can also be used by designers for topology design space exploration before implementation.
Lesley Shannon
FPGA2
2010 Finding System-Level Information and Analyzing Its Correlation to FPGA Placement
abstract
One of the more popular placement algorithms for Field Programmable Gate Arrays (FPGAs) is called Simulated Annealing (SA). This algorithm tries to create a good quality placement from a flattened design that no longer contains any high-level information related to the original design hierarchy. Unfortunately, placement is an NP-hard problem and as the size and complexity of designs implemented on FPGAs increases, SA does not scale well to find good solutions in a timely fashion. As modern FPGAs can be used to implement Systems- and Networks-on-Chip, designers are required to spend an increasing amount of time waiting for place and route tools to complete that is not being matched by an increase in the power of computing work stations. In this paper, we investigate if system-level information can be reconstructed from a flattened netlist and evaluate how that information is realized in terms of its locality in the final placement. If there is a strong relationship between good quality placements and system-level information, then it may be possible to divide a large design into smaller components and improve the time needed to create a good quality placement. Our preliminary results suggest that the locality property of the information embedded in the system-level HDL structure (i.e. “module”, “always”, and “if” statements) is greatly affected by both the designer and the design itself. A reconstructive algorithm, called affinity propagation, is also considered as a possible method of generating a meaningful coarse grain picture of the design.
Farnaz Gharibian, Lesley Shannon, Peter Jamieson
FPL2
2010 A configurable framework for investigating workload execution
abstract
Processor systems contain a limited number of hardware counters that provide some visibility for certain types of interactions, but do not support sophisticated analysis due to limited resources. By contrast, system software simulators provide multidimensional runtime data, but slowdown application execution, often resulting in an inaccurate picture of hardware/ software interactions. The ideal solution to this problem is to create a dedicated hardware unit to “watch” the processor for these types of behaviours. In this paper, we present a hardware framework that leverages an FPGA's reconfigurable fabric to investigate of workload execution behaviours on processors using a hardware-Based Analyzer for the Characterization of User Software (ABACUS). ABACUS is currently able to interface with the LEON3 processor using 1367 FFs, 1504 LUTs and 1 Block RAM on a Virtex 2Pro running at 144 MHz.
Eric Matthews, Lesley Shannon, Alexandra Fedorova
FPT2
2008 Extending the SIMPPL SoC architectural framework to support application-specific architectures on multi-FPGA platforms
abstract
Process technology has reduced in size such that it is possible to implement complete application-specific architectures as systems-on-chip (SoCs) using both application-specific integrated circuits (ASICs) and field programmable gate arrays (FPGAs). However, the reconfigurable nature of an FPGA results in lower logic density, such that large, complex applications require multi-FPGA implementation platforms. Although designing SoCs is challenging, SoC models such as systems integrating modules with predefined physical links (SIMPPL) exist to facilitate the design process. SIMPPL leverages defined physical interfaces and communication protocols to enable rapid system-level integration for application-specific architectures. This paper presents a ldquoSIMPPL repeaterrdquo that enables the SIMPPL SoC architectural framework to be used for systems spanning multiple FPGAs. The SIMPPL repeater abstracts inter-chip communication, allowing designers to treat a multi-FPGA platform as a single large reconfigurable fabric and focus on their application-specific architecture.
David Dickin, Lesley Shannon
ASAP2
2008 A multi-FPGA application-specific architecture for accelerating a floating point Fourier Integral Operator
abstract
Many complex systems require the use of floating point arithmetic that is exceedingly time consuming to perform on personal computers. However, floating point operators are also hardware resource intensive and require longer latencies than fixed point operators to complete. Due to the reduced logic density of FPGAs relative to ASICs, it is often only possible to accelerate a portion of a floating point application in hardware. This paper presents an application-specific architecture for the hardware acceleration of a complete Fourier Integral Operator (FIO) kernel used in seismic imaging on a multi-FPGA platform. The design utilizes several floating point computing elements (CEs) to calculate the FIO kernel in parallel stages on multiple FPGAs. A detailed study of floating point CEs, including a Fast Fourier Transform (FFT) CE, and a complete FIO prototype implementation on the BEE2 platform is described. The prototype implementation has a 12.4x increase in throughput over an optimized software implementation, and a predicted 15.8x increase in throughput on the BEE3 platform.
Lesley Shannon, Matthew J. Yedlin, Gary F. Margrave
ASAP2
2008 Facilitating Processor-Based DPR Systems for non-DPR Experts
abstract
Currently, only Xilinx field programmable gate arrays (FPGAs) support dynamic partial reconfiguration (DPR). While there is currently some computer aided design (CAD) tool support for ISE-based DPR designs, none exists for microprocessor-based designs created in EDK. Creating DPR systems with the limited tool support currently available for ISE-based systems is already a challenging and complex process for novice DPR designers. These difficulties are severely compounded for potential microprocessor-based designs requiring a significant learning curve for novice DPR designers before they can successfully create their first working DPR system. This paper presents preliminary work towards extending the automation in Xilinx®'s current DPR design flow to include microprocessor based systems. The objective is to abstract low level details for novice designers, allowing them to focus on learning how to improve the quality of their design as opposed to how to perform the necessary manual transformations to generate a preliminary functional design. A case study demonstrated that the learning curve required to implement a first working design could be reduced by more than a factor of 15 times by improving the current automation available for microprocessor-based EDK designs.
William A. Gruver, Dorian Sabaz, Lesley Shannon
FCCM4
2008 An on-chip testbed that emulates runtime traffic and reduces design verification time for FPGA designs
abstract
Field programmable gate arrays (FPGAs) are commonly used as an inexpensive and flexible implementation platform for system-on-chip (SoC) designs. Now that FPGAs are large enough to implement SoCs, the reprogrammable fabric allows a different approach to the design process where on-chip computer aided design (CAD) tools can leverage reconfigurability to reduce design time. Statistics on commercial SoC designs suggest that 50% or more of design time may be spent on testing and verification due to design complexity. In previous work, we have proposed the systems integrating modules with predefined physical links (SIMPPL) SoC architectural framework to improve the design process. The defined communication links and protocols have been used to reduce integration time by an order of magnitude. In this paper, we propose an on-chip testbed that leverages both SIMPPL and an FPGApsilas reconfigurability to enable onchip testing and verification in real time using run time traffic patterns to reduce design time. The proposed testbed requires 331 LUTs and 224 flipflops for the Transmitter and 31 LUTs and 30 flipflops for the receiver. This testbed is able to generate a variety of possible run time traffic patterns that may be used to verify the operation of the CE.
Wayne Chen, Lesley Shannon
FPT2
2008 A new flexible PR domain model to replace the fixed multi-PR region model for DPR systems
abstract
Currently Xilinxpsilas dynamic partial reconfiguration (DPR) model requires the size and number of partially reconfigurable (PR) regions to be fixed during the design phase. Only one PR module may be active per PR region at any given time. This paper presents a new DPR model that replaces the current multi-PR region model with a single partially reconfigurable domain (PR Domain). Multiple PR modules of different sizes are able to concurrently reside in the PR domain. PR modules may be loaded to any portion of the PR domain given that sufficient resources exist. Future PR modules can be designed with dimensions up to those of the PR domain. A case study with multiple PR modules of varied sizes is used to demonstrate the adaptability and flexibility of the PR domain model.
Dorian Sabaz, William A. Gruver, Lesley Shannon
FPT4
2007 A Multiprocessor System-on-Chip Implementation of a Laser-based Transparency Meter on an FPGA
abstract
Modern FPGAs are large enough to implement multi-processor systems-on-chip (MPSoCs). Commercial FPGA companies also provide system design tools that abstract sufficient low-level system details to allow non-FPGA experts to design these systems for new applications. The application presented herein was designed by photomask researchers to implement a new technique for measuring the transparency of bimetallic grayscale masks using an FPGA platform. Production of the bimetallic grayscale masks requires a direct-write laser system. Previously, system calibration was determined by writing large rectangles of varying transparency on a mask and then measuring them using a spectrometer. The proposed technique uses the same mask-writing system but adds photodiode sensors connected to a multiprocessor computing system implemented on an FPGA. The added sensors combined with the laser beam's smaller focal point allows the calibration rectangles to be up to 5000 times smaller than those required by the spectrometer. This allows for direct mask verification on a mum-sized scale. Furthermore, the MPSoC design on the FPGA is easily scalable to support an increased number of photodiodes for the future addition of a feedback approach to the project.
James Dykes, Paulman Chan, Glenn H. Chapman, Lesley Shannon
FPT4
2007 Routability of Network Topologies in FPGAs
abstract
A fundamental difference between application-specific integrated circuits (ASICs) and field-programmable gate arrays (FPGAs) is that the wires in ASICs are designed to match the requirements of a particular design. Conversely, in an FPGA, the area is fixed and the routing resources exist whether or not they are used. In this paper, we investigate how well several common network topologies map onto a modern FPGA routing fabric. Different multiprocessor network topologies with between 8 and 64 nodes are mapped to a single large FPGA. Except for the fully-connected networks, it is observed that the difference in logic resources used and routing overhead among these topologies is insignificant for the systems tested. Fully-connected networks up to about 22 nodes are also feasible on the same FPGA although the logic and routing utilization clearly grows much faster. The conclusion is that a modern FPGA fabric is very rich in resources and capable of supporting highly interconnected topologies. For systems with a modest number of nodes implemented on current large FPGAs, it is not necessary to use the connectivity-limited topologies typically used for networks-on-chip. Rather, direct point-to-point connections between all communicating nodes can be considered.
Manuel Saldaña, Lesley Shannon, Jia Shuo Yue, Sikang Bian, John Craig, Paul Chow
IEEE Trans. Very Large Scale Integr. Syst.2
2007 SIMPPL: An Adaptable SoC Framework Using a Programmable Controller IP Interface to Facilitate Design Reuse
abstract
As the complexity of designing system-on-chips increases, so does the need to abstract low-level design issues to improve designer productivity. The reuse of previously designed Intellectual Property (IP) modules is a common form of abstraction used to reduce design time. However, different applications typically use a variety of physical interfaces, communication protocols, and global system-level control for IP modules, which complicates design reuse. In this paper, we describe the SIMPPL system model and an abstraction for IP modules, called the computing element (CE), that facilitate the SoC design for both field-programmable gate array (FPGA) and application-specific integrated circuit (ASIC) platforms. The CE abstraction decouples the datapath and system-level communication from the application-specific control to promote design reuse by localizing control redesign of IP for new applications. The SIMPPL model facilitates multi-clock domain SoC designs and expedites system integration by defining the intermodule links and communication protocols
Lesley Shannon, Paul Chow
IEEE Trans. Very Large Scale Integr. Syst.1
2006 The routability of multiprocessor network topologies in FPGAs
abstract
A fundamental difference between ASICs and FPGAs is that wires in ASICs are designed such that they match the requirements of a particular design. Wire parameters such as length, width, layout and the number of wires can be varied to implement a desired circuit. Conversely, in an FPGA, area is fixed and routing resources exist whether or not they are used, so the goal becomes implementing a circuit within the limits of available resources. The architecture for existing routing structures in FPGAs has evolved over time to suit the requirements of large, localized digital circuits. However, FPGAs now have the capacity to host networks of such circuits, and system-level interconnection becomes a key element of the design process.Following a standard design flow and using commercial tools, we investigate how this fundamental difference in resource usage affects the mapping of various network topologies to a modern FPGA routing structure. By exploring the routability of different multiprocessor network topologies with 8, 16 and 32 nodes on a single FPGA, we show that the difference between resource utilization of a ring, star, hypercube and mesh topologies is not significant up to 32 nodes. We also show that a fully-connected network can be implemented with at least 16 nodes, but with 32 nodes it exceeds the routing resources available on the FPGA. We also derive a cost metric that helps to estimate the impact of the topology selection based on the number of nodes.
Manuel Saldaña, Lesley Shannon, Paul Chow
FPGA2
2006 A System Design Methodology for Reducing System Integration Time and Facilitating Modular Design Verification
abstract
This paper provides a realistic case study of using the previously introduced SIMPPL system architectural model, which fixes the physical interface and communication protocols between processing elements (PEs) using PE-specific SIMPPL controllers. The implementation of a real-time MPEG-1 video decoder using SIMPPL provides a practical demonstration of how the complexity of system-level design issues are reduced by enabling rapid system-level integration and on-chip verification. The adaptation of the MPEG-1 PEs into the SIMPPL framework combined with the system-level integration was accomplished in 72.5 hours, which is only 4.5% of the overall system design time, instead of the more typical system integration times that can be as much as 30% of the design time
Lesley Shannon, Blair Fort, Samir Parikh, Arun Patel, Manuel Saldaña, Paul Chow
FPL1
2005 Simplifying the Integration of Processing Elements in Computing Systems Using a Programmable Controller
abstract
As technology sizes decrease and die area increases, designers are creating increasingly complex computing systems using FPGAs. To reduce design time for new products, the reuse of previously designed intellectual property (IP) cores is essential. However, since no universally accepted interface standards exist for IP cores, there is often a certain amount of redesign necessary before they are incorporated into the new system. Furthermore, the core's functionality may need updating to support the requirements of the new application. This paper demonstrates how the SIMPPL system model allows designers to rapidly implement on-chip systems comprising multiple computing elements (CEs). Furthermore, using a controller-based interface to manage inter-CE transfers enables users to easily adapt the control sequence of individual CEs to suit the needs of new applications without necessitating the redesign of other elements in the system. Two systems using three different hardware modules adapted to CEs are described to illustrate the power and simplicity of the SIMPPL model. It required a total of six hours to implement both designs on-chip once the individual CEs had been designed.
Lesley Shannon, Paul Chow
FCCM1
2005 Leveraging Reconfigurability in the Design Process
abstract
We are investigating an on-chip design methodology for embedded computing systems implemented on FPGAs. The objective is to exploit the platform's reconfigurability, such that system development can occur on the final implementation platform. This is comparable to the software development process where applications are typically designed on a workstation that is representative of the final product's platform, not a simulator that models the processor. Designing on the target technology is also appealing for FPGA designs. An on-chip design methodology better leverages the main advantage of reconfigurability - the user may redesign the system while avoiding non-recurring costs, such as mask redesign costs. It also allows the user to quickly obtain real-time information about system performance.
Lesley Shannon, Paul Chow
FPL1
2005 Designing an FPGA SoC Using a Standardized IP Block Interface
Lesley Shannon, Blair Fort, Samir Parikh, Arun Patel, Manuel Saldaña, Paul Chow
FPT1
2004 Using reconfigurability to achieve real-time profiling for hardware/software codesign
abstract
Embedded systems combine a processor with dedicated logic to meet design specifications at a reasonable cost. The attempt to amalgamate two distinct design environments introduces many problems, one being how to partition a single design for the two platforms to achieve the best performance with the least effort. Since the latest FPGA technology allows the integration of soft or hard CPU cores with dedicated logic on a single chip, this presents new opportunities for addressing hardware/software codesign issues in the FPGA design process by utilizing the reconfigurable environment.This paper introduces SnoopP, a non-intrusive, real time, profiling tool. The user is able to obtain a clock cycle accurate profile of the real time performance of a software program running on a soft-core processor instantiated on an FPGA. SnoopP is an essential tool for hardware/software codesign on a reconfigurable platform. It allows the user to quickly obtain accurate profiling information that may greatly influence the partitioning of the design.
Lesley Shannon, Paul Chow
FPGA1
2004 Maximizing system performance: using reconfigurability to monitor system communications
abstract
Commercial FPGA companies now provide tools that allow users to implement designs comprising soft-core processors and modules of dedicated logic. If a designer chooses to partition a system into multiple processors and hardware modules, tools and techniques for design analysis are necessary to understand system performance. This work introduces WOoDSTOCK, a tool that profiles system performance by adding monitors to the circuit running in real time on the chip. The user is able to generate a system specific profiler tailored to monitor the communication links between the different computing elements. This provides a macroscopic picture of system performance, which highlights the computing elements that cause bottlenecks in the design.
Lesley Shannon, Paul Chow
FPT1
2003 Standardizing the Performance Assessment of Reconfigurable Processor Architectures
abstract
This paper presents the Reconfigurable Architecture TEsting Suite, or RATES, which defines a standard for describing and using benchmarks for reconfigurable architectures. RATES is a set of functional benchmarks, is totally independent from the architecture and language, and usable on any processing platform be it general purpose or reconfigurable. It requires standard algorithms to allow comparisons amongst architectures but allows custom algorithms to highlight specific features.
Lesley Shannon, Paul Chow
FCCM1