Vaughn Betz

dblp:76/1530 · DBLP profile ↗
← Back
120ranked-venue papers
8as first author
31since 2021 · last 2026
0000-0003-0528-6493ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 118 · 8 first-author · 30 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 The Optimal, The Fast, and The Hybrid: Automatic Placement and Routing for AIE Arrays
abstract
Most of the widely deployed deep learning (DL) workloads, such as large language models and convolutional neural networks, are relatively regular compute graphs that exhibit a high degree of compute parallelism. Therefore, they are a natural fit for spatial dataflow accelerator architectures that map computations to an array of many compute cores communicating via shared memory buffers and/or some form of flexible interconnect between them, such as circuit or packet switched networks-on-chip (NoCs). AMD’s adaptive intelligent engine (AIE) arrays in both the Versal FPGAs and Ryzen NPU devices are exemplars of such architectures. Despite their high peak performance, efficiently mapping workloads to these architectures to maximize compute utilization is a challenging task. The application’s compute kernels are first partitioned into logical cores that are then placed at specific physical core locations. Finally, the different types of inter-core communication resources are configured to realize efficient data movement between cores. Each of these steps is a complex optimization problem that determines the ability to find a feasible mapping and directly impacts performance results. The current programming model for AIE arrays relies on manual placement or uses greedy 2D tiling algorithms. These approaches either require significant designer effort or work only for regular 2D-structured computations, but produce poor-quality or unroutable solutions for other cases. To this end, this work presents a versatile automatic placement and routing (PnR) framework for AIE arrays. We evaluate a variety of placement algorithms that guarantee optimality or trade optimality for scalability. We also formulate routing as a modified multi-commodity flow problem that is solved using mixed-integer linear programming. To demonstrate our PnR framework, we integrate it into AMD’s open-source MLIR-AIE toolchain and develop an entire benchmark suite using their programming API for end-to-end performance evaluation. Across 202 synthetic and real-world benchmarks, our PnR framework finds a legal mapping for 200 out of 202 benchmarks (a 99% success rate) compared to the 62% success rate of AMD’s greedy sequential placer in the MLIR-AIE toolchain. It also reduces routing resource usage by 15% compared to manual placement followed by AMD’s router. On-device end-to-end runtime measurements show that our PnR produces solutions that have a 30% speedup over AMD’s placer and are only 7% slower than expert manual placement.
Hang Yan 0014, James Yen, Rongbo Zhang, Andrew Boutros, Vaughn Betz
FCCM5
2025 From Errors to Solutions: LLM-Powered Command Scripting for FPGA Cad Tools
abstract
Computer-aided design (CAD) tools provide hundreds or even thousands of options that control various optimizations throughout the design flow. While this flexibility is powerful, it requires significant experience to be familiar with those options and effectively utilize them. For example, when a design fails, in many cases errors can be resolved by adjusting the CAD tool options rather than modifying the design itself. In this work, we propose VPR-LLM, a tool that utilizes Large Language Models (LLMs) to automate error resolution in the open-source FPGA CAD tool Verilog-to-Routing (VTR) by modifying the command-line options used to run the tool. VPRLLM parses error logs, VTR help messages, and documentation, then utilizes an LLM to generate modified command-line options that resolve the issue. VPR-LLM supports various LLM models and prompting techniques. All these models and techniques are evaluated and compared in terms of efficiency and cost. To evaluate our method, we proposed a dataset of 26 VTR run failures spanning five distinct error categories. The proposed technique successfully resolved 80 % of the cases without requiring any fine-tuning to the LLM model, demonstrating the effectiveness of VPR-LLM. This work represents an initial step toward AI-assisted debugging in CAD flows, where LLMs can enhance productivity by automatically identifying and correcting tool configurations.
Mohamed A. Elgammal, Vaughn Betz
FPL2
2025 Double Duty: FPGA Architecture to Enable Concurrent LUT and Adder Chain Usage
abstract
Flexibility and customization are key strengths of Field-Programmable Gate Arrays (FPGAs) when compared to other computing devices. For instance, FPGAs can efficiently implement arbitrary-precision arithmetic operations, and can perform aggressive synthesis optimizations to eliminate ineffectual operations. Motivated by sparsity and mixed-precision in deep neural networks (DNNs), we investigate how to optimize the current logic block architecture to increase its arithmetic density. We find that modern FPGA logic block architectures prevent the independent use of adder chains, and instead only allow adder chain inputs to be fed by look-up table (LUT) outputs. This only allows one of the two primitives—either adders or LUTs—to be used independently in one logic element and prevents their concurrent use, hampering area optimizations. In this work, we propose the Double Duty logic block architecture to enable the concurrent use of the adders and LUTs within a logic element. Without adding expensive logic cluster inputs, we use 4 of the existing inputs to bypass the LUTs and connect directly to the adder chain inputs. We accurately model our changes at both the circuit and CAD levels using open-source FPGA development tools. Our experimental evaluation on a Stratix-10-like architecture demonstrates area reductions of 21.6% on adder-intensive circuits from the Kratos benchmarks, and 9.3% and 8.2% on the more general Koios and VTR benchmarks respectively. These area improvements come without an impact to critical path delay, demonstrating that higher density is feasible on modern FPGA architectures by adding more flexibility in how the adder chain is used. Averaged across all circuits from our three evaluated benchmark set, our Double Duty FPGA architecture improves area-delay product by 9.7%.
Junius Pun, Xilai Dai, Grace Zgheib, Mahesh A. Iyer, Andrew Boutros, Vaughn Betz, Mohamed S. Abdelfattah
FPL6
2025 Field-Programmable Gate Array Architecture for Deep Learning: Survey and Future Directions
abstract
Deep learning (DL) is becoming the cornerstone of numerous applications both in large-scale datacenters and at the edge. Specialized hardware is often necessary to meet the performance requirements of state-of-the-art DL models, but the rapid pace of change in DL models and the wide variety of systems integrating DL make it impossible to create custom computer chips for all but the largest markets. Field-programmable gate arrays (FPGAs) present a unique blend of reprogrammability and direct hardware execution that make them suitable for accelerating DL inference. They offer the ability to customize processing pipelines and memory hierarchies to achieve lower latency and higher energy efficiency compared to general-purpose central processing units (CPUs) and graphics processing units (GPUs), at a fraction of the development time and cost of custom chips. Their diverse and high-speed inputs/outputs (IOs) also enable directly interfacing the FPGA to the network and/or a variety of external sensors, making them suitable for both datacenter and edge use cases. As DL has become an ever more important workload, FPGA architectures are evolving to enable higher DL performance. In this article, we survey both academic and industrialFPGA chip architectureenhancements for DL. First, we give a brief introduction on the basics of FPGA architecture and how its components lead to strengths and weaknesses for DL applications. Next, we discuss differentdesign stylesof DL inference accelerators implemented on FPGAs that achieve state-of-the-art performance and productive development flows, ranging from model-specific dataflow styles to software-programmable overlay styles. We survey DL-specific enhancements to traditional FPGA building blocks including the logic blocks (LBs), arithmetic circuitry, and on-chip memories, as well as new DL-specialized blocks that integrate into the FPGA fabric to accelerate tensor computations. Finally, we discuss hybrid devices that combine processors and coarse-grained accelerator blocks with FPGA-like interconnect and networks-on-chip (NoCs), and highlight promising future research directions.
Andrew Boutros, Aman Arora 0001, Vaughn Betz
Proc. IEEE3
2025 Editorial: A Message from the New Editor-in-Chief
abstract
No abstract available.
Vaughn Betz
ACM Trans. Reconfigurable Technol. Syst.1
2025 VTR 9: Open-Source CAD for Fabric and Beyond FPGA Architecture Exploration
abstract
This work details the capabilities of a major new release of the Verilog-to-Routing (VTR) open source FPGA CAD tool flow. Enhancements include generalizations of VTR’s architecture modeling language and optimizers to enable a more diverse set of programmable routing fabrics, FPGAs with embedded hard Networks-on-Chip (NoCs) and three-dimensional 3D FPGA systems that leverage stacked silicon integration. The new Parmys logic synthesis flow improves language coverage and result quality, and the physical implementation flow includes a more efficient placement engine, floorplanning constraints to guide placement, the ability to perform single-stage (flat) routing to improve quality, and parallel routing algorithms to reduce CPU time. This release also includes new architecture captures of recent commercial devices (Xilinx’s 7-series and Altera’s Stratix 10) and new benchmark suites (Titanium25 and Hermes) to aid FPGA architecture investigation. Verilog language coverage is greatly improved with the new Parmys logic synthesis flow, enabling more designs to be used with VTR. Finally, the placement and routing engines have beeenbeen sped up by 4 \(\times\) and 2.2 \(\times\) vs. VTR 8, respectively, leading to an overall physical implementation flow CPU time reduction of 48% with better result quality on average compared to VTR 8.
Mohamed A. Elgammal, Amin Mohaghegh, Soheil Gholami Shahrouz, Fatemehsadat Mahmoudi, Fahrican Kosar, Kimia Talaei, Joshua Fife, Daniel Khadivi, Kevin E. Murray, Andrew Boutros, Kenneth B. Kent, Jeffrey B. Goeders, Vaughn Betz
ACM Trans. Reconfigurable Technol. Syst.13
2025 Corrigendum: VTR 9: Open-Source CAD for Fabric and Beyond FPGA Architecture Exploration
abstract
This is a corrigendum for the article “VTR 9: Open-Source CAD for Fabric and Beyond FPGA Architecture Exploration” published in ACM Trans. Reconfig. Technol. Syst. 18, 3, Article 39 (August 2025), 53 pages.
Mohamed A. Elgammal, Amin Mohaghegh, Soheil Gholami Shahrouz, Fatemehsadat Mahmoudi, Fahrican Kosar, Kimia Talaei, Joshua Fife, Daniel Khadivi, Kevin E. Murray, Andrew Boutros, Kenneth B. Kent, Jeffrey B. Goeders, Vaughn Betz
ACM Trans. Reconfigurable Technol. Syst.13
2024 Stay Flexible: A High-Performance FPGA NPU Overlay for Graph Neural Networks
abstract
Graph neural networks (GNNs) are a class of deep learning (DL) models widely-used for learning latent representations of graph-structured data for a variety of node/graph-level prediction tasks. Real-time applications of GNNs are evolving in various domains such as 3D object detection from LiDAR point clouds in autonomous vehicles [1] and classifying collected data in particle physics colliders [2]. Typically, these use cases have stringent latency constraints but can still benefit from batch processing of multiple graphs from different input sources. Existing accelerators either rely on preprocessing input graphs [3], [4] or are extremely specialized streaming pipelines which are unable to support dynamically changing workloads for these applications [5]. In this work, we take a different approach by enhancing the neural processing unit (NPU) [6] to accelerate a wide variety of GNN models without sacrificing its flexibility, performance or ability to run any of its originally supported DL workloads (e.g. MLPs, RNNs, GRUs, LSTMs).
Taikun Zhang, Andrew Boutros, Sergey Gribok, Kwadwo Boateng, Vaughn Betz
FCCM5
2024 H2PIPE: High Throughput CNN Inference on FPGAs with High-Bandwidth Memory
abstract
Convolutional Neural Networks (CNNs) combine large amounts of parallelizable computation with frequent memory access. Field Programmable Gate Arrays (FPGAs) can achieve low latency and high throughput CNN inference by implementing dataflow accelerators that pipeline layer-specific hardware to implement an entire network. By implementing a different processing element for each CNN layer, these layer-pipelined accelerators can achieve high compute density, but having all layers processing in parallel requires high memory bandwidth. Traditionally this has been satisfied by storing all weights on chip, but this is infeasible for the largest CNNs, which are often those most in need of acceleration. In this work we augment a state-of-the-art dataflow accelerator (HPIPE) to leverage both High-Bandwidth Memory (HBM) and on-chip storage, enabling high performance layer-pipelined dataflow acceleration of large CNNs. Based on profiling results of HBM’s latency and throughput against expected address patterns, we develop an algorithm to choose which weight buffers should be moved off chip and how deep the on-chip FIFOs to HBM should be to minimize compute unit stalling. We integrate the new hardware generation within the HPIPE domain-specific CNN compiler and demonstrate good bandwidth efficiency against theoretical limits. Compared to the best prior work we obtain speed-ups of at least $19.4 \mathrm{x}, 5.1 \mathrm{x}$ and 10.5 x on ResNet-18, ResNet-50 and VGG-16 respectively.
Mario Doumet, Marius Stan, Mathew Hall, Vaughn Betz
FPL4
2024 Better Together: Combining Analytical and Annealing Methods for FPGA Placement
abstract
Placement is a critical step in the FPGA design implementation flow that strongly impacts routability and timing closure. Recent state-of-the-art academic analytical placers have achieved impressive scalability but are limited to AMD Ultrascale-like architectures and mostly synthetic designs. On the other hand, VPR, the place and route tool within the widely used open-source Verilog-to-Routing (VTR) toolchain, can produce a legal placement for any arbitrary architecture; however, its simulated annealing placer scales poorly. Thus, there is a clear need to bring scalable, high-quality placement to realistic architectures and circuits. In this work, we develop a hybrid framework that combines the strength of a scalable flat analytical placer with the flexibility of simulated annealing techniques to adapt to various architectures and circuits, substantially improving the quality of results. We augment the state-of-theart analytical elfPlace FPGA placer as aug-elfPlace, generalizing its architecture modeling to handle real-world constraints and target different and more complete architectures. We leverage VPR’s legalization capability to integrate with external placers such as aug-elfPlace. VPR’s simulated annealing placer can further optimize the legalized placement, and VPR’s router and timing analysis can provide final quality results. By integrating wirelength-driven aug-elfPlace and VPR, our hybrid framework achieves up to 2% timing improvement with 15% reduction in routed wirelength compared to timing-driven VPR, on average across the large and heterogeneous Titan23 benchmark suite targeting an Intel Stratix-IV-like architecture.
Rachel Selina Rajarathnam, Kate Thurmer, Vaughn Betz, Mahesh A. Iyer, David Z. Pan
FPL3
2024 The Road Less Traveled: Congestion-Aware NoC Placement and Packet Routing for FPGAs
abstract
To help scale to ever-larger and more complex designs, recent FPGA architectures now integrate network-on-chips (NoCs). NoCs help transfer high-bandwidth data over long distances within the chip without using scarce low-delay long routing wire segments. While NoC-enhanced FPGAs aid system integration and design reuse, they also complicate FPGA computer-aided design (CAD) flows by introducing new constraints and metrics. Placement and routing needs to optimize NoC metrics like latency and bandwidth utilization and avoid link oversubscription (congestion), while simultaneously optimizing the programmable routing resource usage of the design modules attached to NoC routers. In this work we develop several new approaches to reduce NoC congestion while minimizing the impact on other design metrics. First, we incorporate a NoC link congestion cost into the placement engine of the open-source CAD flow, versatile place & route (VPR). Second, we integrate turn model NoC routing algorithms into the placement engine to leverage path diversity to further reduce congestion. On average over a suite of 29 benchmarks combining placement congestion modeling with turn model packet routing reduces NoC congestion by 90.7% at the cost of increasing aggregate bandwidth demand by 4%. In cases where the enhanced placement engine and NoC routing fail to fully resolve congestion, we formulate NoC routing as a boolean satisfiability (SAT) problem. This approach yields significant additional improvements; the combined algorithm reduces congestion by 95.1% compared to the baseline placement. Finally, we enhance the reinforcement learning (RL) agent in VPR’s placement engine by introducing a NoC-aware move type, resulting in an 8.8% reduction in wirelength on designs that make extensive use of the NoC.
Soheil Gholami Shahrouz, Vaughn Betz
FPL2
2024 A Software-Programmable Neural Processing Unit for Graph Neural Network Inference on FPGAs
abstract
Graph neural networks (GNNs) are a widely-used class of deep learning (DL) models for learning latent representations of graph-structured data for a variety of node/graph-level prediction tasks, some of which require real-time low latency inference. Most existing GNN accelerators rely on preprocessing input graphs on a host/embedded CPU to parallelize computations on different sub-graphs, making them unsuitable for real-time use cases. Others are extremely specialized streaming pipelines for only a specific type of GNN and therefore suffer from long FPGA bitstream compile times when the model is updated and cannot be used in applications that combine GNNs with other classes of DL models. In this work, we enhance the neural processing unit (NPU) FPGA overlay architecture, instruction set, and software stack to support a variety of GNN models. We achieve this without sacrificing the NPU flexibility; our enhanced NPU can be programmed purely through software to accelerate different GNNs or any of its originally supported DL workloads (e.g. MLPs, RNNs, GRUs, LSTMs). In addition, this flexibility enables our NPU software compiler to generate GNN kernels with different performance targets (throughput-optimized vs. latency-optimized) by exploiting different dimensions of compute parallelism on the same overlay architecture. Besides the flexibility benefits, our NPU implemented on an Intel Stratix 10 NX (14 nm) FPGA can process $7.8 \times$ more graphs per second at a similar latency on average compared to a state-of-the-art model-specific FPGA accelerator targeting real-time applications on an AMD Ultrascale+ same-generation FPGA. It also achieves 5.8 $\times$ higher throughput compared to an Nvidia RTX A6000 GPU (8 nm) and $2.6 \times$ lower latency than a state-of-the-art accelerator that combines CPU-based graph preprocessing with AMD Versal (7 nm) fabric and AI engine compute. Finally, we present a case study for using our enhanced NPU in real-time GNNbased multi-input multi-output (MIMO) antenna scheduling, highlighting that it meets the latency requirements of this task in 5G communication networks.
Taikun Zhang, Andrew Boutros, Sergey Gribok, Kwadwo Boateng, Vaughn Betz
FPL5
2024 High Throughput FPGA-Based Object Detection via Algorithm-Hardware Co-Design
abstract
Object detection and classification is a key task in many computer vision applications such as smart surveillance and autonomous vehicles. Recent advances in deep learning have significantly improved the quality of results achieved by these systems, making them more accurate and reliable in complex environments. Modern object detection systems make use of lightweight convolutional neural networks (CNNs) for feature extraction, coupled with single-shot multi-box detectors (SSDs) that generate bounding boxes around the identified objects along with their classification confidence scores. Subsequently, a non-maximum suppression (NMS) module removes any redundant detection boxes from the final output. Typical NMS algorithms must wait for all box predictions to be generated by the SSD-based feature extractor before processing them. This sequential dependency between box predictions and NMS results in a significant latency overhead and degrades the overall system throughput, even if a high-performance CNN accelerator is used for the SSD feature extraction component. In this paper, we present a novel pipelined NMS algorithm that eliminates this sequential dependency and associated NMS latency overhead. We then use our novel NMS algorithm to implement an end-to-end fully pipelined FPGA system for low-latency SSD-MobileNet-V1 object detection. Our system, implemented on an Intel Stratix 10 FPGA, runs at 400 MHz and achieves a throughput of 2,167 frames per second with an end-to-end batch-1 latency of 2.13 ms. Our system achieves 5.3× higher throughput and 5× lower latency compared to the best prior FPGA-based solution with comparable accuracy.
Anupreetham Anupreetham, Mohamed Ibrahim 0005, Mathew Hall, Andrew Boutros, Ajay Kuzhively, Abinash Mohanty, Eriko Nurvitadhi, Vaughn Betz, Yu Cao 0001, Jae-sun Seo
ACM Trans. Reconfigurable Technol. Syst.8
2023 Placement Optimization for NoC-Enhanced FPGAs
abstract
Field-programmable gate array (FPGA) architectures have recently incorporated hardened networks-on-chip (NoCs) to enable more efficient and easier system-level integration. However, the embedding of hard NoCs presents a new challenge for FPGA computer-aided design (CAD); the tools need to optimize the placement of circuit netlist primitives to not only minimize total wirelength and critical path delay, but also consider the NoC traffic patterns between modules to minimize their aggregate bandwidth and/or meet latency constraints. This work enables flexible modeling of FPGA architectures with hard NoCs in the open-source versatile place & route (VPR) CAD flow, facilitating both CAD and architecture research. We enhance the placement engine in VPR to co-optimize traditional circuit implementation metrics (e.g. wirelength, critical path delay) and NoC performance metrics (e.g. congestion, bandwidth utilization, latency) when mapping an application design with NoC-attached modules to a candidate NoC-enhanced FPGA architecture. We test our VPR enhancements using a variety of synthetic benchmarks and verify that the placement engine can effectively optimize NoC aggregate bandwidth and meet specified latency constraints. Then, we present a complete flow that integrates VPR with a high-level SystemC architecture simulator, RAD-Sim, that can capture the NoC traffic flows of complete application designs and use it to drive VPR's placement optimizations. We showcase this combined flow using a real application design from the deep learning domain. The results show that our NoC-enhanced VPR flow can result in 2x reduction in NoC aggregate bandwidth (on average) compared to a NoC-agnostic flow, without affecting the design's wirelength or critical path delay.
Srivatsan Srinivasan, Andrew Boutros, Fatemehsadat Mahmoudi, Vaughn Betz
FCCM4
2023 Open-source and FPGAs: Hardware, Software, Both or None?
abstract
Following the footsteps of the open-source software movement that is at the foundation of many fundamental infrastructures today, e.g., Linux, the internet, etc., a growing amount of open-source hardware initiatives have been impacting our field, e.g., the RISC-V ISA, Open chiplet standards, etc.
Dana How, Tim Ansell, Vaughn Betz, Chris Lavin, Ted Speers, Pierre-Emmanuel Gaillardon
FPGA3
2023 A Whole New World: How to Architect Beyond-FPGA Reconfigurable Acceleration Devices?
abstract
Field-programmable gate arrays (FPGAs) have evolved beyond a fabric of soft logic and hard blocks surrounded by programmable routing to also incorporate high-performance networks-on-chip (NoCs), general-purpose processor cores and application-specific accelerators. These new reconfigurable acceleration devices (RADs) open up a myriad of architecture research questions, but require new computer-aided design tools for quantitative evaluation. In this work, we first enhance an existing RAD architecture simulator, RAD-Sim, to model high-bandwidth memory (HBM) and conventional DDR interfaces. We also introduce RAD-Gen which evaluates the silicon area and performance of the different components of a candidate RAD. We showcase the complete flow through a case study on accelerating deep learning recommendation models (DLRMs). Using RAD-Sim and RAD-Gen, we compare traditional FPGAs to RADs that incorporate hard NoCs and matrix-vector multiplication accelerators. This study demonstrates the utility of these tools in evaluating both the performance and implementation feasibility of different combinations of NoC and accelerator architecture parameters. The resulting RAD achieves an order of magnitude improvement in DLRM inference throughput and latency compared to prior FPGA implementations.
Andrew Boutros, Stephen More, Vaughn Betz
FPL3
2023 VPR-Gym: A Platform for Exploring AI Techniques in FPGA Placement Optimization
abstract
With the increasing complexity and capacity of modern Field-Programmable Gate Arrays (FPGAs), there is a growing demand for efficient FPGA computer-aided design (CAD) tools, particularly in the placement stage. While some previous works, such as RLPlace, have explored the efficacy of single-state Reinforcement Learning (RL) to optimize FPGA placement by framing it as a multi-armed bandit (MAB) problem, numerous AI techniques remain unexplored due to the outstanding engineering challenges to integrate them into the FPGA CAD flow which is based on C++. In this paper, we present VPR-Gym, a Python environment built on OpenAI Gym, that allows seamless integration with various machine learning libraries including PyTorch, TensorFlow, and Nevergrad while enabling the comparison between different AI techniques for FPGA placement. Moreover, we introduce a learning objective that reformulates the FPGA placement task as an optimization problem, thereby expanding the range of AI techniques that can be investigated beyond those for MAB problems. To showcase the capabilities of our platform, we conduct experiments comparing the performance of various MAB algorithms and evolution strategy (ES) algorithms. Our findings demonstrate that the ES approaches exhibit superior performance over the existing MAB approaches, highlighting the effectiveness of VPR-Gym in facilitating AI research to enhance FPGA placement.
Ruichen Chen, Shengyao Lu, Mohamed A. Elgammal, Peter Chun, Vaughn Betz, Di Niu 0002
FPL5
2023 Titan 2.0: Enabling Open-Source CAD Evaluation with a Modern Architecture Capture
abstract
This paper presents an updated version of the Titan (Quartus + VPR) flow, and a VPR-compatible architecture description of Intel's Stratix 10 device. Together these components enable large designs to be targeted at a complex and realistic architecture, facilitating improved benchmarking and optimization of open-source CAD flows. Additionally, the Stratix 10 architecture capture is a useful baseline architecture on which proposed new FPGA features can be evaluated. We capture the device primitives, intra-block connectivity, device floorplans, routing architecture and timing with reasonable fidelity in the human-readable VTR architecture format, enabling easy modification by researchers. Using the updated Titan flow and Stratix 10 capture, we compare the quality of results (QoR), runtime, and resource utilization of VPR to Quartus Prime. The results show that VPR's overall runtime is comparable and its memory footprint is smaller than that of Quartus, but the packer runtime is higher and Quartus achieves better wirelength and frequency. We identify causes of the slower packing runtime in VPR and suggest directions for future improvements.
Kimia Talaei Khoozani, Arash Ahmadian Dehkordi, Vaughn Betz
FPL3
2023 Tear Down The Wall: Unified and Efficient Intra-and Inter-Cluster Routing for FPGAs
abstract
Routing is one of the most time-consuming phases of the FPGA CAD flow, and its quality of result greatly impacts design speed and whether a successful implementation is achieved with the fixed wiring resources available. Routing is also strongly affected by the FPGA programmable interconnect architecture, making data-driven algorithms crucial to explore new fabrics. The VTR CAD flow has historically split the routing problem into two parts: within the logic cluster and between clusters. This allows a flexible specification of the programmable interconnect of each of these components and also simplifies the routing problem by splitting it into two sub-problems. However, for some architectures, this split can significantly reduce result quality, so in this work, we develop a single-stage router, run-flat, that can simultaneously optimize the intra-cluster and inter-cluster routing. We show that this new router increases the generality of architectures VTR can target and improves the quality of results at the cost of a run time increase. We detail router enhancements that reduce memory by removing non-essential nodes from the routing architecture, reduce run time by accurately ranking partial routing choices with an enhanced lookahead and efficiently resolve resource overuse at choke points with a new net-aware congestion model. Over the VTR benchmarks on an architecture with partial crossbars within the logic clusters, run-flat reduces the minimum channel width by 14% and critical path delay by 2% on average at the cost of 2.7 xthe router run time. On the Titan benchmark suite on a Stratix-IV-like architecture, run-flat reduces wirelength by 6% while taking 1.36 x the router run time.
Amin Mohaghegh, Vaughn Betz
FPL2
2023 Koios 2.0: Open-Source Deep Learning Benchmarks for FPGA Architecture and CAD Research
abstract
the prevalence of deep learning (DL) in many applications, researchers are investigating different ways of optimizing field-programmable gate array (FPGA) architecture and CAD to achieve better quality-of-results (QoRs) on DL-based workloads. In this optimization process, benchmark circuits are an essential component; the QoR achieved on a set of benchmarks is the main driver for architecture and CAD design choices. However, current academic benchmark suites are inadequate, as they do not capture any designs from the DL domain. This work presents the second version of our suite of DL acceleration benchmark circuits for FPGA architecture and CAD research, called Koios. This suite of 40 circuits covers a wide variety of accelerated neural networks, design sizes, implementation styles, abstraction levels, and numerical precisions. These benchmarks include 32 DL designs and eight synthetic (proxy) benchmarks. The Koios benchmarks are larger, more data parallel, more heterogeneous, more deeply pipelined, and utilize more FPGA architectural features compared to existing open-source benchmarks. This enables researchers to pinpoint architectural inefficiencies for this class of workloads and optimize CAD tools on more representative benchmarks that stress the CAD algorithms in different ways. In this article, we describe the Koios designs, compare their characteristics to prior FPGA benchmark suites, and present results of running them through the verilog-to-routing (VTR) flow using a recent FPGA architecture model. Finally, we present case studies showing how exploration of DL-optimized FPGA architecture and CAD algorithms can be performed using our new benchmark suite.
Aman Arora 0001, Andrew Boutros, Seyed Alireza Damghani, Karan Mathur, Vedant Mohanty, Tanmay Anand, Mohamed A. Elgammal, Kenneth B. Kent, Vaughn Betz, Lizy Kurian John
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.9
2023 Toward Software-like Debugging for FPGAs via Checkpointing and Transaction-based Co-Simulation
abstract
Checkpoint-based debugging flows have recently been developed that allow the user to move the design state back and forth between an FPGA and a simulator. They provide a softwarelike debugging experience by combining the speed of hardware execution and the full visibility of simulation. However, they assume the entire system state can be moved to a simulator, limiting them to self-contained systems. In this article, we present StateLink, a transaction-based co-simulation framework that allows part of the system (the task) to run in a simulator and still interact with other system components that reside in hardware. StateLink allows tasks to remain connected to and active in the overall hardware system after their state is moved to a simulator. This extends the functionality of checkpoint-based debugging frameworks to designs with external I/Os and significantly speeds up the simulation of tasks that are part of a large system. StateLink typically adds no timing overhead and a modest hardware area overhead. The total area overhead of using the proposed flow on a Memcached system is only 13%. This flow allows the user to benefit from both the hardware speedup of ∼1M× and the StateLink speedup of up to 44× versus full system simulation.
Sameh Attia, Vaughn Betz
ACM Trans. Reconfigurable Technol. Syst.2
2022 RAD-Sim: Rapid Architecture Exploration for Novel Reconfigurable Acceleration Devices
abstract
With the continued growth in field-programmable gate array (FPGA) capacity and their incorporation into new environments such as datacenters, we have witnessed the introduction of a new class of reconfigurable acceleration devices (RADs) that go beyond conventional FPGA architectures. These devices combine a reconfigurable fabric with coarse-grained domain-specialized accelerator blocks all connected via a high-performance packet-switched network-on-chip (NoC) for efficient system-wide communication. However, we lack the tools necessary to efficiently explore the huge design space for RADs, study the complex interactions between their different components and evaluate various combinations of design choices. In this work, we develop RAD-Sim, a cycle-level architecture simulator that allows rapid application-driven exploration of the design space of novel RADs. To showcase the capabilities of RAD-Sim, we map and simulate a state-of-the-art deep learning (DL) inference overlay on a RAD instance incorporating an FPGA fabric and a complex of hard matrix-vector multiplication engines, communicating over a system-wide NoC. Through this example, we show how RAD-Sim can help architects quantify the effect of changing specific architecture parameters on end-to-end application performance.
Andrew Boutros, Eriko Nurvitadhi, Vaughn Betz
FPL3
2022 Quality & Generality: A Flexible FPGA Re-Clustering Technique to Improve Packing and Placement
abstract
The Packing and Placement stages are two major steps in the FPGA backend flow which greatly affect the Quality-of-Results (QoR) of design implementation. While these problems have been extensively studied in the literature, most approaches have either sacrificed generality by targeting specific and simplified FPGAs with few “block packing” legality constraints, or sacrificed quality by making irreversible packing decisions early in the flow and hence constraining the optimizations available to the subsequent placement stage. In this paper, we propose a new (re-clustering API) that can be used to update the packed netlist at different points throughout the packing and placement stages. This API can be used in our proposed flow to improve the QoR while preserving the generality and flexibility of the flow and ensuring the legality of the solution for any proposed FPGA architecture.
Mohamed A. Elgammal, Vaughn Betz
FPT2
2022 HPIPE NX: Boosting CNN Inference Acceleration Performance with AI-Optimized FPGAs
abstract
With the ever-increasing compute demands of artificial intelligence (AI) workloads, there is extensive interest in leveraging field-programmable gate-arrays (FPGAs) to quickly deploy hardware accelerators for the latest convolutional neural networks (CNNs). Recent FPGA architectures are also evolving to better serve the needs of AI, but accelerators need extensive re-design to leverage these new features. The Stratix 10 NX chip by Intel is a new FPGA that replaces traditional DSP blocks with in-fabric AI tensor blocks that provide 15x more multipliers and up to 143 TOPS of performance, at the cost of lower precision (INT8) and significant restrictions on how many operands can be fed to the multipliers from the programmable routing. In this paper, we explore different CNN accelerator structures to leverage the tensor blocks, considering the various tensor block modes, operand bandwidth restrictions, and on-chip memory restrictions. We incorporate the most performant techniques into HPIPE, a layer-pipelined and sparse-aware CNN accelerator for FPGAs. We enhance HPIPE's software compiler to restructure the CNN computations and on-chip memory layout to take advantage of the additional multipliers offered by the new tensor block architecture, while also avoiding stalls due to data loading restrictions. We achieve cycle-by-cycle speedups in tensor mode of up to$\mathbf{8}.\mathbf{3}\mathbf{x}$for Mobilenet-v1 versus the original HPIPE design using conventional DSPs. On the FPGA, we achieve a throughput of 28,541 and 29,429 images/s on Mobilenet-v1 and Mobilenet-v2 respectively, outperforming all previous FPGA accelerators by at least 4.0x, including one on an AI-optimized Xilinx chip. We also outperform NVIDIA's V100 GPU, a machine learning targeted GPU on a similar process node with a$\mathbf{1}.\mathbf{7}\mathbf{x}$larger die size, by up to 17x with a batch size of one and 1.3x with NVIDIA's largest reported batch size of 128.
Marius Stan, Mathew Hall, Vaughn Betz
FPT4
2022 Stop and Look: A Novel Checkpointing and Debugging Flow for FPGAs
abstract
Hardware checkpointing enables live migration, fault recovery, and context switching, but has been difficult to achieve for FPGA applications. We detail techniques to checkpoint complex FPGA designs and develop StateMover, a new checkpoint-based debugging flow for FPGAs that combines the speed of hardware execution with the full observability and controllability of simulation. StateMover can safely stop a running design and seamlessly move its state back and forth between an FPGA and a simulator. StateMover can create complete design checkpoints even for designs that have multi-cycle I/O interfaces, contain buried state that is not accessible by FPGA readback, and use external memories. StateMover and its associated IPs allow a designer to quickly make a design checkpointable, with a small area overhead. Moving the state from/to an FPGA to/from a simulator can be performed in a few seconds for large Xilinx UltraScale FPGAs.
Sameh Attia, Vaughn Betz
IEEE Trans. Computers2
2022 RLPlace: Using Reinforcement Learning and Smart Perturbations to Optimize FPGA Placement
abstract
Simulated annealing (SA) is one of the most common FPGA placement techniques, and is used both as a standalone algorithm and to improve an initial analytical placement. While SA-based placers can achieve high-quality results, they suffer from long runtimes. In this article, we introduceRLPlace, a novel SA-based FPGA placer that utilizes both reinforcement learning (RL) and targeted perturbations (directed moves). The proposed moves target both wirelength and timing optimization and explore the solution space more efficiently than traditional random moves while preventing oscillation in the Quality of Results (QoR). RL techniques are used to dynamically select the most effective move types as optimization progresses. The experimental results show thatRLPlaceoutperforms the widely used VTR 8 placer across all runtime/quality tradeoff points, achieving better QoR placement solutions in less runtime. On average, across the Titan23 suite of large FPGA benchmarks, RLPlace can reduce CPU time by$2.5\times $with result quality comparable to VTR 8, or improve wirelength by 8% (at a high CPU time budget) −26% (at a low CPU time budget) versus VTR 8.0 given the same CPU time.
Mohamed A. Elgammal, Kevin E. Murray, Vaughn Betz
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 Tensor Slices: FPGA Building Blocks For The Deep Learning Era
abstract
FPGAs are well-suited for accelerating deep learning (DL) applications owing to the rapidly changing algorithms, network architectures and computation requirements in this field. However, the generic building blocks available on traditional FPGAs limit the acceleration that can be achieved. Many modifications to FPGA architecture have been proposed and deployed including adding specialized artificial intelligence (AI) processing engines, adding support for smaller precision math like 8-bit fixed point and IEEE half-precision (fp16) in DSP slices, adding shadow multipliers in logic blocks, etc. In this paper, we describe replacing a portion of the FPGA’s programmable logic area with Tensor Slices. These slices have a systolic array of processing elements at their heart that support multiple tensor operations, multiple dynamically-selectable precisions and can be dynamically fractured into individual multipliers and MACs (multiply-and-accumulate). These slices have a local crossbar at the inputs that helps with easing the routing pressure caused by a large block on the FPGA. Adding these DL-specific coarse-grained hard blocks to FPGAs increases their compute density and makes them even better hardware accelerators for DL applications, while still keeping the vast majority of the real estate on the FPGA programmable at fine-grain.
Aman Arora 0001, Moinak Ghosh, Samidh Mehta, Vaughn Betz, Lizy Kurian John
ACM Trans. Reconfigurable Technol. Syst.4
2021 Tensor Slices to the Rescue: Supercharging ML Acceleration on FPGAs
abstract
FPGAs are well-suited for accelerating deep learning (DL) applications owing to the rapidly changing algorithms, network architectures and computation requirements in this field. However, the generic building blocks available on traditional FPGAs limit the acceleration that can be achieved. Many modifications to FPGA architecture have been proposed and deployed including adding specialized artificial intelligence (AI) processing engines, adding support for IEEE half-precision (fp16) math in DSP slices, adding hard matrix multiplier blocks, etc. In this paper, we describe replacing a small percentage of the FPGA's programmable logic area with Tensor Slices. These slices are arrays of processing elements at their heart that support multiple tensor operations, multiple dynamically-selectable precisions and can be dynamically fractured into individual adders, multipliers and MACs (multiply-and-accumulate). These tiles have a local crossbar at the inputs that helps with easing the routing pressure caused by a large slice. By spending ~3% of FPGA's area on Tensor Slices, we observe an average frequency increase of 2.45x and average area reduction by 0.41x across several ML benchmarks, including a TPU-like design, compared to an Intel Agilex-like baseline FPGA. We also study the impact of spending area on Tensor slices on non-ML applications. We observe an average reduction of 1% in frequency and an average increase of 1% in routing wirelength compared to the baseline, across the non-ML benchmarks we studied. Adding these ML-specific coarse-grained hard blocks makes the proposed FPGA a much efficient hardware accelerator for ML applications, while still keeping the vast majority of the real estate on the FPGA programmable at fine-grain.
Aman Arora 0001, Samidh Mehta, Vaughn Betz, Lizy Kurian John
FPGA3
2021 End-to-End FPGA-based Object Detection Using Pipelined CNN and Non-Maximum Suppression
abstract
Object detection is an important computer vision task, with many applications in autonomous driving, smart surveillance, robotics, and other domains. Single-shot detectors (SSD) coupled with a convolutional neural network (CNN) for feature extraction can efficiently detect, classify and localize various objects in an input image with very high accuracy. In such systems, the convolution layers extract features and predict the bounding box locations for the detected objects as well as their confidence scores. Then, a non-maximum suppression (NMS) algorithm eliminates partially overlapping boxes and selects the bounding box with the highest score per class. However, these two components are strictly sequential; a conventional NMS algorithm needs to wait for all box predictions to be produced before processing them. This prohibits any overlap between the execution of the convolutional layers and NMS, resulting in significant latency overhead and throughput degradation. In this paper, we present a novel NMS algorithm that alleviates this bottleneck and enables a fully-pipelined hardware implementation. We also implement an end-to-end system for low-latency SSD-MobileNet-V1 object detection, which combines a state-of-the-art deeply-pipelined CNN accelerator with a custom hardware implementation of our novel NMS algorithm. As a result of our new algorithm, the NMS module adds a minimal latency overhead of only 0.13μ s to the SSD-MobileNet-V1 convolution layers. Our end-to-end object detection system implemented on an Intel Stratix 10 FPGA runs at a maximum operating frequency of 350 MHz, with a throughput of 609 frames-per-second and an end-to-end batch-1 latency of 2.4 ms. Our system achieves 1.5× higher throughput and 4.4× lower latency compared to the current state-of-the-art SSD-based object detection systems on FPGAs.
Anupreetham Anupreetham, Mohamed Ibrahim 0005, Mathew Hall, Andrew Boutros, Ajay Kuzhively, Abinash Mohanty, Eriko Nurvitadhi, Vaughn Betz, Yu Cao 0001, Jae-sun Seo
FPL8
2021 Koios: A Deep Learning Benchmark Suite for FPGA Architecture and CAD Research
abstract
With the prevalence of deep learning (DL) in many applications, researchers are investigating different ways of optimizing FPGA architecture and CAD to achieve better quality-of-results (QoR) on DL-based workloads. In this optimization process, benchmark circuits are an essential component; the QoR achieved on a set of benchmarks is the main driver for architecture and CAD design choices. However, current academic benchmark suites are inadequate, as they do not capture any designs from the DL domain. This work presents a new suite of DL acceleration benchmark circuits for FPGA architecture and CAD research, called Koios. This suite of 19 circuits covers a wide variety of accelerated neural networks, design sizes, implementation styles, abstraction levels, and numerical precisions. These designs are larger, more data parallel, more heterogeneous, more deeply pipelined, and utilize more FPGA architectural features compared to existing open-source benchmarks. This enables researchers to pin-point architectural inefficiencies for this class of workloads and optimize CAD tools on more realistic benchmarks that stress the CAD algorithms in different ways. In this paper, we describe the designs in our benchmark suite, present results of running them through the Verilog-to-Routing (VTR) flow using a recent FPGA architecture model, and identify key insights from the resulting metrics. On average, our benchmarks have 3.7× more netlist primitives, 1.8× and 4.7× higher DSP and BRAM densities, and 1.7× higher frequency with 1.9× more near-critical paths compared to the widely-used VTR suite. Finally, we present two example case studies showing how architectural exploration for DL-optimized FPGAs can be performed using our new benchmark suite.
Aman Arora 0001, Andrew Boutros, Daniel Rauch, Aishwarya Rajen, Aatman Borda, Seyed Alireza Damghani, Samidh Mehta, Sangram Kate, Pragnesh Patel, Kenneth B. Kent, Vaughn Betz, Lizy Kurian John
FPL11
2021 StateLink: FPGA System Debugging via Flexible Simulation/Hardware Integration
abstract
Checkpoint-based debugging flows that allow moving the design state between an FPGA and a simulator have recently emerged. These flows combine the speed of hardware execution and the full observability and controllability of HDL simulation. However, they assume the entire system state can be moved to a simulator, limiting them to self-contained systems and precluding their use in network or CPU-attached FPGAs. In this paper, we present StateLink, a co-simulation framework that allows a design-under-test (DUT) running in a simulator to interact with other design elements that reside in hardware. StateLink creates links between DUT interfaces in the HDL simulation and their equivalents in hardware, thereby allowing the DUT to remain connected to and active in the overall hardware system after its state is moved to a simulator. This extends the functionality of checkpoint-based debugging frameworks to designs with external I/Os such as DRAM and Ethernet, and to designs that contain components with no simulation models. It also significantly decreases the simulation time of DUTs that are part of a large system. For example, it speeds up the HDL simulation of designs that interface with DRAM by up to 25 ×. Incorporating StateLink in a design typically adds no timing overhead and a modest hardware area overhead; for example, StateLink adds 916 LUTs to a 32-bit AXI memory-mapped and 1423 LUTs to a 32-bit AXI streaming interface.
Sameh Attia, Vaughn Betz
FPT2
2020 AIR: A Fast but Lazy Timing-Driven FPGA Router
abstract
Routing is a key step in the FPGA design process, which significantly impacts design implementation quality. Routing is also very time-consuming, and can scale poorly to very large designs. This paper describes the Adaptive Incremental Router (AIR), a high-performance timing-driven FPGA router. AIR dynamically adapts to the routing problem, which it solves `lazily' to minimize work. Compared to the widely used VPR 7 router, AIR significantly reduces route-time ($7.1 \times$ faster), while also improving quality (15% wirelength, and 18% critical path delay reductions). We also show how these techniques enable efficient incremental improvement of existing routing.
Kevin E. Murray, Vaughn Betz
ASP-DAC3
2020 StateMover: Combining Simulation and Hardware Execution for Efficient FPGA Debugging
abstract
Debugging consumes a large portion of FPGA design time, and with the growing complexity of traditional FPGA systems and the additional verification challenges posed by multiple FPGAs interacting within data centers, debugging productivity is becoming even more important. Current debugging flows either depend on simulation, which is extremely slow but has full visibility, or on hardware execution, which is fast but provides very limited control and visibility. In this paper, we present StateMover, a checkpointing-based debugging framework for FPGAs, which can move design state back and forth between an FPGA and a simulator in a seamless way. StateMover leverages the speed of hardware execution and the full visibility and ease-of-use of a simulator. This enables a novel debugging flow that has a software-like combination of speed with full observability and controllability. StateMover adds minimal hardware to the design to safely stop the design under test so that its state can be extracted or modified in an orderly manner. The added hardware has no timing overhead and a very small area overhead. StateMover currently supports Xilinx UltraScale devices, and its underlying techniques and tools can be ported to other device families that support configuration readback. Moving the state from/to an FPGA to/from a simulator can be performed in a few seconds for large FPGAs, enabling a new debugging flow.
Sameh Attia, Vaughn Betz
FPGA2
2020 HPIPE: Heterogeneous Layer-Pipelined and Sparse-Aware CNN Inference for FPGAs
abstract
This poster presents a novel cross-layer-pipelined Convolutional Neural Network accelerator architecture, and network compiler, that make use of precision minimization and parameter pruning to fit ResNet-50 entirely into on-chip memory on a Stratix 10 2800 FPGA. By statically partitioning the hardware across each of the layers in the network, our architecture enables full DSP utilization and reduces the soft logic per DSP ratio by roughly 4x over prior work on sparse CNN accelerators for FPGAs. This high DSP utilization, a frequency of 420MHz, and skipping zero weights enable our architecture to execute a sparse ResNet-50 model at a batch size of 1 at 3300 images/s, which is nearly 3x higher throughput than NVIDIA's fastest machine learning targeted GPU, the V100. We also present a network compiler and a flexible hardware interface that make it easy to add support for new types of neural networks, and to optimize these networks for FPGAs with different on-chip resources.
Mathew Hall, Vaughn Betz
FPGA2
2020 Using OpenCL to Enable Software-like Development of an FPGA-Accelerated Biophotonic Cancer Treatment Simulator
abstract
The simulation of light propagation through tissues is important for medical applications, such as photodynamic therapy (PDT) for cancer treatment. To optimize PDT an inverse problem, which works backwards from a desired distribution of light to the parameters that caused it, must be solved. These problems have no closed-form solution and therefore must be solved numerically using an iterative method. This involves running many forward light propagation simulations which is time-consuming and computationally intensive.
Tanner Young-Schultz, Lothar Lilge, Stephen Brown 0003, Vaughn Betz
FPGA4
2020 Feel Free to Interrupt: Safe Task Stopping to Enable FPGA Checkpointing and Context Switching
abstract
Saving and restoring an FPGA task state in an orderly manner is essential to enable hardware checkpointing, which is highly desirable to improve the ability to debug cloud-scale hardware services, and context switching, which allows multiple users to share FPGA resources. However, these features require task interruption, and stopping a task at an arbitrary time can cause several hazards including deadlock and data loss. In this article, we build a context saving and restoring simulator to simulate and identify these hazards. In addition, we derive design rules that should be followed to achieve safe task interruption. Finally, we propose task wrappers that can be placed around an FPGA task to implement these rules. The timing and area overheads added by these wrappers are very small; they add 1.8% area and no timing overhead to a full Memcached system. Taken together, these design rules and wrappers enable safe checkpointing and context switching in a wide variety of FPGA tasks, including those with multiple clocks, multi-cycle I/O transactions, and interface dependencies.
Sameh Attia, Vaughn Betz
ACM Trans. Reconfigurable Technol. Syst.2
2020 FPGA Logic Block Architectures for Efficient Deep Learning Inference
abstract
Reducing the precision of deep neural network (DNN) inference accelerators can yield large efficiency gains with little or no accuracy degradation compared to half or single precision floating-point by enabling more multiplication operations per unit area. A wide range of precisions fall on the pareto-optimal curve of hardware efficiency vs. accuracy with no single precision dominating, making the variable precision capabilities of FPGAs very valuable. We propose three types of logic block architectural enhancements and fully evaluate a total of six architectures that improve the area efficiency of multiplications and additions implemented in the soft fabric. Increasing the LUT fracturability and adding two adders to the ALM (4-bit Adder Double Chain architecture) leads to a 1.5× area reduction for arithmetic heavy machine learning (ML) kernels, while increasing their speed. In addition, this architecture also reduces the logic area of general applications by 6%, while increasing the critical path delay by only 1%. However, our highest impact option, which adds a 9-bit shadow multiplier to the logic clusters, reduces the area and critical path delay of ML kernels by 2.4× and 1.2×, respectively. These large gains come at a cost of 15% logic area increase for general applications.
Mohamed Eldafrawy, Andrew Boutros, Sadegh Yazdanshenas, Vaughn Betz
ACM Trans. Reconfigurable Technol. Syst.4
2020 VTR 8: High-performance CAD and Customizable FPGA Architecture Modelling
abstract
Developing Field-programmable Gate Array (FPGA) architectures is challenging due to the competing requirements of various application domains and changing manufacturing process technology. This is compounded by the difficulty of fairly evaluating FPGA architectural choices, which requires sophisticated high-quality Computer Aided Design (CAD) tools to target each potential architecture. This article describes version 8.0 of the open source Verilog to Routing (VTR) project, which provides such a design flow. VTR 8 expands the scope of FPGA architectures that can be modelled, allowing VTR to target and model many details of both commercial and proposed FPGA architectures. The VTR design flow also serves as a baseline for evaluating new CAD algorithms. It is therefore important, for both CAD algorithm comparisons and the validity of architectural conclusions, that VTR produce high-quality circuit implementations. VTR 8 significantly improves optimization quality (reductions of 15% minimum routable channel width, 41% wirelength, and 12% critical path delay), run-time (5.3× faster) and memory footprint (3.3× lower). Finally, we demonstrate VTR is run-time and memory footprint efficient, while producing circuit implementations of reasonable quality compared to highly-tuned architecture-specific industrial tools—showing that architecture generality, good implementation quality, and run-time efficiency are not mutually exclusive goals.
Kevin E. Murray, Oleg Petelin, Jia Min Wang, Mohamed Eldafrawy, Jean-Philippe Legault, Eugene Sha, Aaron Graham, Jean Wu, Matthew J. P. Walker, Hanqing Zeng, Panagiotis Patros, Jason Luu, Kenneth B. Kent, Vaughn Betz
ACM Trans. Reconfigurable Technol. Syst.15
2020 Optimizing FPGA Logic Circuitry for Variable Voltage Supplies
abstract
Unlike central processing units (CPUs), field-programmable gate arrays (FPGAs) have conventionally been powered with a fixed supply voltage (Vdd). However, recent efforts have shown that adopting dynamic voltage scaling reduces FPGA power consumption significantly. In this article, we analyze the delay sensitivity of different FPGA circuit elements to supply voltage changes and determine that conventional lookup table (LUT) designs greatly impact variable Vdd operation. To build FPGAs with lower delay sensitivity to Vdd, we propose several new LUT designs, including gate boosting the LUT, decoding the slowest two inputs of the LUT, and using separate voltage islands for the FPGA LUTs and routing. Our fastest proposed design (decode driver island) reduces the area-delay product of the FPGA logic plus routing tile compared to a conventional design by 12% and 52% at Vdd values of 0.8 V (the nominal voltage) and 0.6 V, respectively. Since our proposed FPGA tile designs are faster and have lower delay sensitivity to voltage, they offer better Energy-Delay2 product (ED2) than that of the baseline at nominal Vdd and below. Our decode-driver-island FPGA achieves a 26% ED2 reduction over the conventional design at the most efficient ED2 operating point.
Ibrahim Ahmed 0001, Linda L. Shen, Vaughn Betz
IEEE Trans. Very Large Scale Integr. Syst.3
2020 Optimizing FPGA Logic Block Architectures for Arithmetic
abstract
Hardened adder and carry logic is widely used in commercial field-programmable gate arrays (FPGAs) to improve the efficiency of arithmetic functions. There are many design choices and complexities associated with such hardening, including circuit design, FPGA architectural choices, and the computer-aided design (CAD) flow. However, these choices have not been studied much and hence we explore a number of possibilities. We also highlight front-end elaboration optimization that helps ameliorate the restrictions placed on logic synthesis by hardened arithmetic. We show that hard adders and carry chains increase the performance of simple adders by a factor of 4 or more, but on larger benchmark designs that contain arithmetic improve the overall performance by 15%. Our results also show that for complete application circuits simple hardened ripple-carry adders perform as well as more complex carry-lookahead adders. Our best non-fracturable lookup table (non-fLUT) architecture with hardened arithmetic yields 12% better area-delay product than architectures without hardened arithmetic. We also investigate the impact of fLUTs and their interaction with hardened arithmetic. We find that fLUTs offer significant (12%-15%) area reduction, which is complementary to the delay reduction of hardened arithmetic. Therefore, our best fLUT architectures which use two bits of hardened arithmetic achieve 25% better area-delay product than non-fLUT architectures without hardened arithmetic.
Kevin E. Murray, Jason Luu, Matthew J. P. Walker, Conor McCullough, Safeen Huda, Charles Chiasson, Kenneth B. Kent, Jason Helge Anderson, Jonathan Rose, Vaughn Betz
IEEE Trans. Very Large Scale Integr. Syst.12
2019 Safe Task Interruption for FPGAs
abstract
Saving and restoring the state of an FPGA task in an orderly manner is essential for enabling hardware checkpointing and context switching. However, it requires task interruption, and stopping a task at an arbitrary time can cause several hazards including deadlock and data loss. In this work, we build a context switching simulator to simulate and identify these hazards. In addition, we introduce design rules that should be followed to achieve safe task interruption, and propose task wrappers that can be placed around an FPGA task to implement these rules.
Sameh Attia, Vaughn Betz
FCCM2
2019 Fast Voltage Transients on FPGAs: Impact and Mitigation Strategies
abstract
As FPGAs grow in size and speed, so too does their power consumption. Power consumption on recent FPGAs has increased to the point that it is comparable to that of high-end CPUs. To mitigate this problem, power reduction techniques such as dynamic voltage scaling (DVS) and clock gating can potentially be applied to FPGAs. However, it is unclear whether they are safe in the presence of fast voltage transients. These fast voltage transients are caused by large changes in activity which we believe are common in most designs. Previous work has shown that it is these fast voltage transients that produce the largest variations in delay. In our work, we measure the impact transients have on applications and present a mitigation strategy to prevent them from causing timing failures. We create transient generators that are able to significantly reduce an application's measured Fmax, by up to 25. We also show that transients are very fast and produce immediate timing impact and hence transient mitigation must occur within the same clock cycle as the transient. We create a clock edge suppressor that is able to detect when a transient event is happening and delay the clock edge, thus preventing any timing failures. Using our clock edge suppressor, we show that we can run an application at full frequency in the presence of fast voltage transients, thereby enabling more aggressive DVS approaches and larger power savings.
Linda L. Shen, Ibrahim Ahmed 0001, Vaughn Betz
FCCM3
2019 Math Doesn't Have to be Hard: Logic Block Architectures to Enhance Low-Precision Multiply-Accumulate on FPGAs
abstract
Recent work has shown that using low-precision arithmetic in Deep Neural Network (DNN) inference acceleration can yield large efficiency gains with little or no accuracy degradation compared to half or single precision floating-point by enabling more MAC operations per unit area. The most efficient precision is a complex function of the DNN application, structure and required accuracy, which makes the variable precision capabilities of FPGAs very valuable. We propose three logic block architecture enhancements to increase the density and reduce the delay of multiply-accumulate (MAC) operations implemented in the soft fabric. Adding another level of carry chain to the ALM (extra carry chain architecture) leads to a 1.5x increase in MAC density, while ensuring a small impact on general designs as it adds only 2.6% FPGA tile area and a representative critical path delay increase of 0.8%. On the other hand, our highest impact option, which combines our 4-bit Adder architecture with a 9-bit Shadow Multiplier, increases MAC density by 6.1x, at the cost of larger tile area and representative critical path delay overheads of 16.7% and 9.8%, respectively.
Andrew Boutros, Mohamed Eldafrawy, Sadegh Yazdanshenas, Vaughn Betz
FPGA4
2019 Becoming More Tolerant: Designing FPGAs for Variable Supply Voltage
abstract
With the end of Dennard scaling, FPGA power consumption has become a major concern. While FPGAs are conventionally supplied by a fixed supply voltage (Vdd), recent industrial (SmartVID) and academic solutions (dynamic voltage scaling) have shown significant power savings by scaling the FPGA Vdd on a chip-specific or chip-and application-specific basis. However, FPGAs have historically been designed for fixed-Vdd operation, which raises the question of whether we can design FPGA circuitry that is better suited for voltage scaling. In this work, we show that conventional LUTs are more sensitive to voltage than routing, so we design different LUT circuits that are more tolerant to voltage scaling. Compared to a conventional LUT, our fastest proposed LUT reduces the average critical path delay by 14% and 47% at nominal (0.8 V) Vdd and at reduced (0.6 V) Vdd, respectively. This significant reduction in delay comes at a cost of only 8% FPGA tile area increase. Our proposed LUT designs result in lower energy-delay and energy-delay^2 products at nominal Vdd and below.
Ibrahim Ahmed 0001, Linda L. Shen, Vaughn Betz
FPL3
2019 Calculated Risks: Quantifying Timing Error Probability With Extended Static Timing Analysis
abstract
Timing analysis is a key step in the digital design process. By modeling device delay variations statistical static timing analysis (SSTA) reduces pessimism compared to traditional static timing analysis (STA). However, it ignores the circuit's logic which causes some timing paths to never, or only rarely, be sensitized. We introduce a general timing analysis approach and tool to calculate the probability that individual timing paths are sensitized, enabling the calculation of bounding delay distributions over all input combinations. We show how this analysis is related to the well-known #SAT problem and present approaches to improve scalability, achieving, on average, results 75% to 37% less pessimistic than STA while running 569 to 16 times faster than Monte-Carlo timing simulation.
Kevin E. Murray, Andrea Suardi, Vaughn Betz, George A. Constantinides
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2019 FRoC 2.0: Automatic BRAM and Logic Testing to Enable Dynamic Voltage Scaling for FPGA Applications
abstract
In earlier technology nodes, FPGAs had low power consumption compared to other compute chips such as CPUs and GPUs. However, in the 14nm technology node, FPGAs are consuming unprecedented power in the 100+W range, making power consumption a pressing concern. To reduce FPGA power consumption, several researchers have proposed deploying dynamic voltage scaling. While the previously proposed solutions show promising results, they have difficulty guaranteeing safe operation at reduced voltages for applications that use the FPGA hard blocks. In this work, we present the first DVS solution that is able to fully handle FPGA applications that use BRAMs. Our solution not only robustly tests the soft logic component of the application but also tests all components connected to the BRAMs. We extend a previously proposed CAD tool, FRoC, to automatically generate calibration bitstreams that are used to measure the application’s critical path delays on silicon. The calibration bitstreams also include testers that ensure all used SRAM cells operate safely while scaling V dd . We experimentally show that using our DVS solution we can save 32% of the total power consumed by a discrete Fourier transform application running with the fixed nominal supply voltage and clocked at the F max reported by static timing analysis.
Ibrahim Ahmed 0001, Shuze Zhao, James Meijers, Olivier Trescases, Vaughn Betz
ACM Trans. Reconfigurable Technol. Syst.5
2019 COFFE 2: Automatic Modelling and Optimization of Complex and Heterogeneous FPGA Architectures
abstract
FPGAs are becoming more heteregeneous to better adapt to different markets, motivating rapid exploration of different blocks/tiles for FPGAs. To evaluate a new FPGA architectural idea, one should be able to accurately obtain the area, delay, and energy consumption of the block of interest. However, current FPGA circuit design tools can only model simple, homogeneous FPGA architectures with basic logic blocks and also lack DSP and other heterogeneous block support. Modern FPGAs are instead composed of many different tiles, some of which are designed in a full custom style and some of which mix standard cell and full custom styles. To fill this modelling gap, we introduce COFFE 2, an open-source FPGA design toolset for automatic FPGA circuit design. COFFE 2 uses a mix of full custom and standard cell flows and supports not only complex logic blocks with fracturable lookup tables and hard arithmetic but also arbitrary heterogeneous blocks. To validate COFFE 2 and demonstrate its features, we design and evaluate a multi-mode Stratix III-like DSP block and several logic tiles with fracturable LUTs and hard arithmetic. We also demonstrate how COFFE 2’s interface to VTR allows full evaluation of block-routing interfaces and various fracturable 6-LUT architectures.
Sadegh Yazdanshenas, Vaughn Betz
ACM Trans. Reconfigurable Technol. Syst.2
2019 The Costs of Confidentiality in Virtualized FPGAs
abstract
Some modern datacenters are augmenting their compute infrastructure by deploying field-programmable gate arrays (FPGAs) to provide users with specialized accelerators that offer superior compute capability, increased energy efficiency, lower latency, and more programming flexibility than CPUs. However, the higher programming flexibility of FPGAs also gives more capabilities to malicious users to remotely sniff data from other applications running on the same FPGA. This has created a challenge for efficient utilization of FPGAs in datacenters: FPGAs in datacenters are currently not shared between users due to potential security risks. In this paper, we propose different techniques to defeat data-sniffing attacks in datacenter FPGAs by encrypting/decrypting the user application's data. We describe techniques that are appropriate to different trust levels and rigorously evaluate the costs of these data confidentiality techniques in current virtualized FPGAs. In addition, for each trust level, we propose an architectural change to the FPGA to mitigate the costs of providing data confidentiality. We also investigate the role of interconnect in these architectural changes and demonstrate that more efficient security features can be implemented together with the interconnect if the FPGAs use a hard network on chip.
Sadegh Yazdanshenas, Vaughn Betz
IEEE Trans. Very Large Scale Integr. Syst.2
2018 A High-Level Synthesis Case Study on Light Propagation Simulation in Turbid Media
abstract
In this work, we look into the benefit of using High-Level Synthesis (HLS) in building and accelerating complex systems with floating-point operations. We present a highly-optimized Monte-Carlo (MC) simulator for light propagation in 3D voxel-based biological tissue representations using HLS. We show how to utilize HLS in creating efficient structures that help achieve the desired throughput. We use Vivado to implement the design on a Xilinx Kintex Ultrascale FPGA running at 150 MHz. With a design time of 1.5 months, experimental results show a 3x speedup against the fastest software simulator published to date.
Abdul-Amir Yassine, Yasmin Afsharnejad, Omar Ragheb, Vaughn Betz, Paul Chow
FCCM4
2018 Latency Insensitive Design Styles for FPGAs
abstract
Long distance interconnect delays are not scaling well with process technology, thereby leading to long routes strongly impacting the critical path of large FPGA designs. This forces the designer to pipeline long connections, which necessitates time consuming logic redesign in traditional latency-sensitive systems. Latency-insensitive design (LID) is an increasingly attractive alternative as the typical latency of long distance interconnect grows, since LID decouples the design of the interconnect from that of the computational modules. By doing so, LID simplifies timing closure, improves forward compatibility (migration of systems to future FPGAs) and makes automated system-level pipelining feasible. Modern FPGAs, such as Stratix 10 which includes pipelined interconnect, make it difficult to use traditional LID solutions without significant area and frequency overhead. We present two LID styles that are more suitable for FPGAs and compare them to traditional LID. Our best system gained 2x area efficiency and 18% speed efficiency over traditional LID. Additionally, our designs come at a minimal speed overhead of only 3% compared to that of a latency-sensitive design.
Mustafa Abbas, Vaughn Betz
FPL2
2018 Automatic BRAM Testing for Robust Dynamic Voltage Scaling for FPGAs
abstract
Recently FPGA researchers have proposed different approaches to enable dynamic voltage scaling (DVS) for FPGAs. While the proposed approaches have shown that DVS is able to significantly reduce FPGA power consumption, most of these solutions were developed only for the soft fabric of the FPGA and hence cannot be deployed for applications that use the FPGA hard blocks such as block RAMs (BRAMs). In this work, we extend a previously proposed offline calibration-based DVS approach to enable DVS for FPGAs with BRAMs; we build testing circuitry to ensure that all used BRAM cells operate safely while scaling the supply voltage, and we develop testing procedures that are able to measure the delay of timing paths that start or end at BRAMs. We extend the CAD tool FRoC to automatically generate calibration designs with BRAM testers along with soft fabric testers to measure the actual Fmax of each application on any chip under different operating conditions; this information is stored in a calibration table that is then used when the application is running to scale the supply voltage to the minimum value that guarantees safe operation at the desired speed. Using our proposed solution, we show that we can run a discrete Fourier transform core with 32 % and 46 % power reduction compared to the conventional fixed-voltage operation at the reported F_max and at a lower clock frequency, respectively.
Ibrahim Ahmed 0001, Shuze Zhao, James Meijers, Olivier Trescases, Vaughn Betz
FPL5
2018 Embracing Diversity: Enhanced DSP Blocks for Low-Precision Deep Learning on FPGAs
abstract
Use of reduced precisions in Deep Learning (DL) inference tasks has recently been shown to significantly improve accelerator performance and greatly reduce both model memory footprint and the required external memory bandwidth. With appropriate network retuning, reduced precision networks can achieve accuracy close or equal to that of full-precision floating-point models. Given the wide spectrum of precisions used in DL inference, FPGAs' ability to create custom bit-width datapaths gives them an advantage over other acceleration platforms in this domain. However, the embedded DSP blocks in the latest Intel and Xilinx FPGAs do not natively support precisions below 18-bit and thus can not efficiently pack low-precision multiplications, leaving the DSP blocks under-utilized. In this work, we present an enhanced DSP block that can efficiently pack 2× as many 9-bit and 4× as many 4-bit multiplications compared to the baseline Arria-10-like DSP block at the cost of 12% block area overhead which leads to only 0.6% total FPGA core area increase. We quantify the performance gains of using this enhanced DSP block in two state-of-the-art convolutional neural network accelerators on three different models: AlexNet, VGG-16, and ResNet-50. On average, the new DSP block enhanced the computational performance of the 8-bit and 4-bit accelerators by 1.32× and 1.6× and at the same time reduced the utilized chip area by 15% and 30% respectively.
Andrew Boutros, Sadegh Yazdanshenas, Vaughn Betz
FPL3
2018 Tatum: Parallel Timing Analysis for Faster Design Cycles and Improved Optimization
abstract
Static Timing Analysis (STA) is used to evaluate the correctness and performance of a digital circuit implementation. In addition to final sign-off checks, STA is called numerous times during placement and routing to guide optimization. As a result, STA consumes a significant fraction of the time required for design implementation; to make progress reducing FPGA compile times we need faster STA. We evaluate the suitability of both GPU and multi-core CPU platforms for accelerating STA. On core STA algorithms our GPU kernel achieves a 6.2 times kernel speed-up but data transfer overhead reduces this to 0.9 times. Our best CPU implementation achieves a 9.2 times parallel speed-up on 32 cores, yielding a 15.2 times overall speed-up compared to the VPR analyzer, and a 6.9 times larger parallel speed-up than a recent parallel ASIC timing analyzer. We then show how reducing the run-time cost of STA can be leveraged to improve optimization quality, reducing critical path delay by 4%.
Kevin E. Murray, Vaughn Betz
FPT2
2018 Improving Confidentiality in Virtualized FPGAs
abstract
FPGAs are being deployed in modern datacenters to provide users with specialized accelerators that offer superior compute capability, increased energy efficiency, lower latency, and more programming flexibility than CPUs. However, FPGAs are not utilized as efficiently in datacenters: unlike CPUs, FPGAs in datacenters are currently not shared between users due to potential security risks. The higher flexibility that comes with FPGAs also gives more capabilities to malicious users. Several recent studies have demonstrated examples of FPGA user applications capable of remotely sniffing data from other applications running on the same FPGA. In this work, we look at various ways to ameliorate these threats by encrypting/decrypting the user application's data under different trust levels for current virtualized FPGAs. We also discuss the role of interconnect and discuss the potential of more efficient security features that can be implemented together with the interconnect if the FPGAs use a hard network on chip.
Sadegh Yazdanshenas, Vaughn Betz
FPT2
2018 Automatic Application-Specific Calibration to Enable Dynamic Voltage Scaling in FPGAs
abstract
Dynamic voltage scaling (DVS) is one of the most effective ways to reduce integrated circuit power. However, the programmability of field programmable gate arrays (FPGAs) means that the critical paths depend on the application configured into the FPGA and this makes DVS more difficult. We propose a DVS technique that is able to determine the minimum safe Vddof any application for each FPGA chip. For each application, we create multiple calibration bit-streams that are used to generate a calibration table (CT), which stores the actual failing points of that application on a specific FPGA, under various operating conditions. This CT is used to scale Vddwhile the application is running to guarantee safe operation with minimal power consumption. We develop an automated tool (FRoC) that ensures a fast-robust-calibration of the FPGA to any application using it. FRoC makes the calibration process invisible to FPGA users, does not add any extra manual steps to the design process, and uses novel algorithms to minimize the extra flash storage requirements for calibration. Our results show that across a large suite of benchmarks the calibration process requires a geomean of less than four bit-streams and our DVS technique achieves a 33% total power reduction on two large applications.
Ibrahim Ahmed 0001, Shuze Zhao, Olivier Trescases, Vaughn Betz
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2018 You Cannot Improve What You Do not Measure: FPGA vs. ASIC Efficiency Gaps for Convolutional Neural Network Inference
abstract
Recently, deep learning (DL) has become best-in-class for numerous applications but at a high computational cost that necessitates high-performance energy-efficient acceleration. The reconfigurability of FPGAs is appealing due to the rapid change in DL models but also causes lower performance and area-efficiency compared to ASICs. In this article, we implement three state-of-the-art computing architectures (CAs) for convolutional neural network (CNN) inference on FPGAs and ASICs. By comparing the FPGA and ASIC implementations, we highlight the area and performance costs of programmability to pinpoint the inefficiencies in current FPGA architectures. We perform our experiments using three variations of these CAs for AlexNet, VGG-16 and ResNet-50 to allow extensive comparisons. We find that the performance gap varies significantly from 2.8× to 6.3×, while the area gap is consistent across CAs with an 8.7 average FPGA-to-ASIC area ratio. Among different blocks of the CAs, the convolution engine, constituting up to 60% of the total area, has a high area ratio ranging from 13 to 31. Motivated by our FPGA vs. ASIC comparisons, we suggest FPGA architectural changes such as increasing DSP block count, enhancing low-precision support in DSP blocks and rethinking the on-chip memories to reduce the programmability gap for DL applications.
Andrew Boutros, Sadegh Yazdanshenas, Vaughn Betz
ACM Trans. Reconfigurable Technol. Syst.3
2018 Wotan: Evaluating FPGA Architecture Routability without Benchmarks
abstract
FPGA routing architectures consist of routing wires and programmable switches that together account for the majority of the fabric delay and area, making evaluation and optimization of an FPGA’s routing architecture very important. Routing architectures have traditionally been evaluated using a full synthesize, pack, place and route CAD flow over a suite of benchmark circuits. While the results are accurate, a full CAD flow has a long runtime and is often tuned to a specific FPGA architecture type, which limits exploration of different architecture options early in the design process. In this article, we present Wotan, a tool to quickly estimate routability for a wide range of architectures without the use of benchmark circuits. At its core, our routability predictor efficiently counts paths through the FPGA routing graph to (1) estimate the probability of node congestion and (2) estimate the probabilities to successfully route a randomized subset of(source, sink)pairs, which are then combined into an overall routability metric. We describe our predictor and present routability estimates for a range of 6-LUT and 4-LUT architectures using mixes of wire types connected in complex ways, showing a rank correlation of 0.91 with routability results from the full VPR CAD flow while requiring 18× less CPU effort.
Oleg Petelin, Vaughn Betz
ACM Trans. Reconfigurable Technol. Syst.2
2018 Enhancing FPGAs with Magnetic Tunnel Junction-Based Block RAMs
abstract
While plentiful on-chip memory is necessary for many designs to fully utilize an FPGA’s computational capacity, SRAM scaling is becoming more difficult because of increasing device variation. An alternative is to build FPGA block RAM (BRAM) from magnetic tunnel junctions (MTJ), as this emerging embedded memory has a small cell size, low energy usage, and good scalability. We conduct a detailed comparison study of SRAM and MTJ BRAMs that includes cell designs that are robust with device variation, transistor-level design and optimization of all the required BRAM-specific circuits, and variation-aware simulation at the 22nm node. At a 256Kb block size, MTJ-BRAM is 3.06× denser and 55% more energy efficient and its F max is 274MHz, which is adequate for most FPGA system clock domains. We also detail further enhancements that allow these 256 Kb MTJ BRAMs to operate at a higher speed of 353MHz for the streaming FIFOs, which are very common in FPGA designs and describe how the non-volatility of MTJ BRAM enables novel on-chip configuration and power-down modes. For a RAM architecture similar to the latest commercial FPGAs, MTJ-BRAMs could expand FPGA memory capacity by 2.95× with no die size increase.
Kosuke Tatsumura, Sadegh Yazdanshenas, Vaughn Betz
ACM Trans. Reconfigurable Technol. Syst.3
2018 High-Performance Instruction Scheduling Circuits for Superscalar Out-of-Order Soft Processors
abstract
Soft processors have a role to play in simplifying field-programmable gate array (FPGA) application design as they can be deployed only when needed, and it is easier to write and debug single-threaded software code than create hardware. The breadth of this second role increases when the performance of the soft processor increases, yet the sophisticated out-of-order superscalar approaches that arrived in the mid-1990s are not employed, despite their area cost now being easily tolerable. In this article, we take an important step toward out-of-order execution in soft processors by exploring instruction scheduling in an FPGA substrate. This differs from the hard-processor design problem because the logic substrate is restricted to LUTs, whereas hard processor scheduling circuits employ CAM and wired-OR structures to great benefit. We discuss both circuit and microarchitectural trade-offs and compare three circuit structures for the scheduler, including a new structure called a fused-logic matrix scheduler . Using our optimized circuits, we show that four-issue distributed schedulers with up to 54 entries can be built with the same cycle time as the commercial Nios II/f soft processor (240MHz). This careful design has the potential to significantly increase both the IPC and raw compute performance of a soft processor, compared to current commercial soft processors.
Henry Wong, Vaughn Betz, Jonathan Rose
ACM Trans. Reconfigurable Technol. Syst.2
2017 Quantifying error: Extending static timing analysis with probabilistic transitions
abstract
Timing analysis is a cornerstone of the digital design process. Statistical Static Timing Analysis was introduced to reduce pessimism by modelling device delay variations. However it ignores circuit logic, which may cause some timing paths to never or only rarely be sensitized. We introduce a general timing analysis approach and tool to calculate the probability that individual timing paths are sensitized, enabling the calculation of bounding delay distributions over all input combinations. We show the connection to the well-known #SAT problem and present approaches to improve scalability, achieving average results 46 to 32% less pessimistic than Static Timing Analysis while running 14.6 to 44.0 times faster than Monte-Carlo timing simulation.
Kevin E. Murray, Andrea Suardi, Vaughn Betz, George A. Constantinides
DATE3
2017 Don't Forget the Memory: Automatic Block RAM Modelling, Optimization, and Architecture Exploration
Sadegh Yazdanshenas, Kosuke Tatsumura, Vaughn Betz
FPGA3
2017 Find the real speed limit: FPGA CAD for chip-specific application delay measurement
abstract
Process variation is increasing with each successive technology node, and it has reached the point where the worst-case timing modelling employed by current FPGA CAD tools is significantly underutilizing the available silicon. Previous studies have proposed exploiting FPGA reconfigurability to reduce this underutilization using techniques such as late binding and dynamic voltage scaling. Most of the proposed solutions require the ability to measure the target application's delay on each configured chip. To accurately measure the delay of an application on a certain chip, we must measure the delay of its speed limiting paths on this specific chip. In this paper, we present a variation-aware CAD tool that automatically generates calibration bitstreams to measure the delay of any input application. Our tool identifies the statistically critical paths of the circuit and optimally selects which paths to test such that it minimizes the chances of reporting an optimistic delay, under a constraint on the number of allowed calibration bitstreams. Experimental results across a suite of benchmarks show that with one calibration bitstream we achieve 16× lower probability of reporting an optimistic delay compared to a greedy approach. With three calibration bitstreams, we reduce the probability of optimism to two chips in a million, approximately 6,000 × lower than a greedy approach.
Ibrahim Ahmed 0001, Shuze Zhao, Olivier Trescases, Vaughn Betz
FPL4
2017 Quantifying and mitigating the costs of FPGA virtualization
abstract
FPGAs are being incorporated into contemporary datacenters in order to improve computational capacity, power consumption, and processing latency. Efficiently integrating FP-GAs in datacenters is, however, quite challenging. Ideally, smaller tasks could share a device and the cloud management layer would be able to partially reconfigure the device to allocate its free resources to incoming tasks. Moreover, to facilitate FPGA hardware upgrades without undue porting effort for previously developed accelerator tasks, the complexities associated with board-specific system-level integration should be abstracted away from designers. By meeting these requirements, FPGAs in the cloud would become multi-user virtualized resources with increased availability and elasticity. The virtualization of FPGAs, however, comes with two major costs in current FPGAs: lower application operating frequency, and extravagant use of routing resources. In this paper, we quantify the costs of FPGA virtualization and demonstrate that for an FPGA that supports four independent tasks, virtualization reduces the task average frequency by 18% to 46% and increases wire usage to 2.6×. We also investigate the cause of these costs and show that the use of hard NoCs in future datacenter-optimized FPGAs would facilitate FPGA virtualization without sacrificing operating frequency or routing resources.
Sadegh Yazdanshenas, Vaughn Betz
FPL2
2017 Automatic circuit design and modelling for heterogeneous FPGAs
abstract
Contemporary FPGAs are composed of a mix of full custom and standard cell-based circuitry, organized into many heterogeneous blocks and the programmable routing. To explore new FPGA architectures, in particular those incorporating new hard blocks, we must estimate the area, power and delay of any new block of interest, and would like to do so efficiently so that many ideas can be evaluated. Unfortunately, existing open source academic FPGA circuit modelling tools cannot target such a wide range of blocks and cannot mix custom and standard cell circuitry. In this work, we present an enhanced (hybrid) COFFE flow that automatically optimizes and models a wide variety of blocks using an intelligent mix of full custom transistor sizing and standard cell flows. Hybrid COFFE can generate models for both advanced logic blocks enriched with hard arithmetic and fracturable LUTs using its full custom flow and arbitrary hard blocks (such as DSP blocks) described in standard Hardware Description Languages (HDLs) that are fabricated using a mix of standard cell and full custom flows. To validate this hybrid flow we model a Stratix-III like DSP block in 65 nm CMOS and find good agreement with published commercial data. The resulting block and routing models are output in the Verilog-To-Routing (VTR) architecture format to facilitate architectural exploration of advanced FPGA blocks.
Sadegh Yazdanshenas, Vaughn Betz
FPT2
2017 Design and Applications for Embedded Networks-on-Chip on FPGAs
abstract
Field-programmable gate-arrays (FPGAs) have evolved to include embedded memory, high-speed I/O interfaces and processors, making them both more efficient and easier-to-use for compute acceleration and networking applications. However, implementing on-chip communication is still a designer's burden wherein custom system-level buses are implemented using the fine-grained FPGA logic and interconnect fabric. Instead, we propose augmenting FPGAs with an embedded network-on-chip (NoC) to implement system-level communication. We design custom interfaces to connect a packet-switched NoC to the FPGA fabric and I/Os in a configurable and efficient way and then define the necessary conditions to implement common FPGA design styles with an embedded NoC. Four application case studies highlight the advantages of using an embedded NoC. We show that access latency to external memory can be ~1.5× lower. Our application case study with image compression shows that an embedded NoC improves frequency by 10-80%, reduces utilization of scarce long wires by 40% and makes design easier and more predictable. Additionally, we leverage the embedded NoC in creating a programmable Ethernet switch that can support up to 819 Gb/s-5× more switching bandwidth and 3× lower area compared to previous work. Finally, we design a 400 Gb/s NoC-based packet processor that is very flexible and more efficient than other FPGA-based packet processors.
Mohamed S. Abdelfattah, Andrew Bitar, Vaughn Betz
IEEE Trans. Computers3
2016 High Performance Instruction Scheduling Circuits for Out-of-Order Soft Processors
abstract
Soft processors have a role to play in easing the difficulty of designing applications into FPGAs for two reasons: first, they can be deployed only when needed, unlike permanent on-die hard processors. Second, for the portions of an application that can function sufficiently fast on a soft processor, it is far easier to write and debug single-threaded software code than to create hardware. The breadth of this second role increases when the performance of the soft processor increases, yet there has been little progress in the performance of soft processors since their commercial inception -- in particular, the sophisticated out-of-order superscalar approaches that arrived in the mid 1990s are not employed, despite the fact that their area cost is now easily tolerable. In this paper we take an important step towards out-of-order execution in soft processors by exploring instruction scheduling in an FPGA substrate. This differs from the hard-processor design problem because the logic substrate is restricted to LUTs, whereas hard processor scheduling circuits employ CAM and wired-OR structures to great benefit. We discuss both circuit and microarchitectural trade-offs, and compare three circuit structures for the scheduler, including a new structure called a fused-logic matrix scheduler. With this circuit, large schedulers up to 40 entries can be built with the same cycle time as the commercial Nios II/f soft processor (240~MHz). This careful design has the potential to significantly increase both the IPC and raw compute performance of a soft processor, compared to current commercial soft processors.
Henry Wong, Vaughn Betz, Jonathan Rose
FCCM2
2016 LYNX: CAD for FPGA-based networks-on-chip
abstract
We present a computer-aided design (CAD) tool that automatically connects an FPGA application using an embedded network-on-chip (NoC). After discussing the CAD flow steps, we delve into the details of implementing transaction communication using our CAD tool. This request-reply type of communication requires special consideration on FPGAs, for example: low round-trip latency, fair arbitration and correct ordering. We show how to implement transaction communication using embedded NoCs, and show that we can improve latency, throughput and efficiency compared to soft buses generated by a commercial CAD tool.
Mohamed S. Abdelfattah, Vaughn Betz
FPL2
2016 Measure twice and cut once: Robust dynamic voltage scaling for FPGAs
abstract
Although dynamic voltage scaling (DVS) is a popular power reduction solution that has been widely used by processors and ASICs, it is still not commercially adopted by FPGAs. A unique feature of FPGAs that leads to challenges in adopting DVS is that the critical path and hence the minimum safe Vdddepends on the configured application. We present a robust DVS technique that solves these challenges. For each application, we generate a calibration table (CT) that stores the actual failing points of that application on a specific FPGA, under various operating conditions. This CT is used to scale Vddwhile the application is running to guarantee safe operation with minimal power consumption. We develop an automated tool (FRoC) that ensures a Fast-Robust-Calibration of the FPGA to any application using it. FRoC ensures that the calibration process is invisible to FPGA users and does not add any extra manual steps to the design process. We show that our proposed DVS technique achieves a 33% total power reduction on two large applications.
Ibrahim Ahmed 0001, Shuze Zhao, Olivier Trescases, Vaughn Betz
FPL4
2016 The speed of diversity: Exploring complex FPGA routing topologies for the global metal layer
abstract
The rapid growth of wire RC delay with technology scaling has put increasing pressure on FPGA architects to make more efficient use of the different layers available in the metal stack. While commercial FPGA architectures have implemented the majority of inter-logic-block wiring on the lower metal layers and a small fraction of wires on the least-resistive upper metal layers, published explorations have largely ignored the question of how to exploit the different layers of the metal stack, focusing instead on very simple interconnect topologies and physical models. We generate VPR architectures and detailed area and delay models at the 22nm node and present enhancements to VPR that enable us to describe and evaluate complex interconnect topologies. We use our new architectures and tool enhancements to explore complex interconnect patterns suitable for modern unidirectional architectures and suggest topologies to connect wires on the semi-global and global metal layers. The proposed topologies improve the critical path routing delay by 17% compared to architectures with no global layer wires, and by 5-13% compared to architectures with global layer wires using the default VPR switch pattern.
Oleg Petelin, Vaughn Betz
FPL2
2016 High density, low energy, magnetic tunnel junction based block RAMs for memory-rich FPGAs
abstract
Many important applications demand large amounts of on-chip memory both to fully utilize an FPGA's computational capacity and to minimize energy-consuming off-chip memory accesses, leading some recent commercial FPGAs to add higher-capacity on-chip block RAMs (BRAMs). While memory is becoming more important to FPGA designs, SRAM scaling is becoming more difficult because of increasing device variation. An alternative is to build FPGA BRAM from magnetic tunnel junction (MTJ) cells as this emerging embedded memory features a small cell size, low energy usage, and good scalability. In this work, we conduct a detailed comparison study of SRAM and MTJ BRAMs that includes cell designs that are robust with device variation, transistor-level design and optimization of all the required BRAM-specific circuits, and variation-aware simulation at the 22nm node. We find that as the capacity of a BRAM increases, the MTJ benefits of high-density and low-energy increase and its drawback of lower speed is mitigated. At a 256 Kb block size, MTJ-BRAM is 3.06× denser and 55% more energy efficient and its Fmaxis 274 MHz, which is adequate for most FPGA system clock domains. We detail how the non-volatility of an MTJ-BRAM saves energy, especially for narrow write operations which are common for the width-configurable BRAMs of FPGAs. For a RAM architecture similar to the latest commercial FPGAs, MTJ-based block RAMs reduce the FPGA fabric area by 28%, or alternatively could expand FPGA memory capacity by 2.95× with no die size increase.
Kosuke Tatsumura, Sadegh Yazdanshenas, Vaughn Betz
FPT3
2016 Microarchitecture and Circuits for a 200 MHz Out-of-Order Soft Processor Memory System
abstract
Although FPGAs have grown in capacity, FPGA-based soft processors have grown very little because of the difficulty of achieving higher performance in exchange for area. Superscalar out-of-order processors promise large performance gains, and the memory subsystem is a key part of such a processor that must help supply increased performance. In this article, we describe and explore microarchitectural and circuit-level tradeoffs in the design of such a memory system. We show the significant instructions-per-cycle wins for providing various levels of out-of-order memory access and memory dependence speculation (1.32 × SPECint2000) and for the addition of a second-level cache (another 1.60 × ). With careful microarchitecture and circuit design, we also achieve a L1 translation lookaside buffers and cache lookup with 29% less logic delay than the simpler Nios II/f memory system.
Henry Wong, Vaughn Betz, Jonathan Rose
ACM Trans. Reconfigurable Technol. Syst.2
2016 Power Analysis of Embedded NoCs on FPGAs and Comparison With Custom Buses
abstract
We propose embedding networks-on-chip (NoCs) on field-programmable gate-arrays (FPGAs) to implement system-level communication. Amongst other benefits, this can alleviate the current challenge of connecting the FPGA's fabric to high-speed I/O and memory interfaces, which are a crucial component of FPGA designs. Our mixed and hard embedded NoCs add only ~1% area to large FPGAs and can run much faster than the core logic, thus keeping up with the speed of I/O and memory interfaces. A detailed power analysis, per NoC component, shows that routers consume 14× less power when implemented hard compared with soft, and whether hard or soft most of the router's power is consumed in the input modules for buffering. For complete systems, hard NoCs consume <;6% (and as low as 3%) of the FPGA's dynamic power budget to support 100 GB/s of communication bandwidth. We find that, depending on design choices, hard NoCs consume 4.5-10.4 mJ of energy per gigabyte of data transferred. Surprisingly, this is comparable with the energy efficiency of the simplest traditional interconnect on an FPGA-soft point-to-point links require 4.7 mJ/GB. When comparing a hard NoC against soft buses that are currently used for interconnection, we find that a typical system is 4× smaller, and uses 23% less energy when implemented using the hard NoC even though it is only 43% utilized.
Mohamed S. Abdelfattah, Vaughn Betz
IEEE Trans. Very Large Scale Integr. Syst.2
2016 Multiple Dice Working as One: CAD Flows and Routing Architectures for Silicon Interposer FPGAs
abstract
Large field-programmable gate array (FPGA) systems with multiple dice connected by a silicon interposer are now commercially available. However, many questions remain concerning their key architecture parameters and efficiency, as the signal count between dice is reduced and the delay between the dice is increased compared with a monolithic FPGA. We modify the versatile place and route (VPR) to target interposer-based FPGAs and investigate placement and routing changes and incorporating partitioning into the flow to improve results. Our best computer-aided design (CAD) flow reduces the routing demand for interposer FPGAs with realistic connectivity between dice by 47% and improves the circuit speed by 13% on average. Architecture modifications to add routing flexibility when crossing the interposer are very beneficial and improve routability by a further 11%. With these CAD and architecture enhancements, we find that if an interposer supplies (between dice) 20% of the routing capacity that the normal (within-die) FPGA routing channels supply, there is only a modest impact on circuit routability. Smaller interposer-routing capacities do impact routability; however, minimum channel width increases by 70% when an interposer supplies only 10% of the within-die routing. The interposer also impacts delay, increasing circuit delay by 11% on average for a 1-ns interposer signal delay and a two-die system.
Ehsan Nasiri, Javeed Shaikh, André Hahn Pereira, Vaughn Betz
IEEE Trans. Very Large Scale Integr. Syst.4
2015 Take the Highway: Design for Embedded NoCs on FPGAs
abstract
We explore the addition of a fast embedded network-on-chip (NoC) to augment the FPGA's existing wires and switches, and help interconnect large applications. A flexible interface between the FPGA fabric and the embedded NoC allows modules of varying widths and frequencies to transport data over the NoC. We study both latency-insensitive and latency-sensitive design styles and present the constraints for implementing each type of communication on the embedded NoC. Our application case study with image compression shows that an embedded NoC improves frequency by 10-80%, reduces utilization of scarce long wires by 40% and makes design easier and more predictable. Additionally, we leverage the embedded NoC in creating a programmable Ethernet switch that can support up to 819 Gb/s on FPGAs.
Mohamed S. Abdelfattah, Andrew Bitar, Vaughn Betz
FPGA3
2015 Design and simulation tools for Embedded NOCs on FPGAs
abstract
We propose embedding hard NoCs on FPGAs to improve system-level communication as detailed in our previous studies [1-6]. This demo paper outlines the three main design and simulation tools that we have been using to experiment with Embedded NoCs on FPGAs.
Mohamed S. Abdelfattah, Andrew Bitar, Ange Yaghi, Vaughn Betz
FPL4
2015 Wotan: A tool for rapid evaluation of FPGA architecture routability without benchmarks
abstract
FPGA routing architectures consist of routing wires and programmable switches which together account for a significant portion of the fabric delay and area. Routing architectures have traditionally been evaluated using a full CAD flow with a suite of benchmark circuits. While the results of such a flow can be accurate, CAD tools are often tuned to a specific architecture type and can take a long time to run which prohibits quick exploration of different architectures early in the design process. In this paper we present an alternative approach that quickly estimates routability for a wide range of architectures without the use of benchmark circuits. Our new routability predictor first assigns congestion probabilities to the architecture's routing resources based on demand estimates found via efficient path enumeration through the routing graph. Next, we compute the probabilities of successfully routing different source/sink connections and finally we combine them to assign an overall routability score. We describe our predictor and present routability estimates for a range of 6-LUT and 4-LUT architectures, showing reasonable agreement with routability results from the full VPR CAD flow in much less CPU time.
Oleg Petelin, Vaughn Betz
FPL2
2015 Bringing programmability to the data plane: Packet processing with a NoC-enhanced FPGA
abstract
Modern computer networks need components that can evolve to support both the latest bandwidth demands and new protocols and features. To address this need, we propose a new programmable packet processor architecture built from an FPGA containing an embedded Network-on-Chip (NoC). The architecture is highly flexible, providing more programmability than is possible in an ASIC-based design, while supporting throughputs of 400 and 800 Gb/s. Additionally, we show that our design is 1.7× and 3.2× more area efficient, and achieves 1.5× and 3.7× lower latency than the best previously proposed FPGA-based packet processor on complex and simple applications, respectively. Lastly, we explore various ways a designer can take advantage of the flexibility available in this architecture.
Andrew Bitar, Mohamed S. Abdelfattah, Vaughn Betz
FPT3
2015 HETRIS: Adaptive floorplanning for heterogeneous FPGAs
abstract
Floorplanning is an approach to improve the scalability of existing CAD algorithms, facilitate team-based design, and also plays an important role in partial reconfiguration. This work introduces HETRIS, a new automated floorplanning tool for heterogeneous FPGAs. HETRIS uses an adaptive legality approach to target arbitrary FPGA architectures. It includes enhancements enabling it to run on average 15.6× faster than previous work, while producing denser floorplans than a commercial tool. Using HETRIS we perform the first evaluation of an FPGA floorplanner using real-world benchmarks, allowing us to investigate the relationship between partitioning, floorplanning and FPGA architecture.
Kevin E. Murray, Vaughn Betz
FPT2
2015 Robust Optimization of Multiple Timing Constraints
abstract
Modern field-programmable gate array (FPGA) circuit designs often contain multiple clocks and complex timing constraints, and achieving these constraints requires timing optimization at all stages of the computer-aided design (CAD) flow. To our knowledge, no prior published work has either described or quantitatively evaluated how to compute connection timing criticalities for circuits with multiple timing constraints in order to best guide CAD optimization algorithms. While single-clock techniques have a simple extension to multi-clock circuits, this formulation is not robust for circuits with multiple constraints of different magnitudes, or impossible constraints. We describe a robust method of timing optimization for circuits with multiple timing constraints, implemented in the open-source versatile place and route FPGA CAD tool. Our formulation can optimize multiple constraints well, even in the case where some constraints are impossible, and achieves over 20% greater clock speed with aggressive constraints than a straight-forward extension of single-clock work.
Michael Wainberg, Vaughn Betz
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2015 Timing-Driven Titan: Enabling Large Benchmarks and Exploring the Gap between Academic and Commercial CAD
abstract
Benchmarks play a key role in Field-Programmable Gate Array (FPGA) architecture and CAD research, enabling the quantitative comparison of tools and architectures. It is important that these benchmarks reflect modern large-scale systems that make use of heterogeneous resources; however, most current FPGA benchmarks are both small and simple. In this artile, we present Titan, a hybrid CAD flow that addresses these issues. The flow uses Altera’s Quartus II FPGA CAD software to perform HDL synthesis and a conversion tool to translate the result into the academic Berkeley Logic Interchange Format (BLIF). Using this flow, we created the Titan23 benchmark set, which consists of 23 large (90K--1.8M block) benchmark circuits covering a wide range of application domains. Using the Titan23 benchmarks and an enhanced model of Altera’s Stratix IV architecture, including a detailed timing model, we compare the performance and quality of VPR and Quartus II targeting the same architecture. We found that VPR is at least 2.8 × slower, uses 6.2 × more memory, 2.2 × more wire, and produces critical paths 1.5 × slower compared to Quartus II. Finally, we identified that VPR’s focus on achieving a dense packing and an inability to take apart clusters is responsible for a large portion of the wire length and critical path delay gap.
Kevin E. Murray, Scott Whitty, Suya Liu, Jason Luu, Vaughn Betz
ACM Trans. Reconfigurable Technol. Syst.5
2014 Efficient and programmable ethernet switching with a NoC-enhanced FPGA
abstract
Communications systems make heavy use of FPGAs; their programmability allows system designers to keep up with emerging protocols and their high-speed transceivers enable high bandwidth designs. While FPGAs are extensively used for packet parsing, inspection and classification, they have seen less use as the switch fabric between network ports. However, recent work has proposed embedding a network-on-chip (NoC) as a new "hard" resource on FPGAs and we show that by properly leveraging such a NoC one can create a very efficient yet still highly programmable network switch.
Andrew Bitar, Jeffrey Cassidy, Natalie D. Enright Jerger, Vaughn Betz
ANCS4
2014 Speeding Up FPGA Placement: Parallel Algorithms and Methods
abstract
Placement of a large FPGA design now commonly requires several hours, significantly hindering designer productivity. Furthermore, FPGA capacity is growing faster than CPU speed, which will further increase placement time unless new approaches are found. Multi-core processors are now ubiquitous, however, and some recent processors also have hardware support for transactional memory (TM), making parallelism an increasingly attractive approach for speeding up placement. We investigate methods to parallelize the simulated annealing placement algorithm in VPR, which is widely used in FPGA research. We explore both algorithmic changes and the use of different parallel programming paradigms and hardware, including TM, thread-level speculation (TLS) and lock-free techniques. We find that hardware TM enables large speedups (8.1x on average), but compromises “move fairness” and leads to an unacceptable quality loss. TLS scales poorly, with a maximum 2.2x speedup, but preserves quality. A new dependency checking parallel strategy achieves the best balance: the deterministic version achieves 5.9x speedup and no quality loss, while the non-deterministic, lock-free version can scale to a 34x speedup.
Matthew An, J. Gregory Steffan, Vaughn Betz
FCCM3
2014 Fast, Power-Efficient Biophotonic Simulations for Cancer Treatment Using FPGAs
abstract
Biophotonics, the study of light propagation through living tissue, is important for many medical applications ranging from imaging and detection through therapy for conditions such as cancer. Effective medical use of light depends on simulating its propagation through highly-scattering tissue. Monte Carlo simulation of photon migration has been adopted as the “gold standard” for its ability to capture complicated geometries and model all of the relevant problem physics. This accuracy and generality comes at a high computational cost, which limits the technique's utility. Greatly generalizing previous work, we present the first and only hardware-accelerated Monte Carlo biophotonic simulator that can accept complicated geometries described by tetrahedral meshes. Implemented on an Altera Stratix V FPGA, it achieves high performance (4x) and extremely high energy efficiency (67x) compared to a tightly-optimized multi-threaded CPU implementation, with demonstrated potential to expand the performance gains even further to 15-20x, which would enable important clinical and research applications.
Jeffrey Cassidy, Lothar Lilge, Vaughn Betz
FCCM3
2014 On Hard Adders and Carry Chains in FPGAs
abstract
Hardened adder and carry logic is widely used in commercial FPGAs to improve the efficiency of arithmetic functions. There are many design choices and complexities associated with such hardening, including circuit design, FPGA architectural choices, and the CAD flow. There has been very little study, however, on these choices and hence we explore a number of possibilities for hard adder design. We also highlight optimizations during front-end elaboration that help ameliorate the restrictions placed on logic synthesis by hardened arithmetic. We show that hard adders and carry chains, when used for simple adders, increase performance by a factor of four or more, but on larger benchmark designs that contain arithmetic, improve overall performance by roughly 15%. We measure an average area increase of 5% for architectures with carry chains but believe that better logic synthesis should reduce this penalty. Interestingly, we show that adding dedicated inter-logic-block carry links or fast carry look-ahead hardened adders result in only minor delay improvements for complete designs.
Jason Luu, Conor McCullough, Safeen Huda, Charles Chiasson, Kenneth B. Kent, Jason Helge Anderson, Jonathan Rose, Vaughn Betz
FCCM10
2014 Quantifying the cost and benefit of latency insensitive communication on FPGAs
abstract
Latency insensitive communication offers many potential benefits for FPGA designs, including easier timing closure by enabling automatic pipelining, and easier interfacing with embedded NoCs. However, it is important to understand the costs and trade-offs associated with any new design style. This paper presents optimized implementations of latency insensitive communication building blocks, quantifies their overheads in terms of area and frequency, and provides guidance to designers on how to generate high-speed and area-efficient latency insensitive systems.
Kevin E. Murray, Vaughn Betz
FPGA2
2014 Cad and routing architecture for interposer-based multi-FPGA systems
abstract
Interposer-based multi-FPGA systems are composed of multiple FPGA dice connected through a silicon interposer. Such devices allow larger FPGA systems to be built than one monolithic die can accomodate and are now commercially available. An open question, however, is how efficient such systems are compared to a monolithic FPGA, as the number of signals passing between dice is reduced and the signal delay between dice is increased in an interposer system vs. a monolithic FPGA.
André Hahn Pereira, Vaughn Betz
FPGA2
2014 Comparing performance, productivity and scalability of the TILT overlay processor to OpenCL HLS
abstract
High-Level-Synthesis (HLS) tools translate a software description of an application into custom FPGA logic, increasing designer productivity vs. Hardware Description Language (HDL) design flows. Overlays seek to further improve productivity by reducing application compile times and raising abstraction by enabling the designer to target a software-programmable substrate instead of the underlying FPGA. We compare the performance, development effort and scalability of two C-to-FPGA approaches: our TILT overlay processor and Altera's OpenCL HLS. Our application-customized TILT implementations of five data-parallel benchmarks have from 41 % to 80% of the throughput per unit of layout area achieved by our best OpenCL HLS designs. The time required for initial hardware compilation of these TILT designs and configuration of the target application onto the overlay is roughly comparable to the compile times of the OpenCL HLS designs: 28 and 103 minutes on average respectively. However subsequent reconfigurations due to changes in the application that do not require re-synthesis of the overlay are fast, taking 38 seconds on average. In contrast, OpenCL HLS applications require full recompilation after every code change. TILT also enables smaller, more area-efficient designs than OpenCL HLS when low to moderate throughput is sufficient. For high throughput, the larger spatially pipelined designs of OpenCL HLS are preferable.
Rafat Rashid, J. Gregory Steffan, Vaughn Betz
FPT3
2014 Networks-on-Chip for FPGAs: Hard, Soft or Mixed?
abstract
As FPGA capacity increases, a growing challenge is connecting ever-more components with the current low-level FPGA interconnect while keeping designers productive and on-chip communication efficient. We propose augmenting FPGAs with networks-on-chip (NoCs) to simplify design, and we show that this can be done while maintaining or even improving silicon efficiency. We compare the area and speed efficiency of each NoC component when implemented hard versus soft to explore the space and inform our design choices. We then build on this component-level analysis to architect hard NoCs and integrate them into the FPGA fabric; these NoCs are on average 20--23× smaller and 5--6× faster than soft NoCs. A 64-node hard NoC uses only ∼2% of an FPGA's silicon area and metallization. We introduce a new communication efficiency metric: silicon area required per realized communication bandwidth. Soft NoCs consume 4960 mm 2 /TBps, but hard NoCs are 84× more efficient at 59 mm 2 /TBps. Informed design can further reduce the area overhead of NoCs to 23 mm 2 /TBps, which is only 2.6× less efficient than the simplest point-to-point soft links (9 mm 2 /TBps). Despite this almost comparable efficiency, NoCs can switch data across the entire FPGA while point-to-point links are very limited in capability; therefore, hard NoCs are expected to improve FPGA efficiency for more complex styles of communication.
Mohamed S. Abdelfattah, Vaughn Betz
ACM Trans. Reconfigurable Technol. Syst.2
2014 VTR 7.0: Next Generation Architecture and CAD System for FPGAs
abstract
Exploring architectures for large, modern FPGAs requires sophisticated software that can model and target hypothetical devices. Furthermore, research into new CAD algorithms often requires a complete and open source baseline CAD flow. This article describes recent advances in the open source Verilog-to-Routing (VTR) CAD flow that enable further research in these areas. VTR now supports designs with multiple clocks in both timing analysis and optimization. Hard adder/carry logic can be included in an architecture in various ways and significantly improves the performance of arithmetic circuits. The flow now models energy consumption, an increasingly important concern. The speed and quality of the packing algorithms have been significantly improved. VTR can now generate a netlist of the final post-routed circuit which enables detailed simulation of a design for a variety of purposes. We also release new FPGA architecture files and models that are much closer to modern commercial architectures, enabling more realistic experiments. Finally, we show that while this version of VTR supports new and complex features, it has a 1.5× compile time speed-up for simple architectures and a 6× speed-up for complex architectures compared to the previous release, with no degradation to timing or wire-length quality.
Jason Luu, Jeffrey B. Goeders, Michael Wainberg, Andrew Somerville, Thien Yu, Konstantin Nasartschuk, Miad Nasr, Tim Liu, Nooruddin Ahmed, Kenneth B. Kent, Jason Helge Anderson, Jonathan Rose, Vaughn Betz
ACM Trans. Reconfigurable Technol. Syst.14
2014 Quantifying the Gap Between FPGA and Custom CMOS to Aid Microarchitectural Design
abstract
This paper compares the delay and area of a comprehensive set of processor building block circuits when implemented on custom CMOS and FPGA substrates, then uses these results to show how soft processor microarchitectures should be different from those of hard processors. We find that the ratios of the custom CMOS versus FPGA area for different building blocks varies considerably more than the speed ratios, thus, area ratios have more impact on microarchitecture choices. Complete processor cores on an FPGA use 17-27 × more area (“area ratio”) than the same design implemented in custom CMOS. Building blocks with dedicated hardware support on FPGAs such as SRAMs, adders, and multipliers are particularly area-efficient (2-7×), while multiplexers and content-addressable memories (CAM) are particularly area-inefficient (>100×). Applying these results, we find out-of-order soft processors should use physical register file organizations to minimize CAM size.
Henry Wong, Vaughn Betz, Jonathan Rose
IEEE Trans. Very Large Scale Integr. Syst.2
2013 Are FPGAs suffering from the innovator's dilemna?
abstract
FPGAs constitute a highly profitable industry, with approximately $5 billion of sales per year. High barriers to entry keep most companies away, and enable high profit margins for the incumbents. The industry has grown greatly over the years, but still constitutes a small portion of the overall semiconductor market. This raises the question to be addressed by this panel: is the FPGA community innovating as much as it should, or is a bias to maintain high profit margins and protect the cash flow of the current FPGA market holding us back from exploring new ideas and products that could greatly expand the appeal of and market for FPGA-related technology? This would be a classic case of the innovator's dilemma defined by Clayton Christensen: it is difficult for a company to engage in creative destruction of a cash cow product.
Vaughn Betz, Jason Cong
FPGA1
2013 The power of communication: Energy-efficient NOCS for FPGAS
abstract
Integrating networks-on-chip (NoCs) on FPGAs can improve device scalability and facilitate design by abstracting communication and simplifying timing closure, not only between modules in the FPGA fabric but also with large “hard” blocks such as high-speed I/O interfaces. We propose mixed and hard NoCs that add less than 1% area to large FPGAs and run 5-6 x faster than the soft NoC equivalent. A detailed power analysis, per NoC component, shows that routers consume 14 x less power when implemented hard compared to soft, and whether hard or soft most of the router's power is consumed in the input modules for buffering. For complete systems, hard NoCs consume less than 6% (and as low as 3%) of the FPGA's dynamic power budget to support 100 GB/s of communication bandwidth. We find that, depending on design choices, hard NoCs consume 4.5-10.4 mJ of energy per GB of data transferred. Surprisingly, this is comparable to the energy efficiency of the simplest traditional interconnect on an FPGA - soft point-to-point links require 4.7 mJ/GB. In many designs, communication must include multiplexing, arbitration and/or pipelining. For all these cases, our results indicate that a hard NoC will be more energy efficient than the conventional FPGA fabric.
Mohamed S. Abdelfattah, Vaughn Betz
FPL2
2013 Should FPGAS abandon the pass-gate?
abstract
Pass-transistors have been the key building block for field-programmable gate array (FPGA) circuitry for many years due to the very small switch they enable. However, passtransistor performance and reliability have been degrading with technology scaling. Transmission gates are an alternative to pass-transistors; while larger, they are more robust. We develop a new FPGA circuit optimization flow and use it to investigate the area, delay and power impact of building FPGAs out of transmission gates instead of pass-transistors in a 22nm process. Our results show that transmission gate FPGAs are 15% larger than pass-transistor FPGAs but are 10-25% faster depending on the allowable level of “gate boosting”. Without gate boosting, transmission gate FPGAs are the better option with 14% lower area-delay product. If 200mV of gate boosting is possible however, pass-transistor FPGAs remain the slightly better choice with a 2% better area-delay product. We also show that transmission gates with a separate power supply for their gate terminal enable a low-voltage FPGA with 50% less power and good delay.
Charles Chiasson, Vaughn Betz
FPL2
2013 Titan: Enabling large and complex benchmarks in academic CAD
abstract
Benchmarks play a key role in FPGA architecture and CAD research, enabling the quantitative comparison of tools and architectures. It is important that these benchmarks reflect modern designs which are large scale systems that make use of heterogeneous resources; however, most current FPGA benchmarks are both small and simple. In this paper we present Titan, a hybrid CAD flow that addresses these issues. The flow uses Altera's Quartus II FPGA CAD software to perform HDL synthesis and a conversion tool to translate the result into the academic BLIF format. Using this flow we created the Titan23 benchmark set, which consists of 23 large (90K-1.8M block) benchmark circuits covering a wide range of application domains. Using the Titan23 benchmarks and a detailed model of Altera's Stratix IV architecture we compared the performance and quality of VPR and Quartus II targeting the same architecture. We found that VPR is at least 2.7× slower, uses 5.1× more memory and 2.6× more wire compared to Quartus II. Finally, we identified that VPR's focus on achieving a dense packing is responsible for a large portion of the wire length gap.
Kevin E. Murray, Scott Whitty, Suya Liu, Jason Luu, Vaughn Betz
FPL5
2013 From Quartus to VPR: Converting HDL to BLIF with the Titan flow
abstract
Realistic benchmarks are important for FPGA Architecture and CAD evaluation. This paper provides a demo illustrating how designs described in HDL can be converted to BLIF using the Titan flow, and used in academic CAD tools.
Kevin E. Murray, Scott Whitty, Suya Liu, Jason Luu, Vaughn Betz
FPL5
2013 COFFE: Fully-automated transistor sizing for FPGAs
abstract
In this paper, we present COFFE (Circuit Optimization For FPGA Exploration), a new fully-automated transistor sizing tool for FPGAs. Automated transistor-level CAD tools are an important part of the architecture exploration flow because they provide accurate area and delay estimates of low-level FPGA circuitry, which must be obtained for each architecture. We show that modeling transistors as linear resistances and capacitances as has been done in previous FPGA transistor sizing tools is highly inaccurate for fine-grained transistor-level design in advanced process nodes. Therefore, COFFE's transistor sizing algorithm maintains circuit non-linearities by relying exclusively on HSPICE simulations to measure delay. Area is estimated with a transistor size-based model that incorporates a number of improvements to enhance its accuracy in advanced process technologies versus prior methods. In addition to more accurate area and delay estimation, COFFE considers more layout effects than prior published work by automatically accounting for transistor and wire loads, which are computed based on architectural parameters and layout area. This new FPGA transistor sizing tool requires only several hours to produce high-quality transistor sizing results for an entire FPGA tile; a task that would normally take months of manual effort. We demonstrate COFFE's utility in FPGA architecture studies by investigating an important new architectural question at the logic-to-routing interface.
Charles Chiasson, Vaughn Betz
FPT2
2013 Efficient methods for out-of-order load/store execution for high-performance soft processors
abstract
As FPGAs continue to increase in size, it becomes increasingly feasible and desirable to build higher performance soft processors. Preserving the familiar single-threaded programming model can be done with an out of order processor. The ability to execute memory loads and stores out of order has a large impact on performance, but this is difficult to do because the dependencies between stores and loads are not known until addresses are computed. Out of order memory disambiguation is traditionally done with CAMs in the load queue and store queue, but large CAMs are inefficient on FPGAs. Store Queue Index Prediction (SQIP) and NoSQ propose to replace CAMs with store-load forwarding prediction and load re-execution. We implement four memory disambiguation schemes (in-order, CAM, SQIP, NoSQ) on a Stratix IV FPGA and evaluate the area and delay trade-offs. We find that CAM area and delay degrade quickly with load/store queue size, while SQIP and NoSQ have little degradation with queue size but have area overhead for prediction and predictor training hardware. SQIP and NoSQ use less area than CAMs beyond 32 and 16 load/store queue entries, respectively, and have higher maximum frequency beyond 4 entries.
Henry Wong, Vaughn Betz, Jonathan Rose
FPT2
2012 Design tradeoffs for hard and soft FPGA-based Networks-on-Chip
abstract
Incorporating Networks-on-Chip (NoC) within FPGAs has the potential not only to improve the efficiency of the interconnect, but also to increase designer productivity and reduce compile time by raising the abstraction level of communication. By comparing NoC components on FPGAs and ASICs we quantify the efficiency gap between the two platforms and use the results to understand the design tradeoffs in that space. The crossbar has the largest FPGA vs. ASIC gaps: 85× area and 4.4× delay, while the input buffers have the smallest: 17× area and 2.9× delay. For a soft NoC router, these results indicate that wide datapaths, deep buffers and a small number of ports and virtual channels (VC) are favorable for FPGA implementation. If one hardens a complete state-of-the-art VC router it is on average 30× more area efficient and can achieve 3.6× the maximum frequency of a soft implementation. We show that this hard router can be integrated with the soft FPGA interconnect, and still achieve an area improvement of 22×. A 64-node NoC of hard routers with soft interconnect utilizes area equivalent to 1.6% of the logic modules in the latest FPGAs, compared to 33% for a soft NoC.
Mohamed S. Abdelfattah, Vaughn Betz
FPT2
2012 Portable and scalable FPGA-based acceleration of a direct linear system solver
abstract
FPGAs have the potential to serve as a platform for accelerating many computations including scientific applications. However, the large development cost and short life span for FPGA designs have limited their adoption by the scientific computing community. FPGA-based scientific computing and many kinds of embedded computing could become more practical if there were hardware libraries that were portable to any FPGA-based system with performance that scaled with the size of the FPGA. To illustrate this idea we have implemented one common super-computing library function: the LU factorization method for solving systems of linear equations. This paper describes a method for making the design both portable and scalable that should be illustrative if such libraries are to be built in the future. The design is a software-based generator that leverages both the flexibility of a software programming language and the parameters inherent in an hardware description language. The generator accepts parameters that describe the FPGA capacity and external memory capabilities. We compare the performance of our engine executing on the largest FPGA available at the time of this work (an Altera Stratix III 3S340) to a single processor core fabricated in the same 65nm IC process running a highly optimized software implementation from the processor vendor. For single precision matrices on the order of 10,000 × 10,000 elements, the FPGA implementation is 2.2 times faster and the energy dissipated per useful GFLOP operation is a factor of 5 times less. For double precision, the FPGA implementation is 1.7 times faster and 3.5 times more energy efficient.
Wei Zhang 0222, Vaughn Betz, Jonathan Rose
ACM Trans. Reconfigurable Technol. Syst.2
2011 Comparing FPGA vs. custom cmos and the impact on processor microarchitecture
abstract
As soft processors are increasingly used in diverse applications, there is a need to evolve their microarchitectures in a way that suits the FPGA implementation substrate. This paper compares the delay and area of a comprehensive set of processor building block circuits when implemented on custom CMOS and FPGA substrates. We then use the results of these comparisons to infer how the microarchitecture of soft processors on FPGAs should be different from hard processors on custom CMOS.
Henry Wong, Vaughn Betz, Jonathan Rose
FPGA2
2011 Efficient and Deterministic Parallel Placement for FPGAs
abstract
We describe a parallel simulated annealing algorithm for FPGA placement. The algorithm proposes and evaluates multiple moves in parallel, and has been incorporated into Altera’s Quartus II CAD system. Across a set of 18 industrial benchmark circuits, we achieve geometric average speedups during the quench of 2.7x and 4.0x on four and eight processors, respectively, with individual circuits achieving speedups of up to 3.6x and 5.9x. Over the course of the entire anneal, we achieve speedups of up to 2.8x and 3.7x, with geometric average speedups of 2.1x and 2.4x. Our algorithm is the first parallel placer to optimize for criteria other than wirelength, such as critical path length, and is one of the few deterministic parallel placement algorithms. We discuss the challenges involved in combining these two features and the new techniques we used to overcome them. We also quantify the impact of maintaining determinism on eight cores, and find that while it reduces performance by approximately 15% relative to an ideal speedup of 8.0x, hardware limitations are a larger factor and reduce performance by 30--40%. We then suggest possible enhancements to allow our approach to scale to 16 cores and beyond.
Adrian Ludwin, Vaughn Betz
ACM Trans. Design Autom. Electr. Syst.2
2010 A comprehensive approach to modeling, characterizing and optimizing for metastability in FPGAs
abstract
Metastability is a phenomenon that can cause system failures in digital circuits. It may occur whenever signals are being transmitted across asynchronous or unrelated clock domains. The impact of metastability is increasing as process geometries shrink and supply voltages drop faster than transistor Vts. FPGA technologies are significantly affected since leading edge FPGAs are amongst the first devices to adopt the most recent process nodes. In this paper, we present a comprehensive suite of techniques for modeling, characterizing and optimizing metastability effects in FPGAs. We first discuss a theoretical model of metastability, and verify the predictions using both circuit level simulations and board measurements. Next we show how designers have traditionally dealt with metastability problems and contrast that with the automatic CAD algorithms described in this paper that both analyze and optimize metastability-related issues. Through our detailed experimental results, we show that we can improve the metastability characteristics of a large suite of industrial benchmarks by an average of 268,000 times with our optimization techniques.
Doris Chen, Deshanand P. Singh, Jeffrey Chromczak, David M. Lewis, Ryan Fung, David Neto, Vaughn Betz
FPGA7
2009 FPGA challenges and opportunities at 40nm and beyond
abstract
FPGA companies are amongst the earliest adopters of next-generation process technology. This involves many challenges, including power management, device modeling, increasing I/O performance to match the computational capacity, and enabling very large designs to be completed quickly. Process scaling increases FPGA capacity and allows new features, such as high-performance I/O protocols (e.g. PCI Express). Scaling favours FPGAs over competing technologies, as fewer and fewer ASICs have the volumes to justify the cost of a design in cutting-edge technology. I will give an overview of the challenges in designing a cutting-edge FPGA, and describe the solutions Altera has adopted in its 40 nm Stratix IV FPGAs. I/O bandwidth is crucial. While the density of FPGAs is increasing rapidly, the number of I/O pins is not-we need to move more bits through the same number of pins. Another challenge is to model the timing, power and signal integrity of increasingly complex FPGAs in increasingly variable processes, while still keeping the tools easy to use. Managing power is a third challenge. Each process generation roughly doubles the number of transistors per die and tends to increase leakage power. The power budget per FPGA is roughly constant, so we need to innovate to control power. Ever larger FPGAs also necessitate innovation to keep designers productive. We must both keep the compile time of traditional FPGA CAD tools reasonable, and develop new tools that allow designers to create and verify more complex systems in the same time.
Vaughn Betz
FPL1
2008 High-quality, deterministic parallel placement for FPGAs on commodity hardware
abstract
In this paper, we describe the application of two parallelization strategies to the Quartus II FPGA placer. The first uses a pipelining approach and achieves speedups of 1.3x on two processing cores. The second uses a parallel moves approach and achieves speedups of 2.2x on four cores. Unlike all previous parallel moves algorithms, ours is deterministic and always gives the same answer as the serial version of the algorithm, without any significant reduction in performance.
Adrian Ludwin, Vaughn Betz, Ketan Padalia
FPGA2
2008 Portable and scalable FPGA-based acceleration of a direct linear system solver
abstract
FPGAs are becoming an attractive platform for accelerating many computations including scientific applications. However, their adoption has been limited by the large development cost and short life span of FPGA designs. We believe that FPGA-based scientific computation would become far more practical if there were hardware libraries that were portable to any FPGA with performance that could scale with the resources of the FPGA. To illustrate this idea we have implemented one common supercomputing library function: the LU factorization method for solving linear systems. This paper discusses issues in making the design both portable and scalable. The design is automatically generated to match the FPGA’s capabilities and external memory through the use of parameters. We compared the performance of the design on the FPGA to a single processor core and found that it performs 2.2 times faster, and that the energy dissipated per computation is a factor 5 times less.
Wei Zhang 0222, Vaughn Betz, Jonathan Rose
FPT2
2008 Slack Allocation and Routing to Improve FPGA Timing While Repairing Short-Path Violations
abstract
Abstract—This work presents the first published algorithm to simultaneously optimize both short- and long-path timing constraints in a Field-Programmable Gate Array (FPGA): the Routing Cost Valleys (RCV) algorithm. RCV consists of two components: a new slack allocation algorithm that determines both a minimum and a maximum delay budget for each circuit connection, and a new router that strives to meet and, if possible, surpass these connection delay constraints. RCV improves both long-path and short-path timing slack significantly versus an earlier Computer-Aided Design (CAD) system, showing the importance of an integrated approach that simultaneously optimizes both types of timing constraints. It is able to meet longpath and short-path timing on all 157 Peripheral Component Interconnect (PCI) cores tested, while an earlier algorithm failed to achieve timing on 75 % of the cores. Even in cases where there are no short-path timing constraints, RCV outperforms a stateof-the-art FPGA router and improves the maximum clock speed of circuits by an average of 3.2 % (and up to 24.7%). L Index Terms—FPGA, routing, slack allocation, timing
Ryan Fung, Vaughn Betz, William Chow
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2007 Power-Efficient RAM Mapping Algorithms for FPGA Embedded Memory Blocks
abstract
Contemporary field-programmable gate array (FPGA) design requires a spectrum of available physical resources. As FPGA logic capacity has grown, locally accessed FPGA embedded memory blocks have increased in importance. When targeting FPGAs, application designers often specify high-level memory functions, which exhibit a range of sizes and control structures. These logical memories must be mapped to FPGA embedded memory resources such that physical design objectives are met. In this paper, a set of power-efficient logical-to-physical RAM mapping algorithms is described, which converts user-defined memory specifications to on-chip FPGA memory block resources. These algorithms minimize RAM dynamic power by evaluating a range of possible embedded memory block mappings and selecting the most power-efficient choice. Our automated approach has been validated with both simulation of power dissipation and measurements of power dissipation on FPGA hardware. A comparison of measured power reductions to values determined via simulation confirms the accuracy of our simulation approach. Our power-aware RAM mapping algorithms have been integrated into a commercial FPGA compiler and tested with 34 large FPGA benchmarks. Through experimentation, we show that, on average, embedded memory dynamic power can be reduced by 26% and overall core dynamic power can be reduced by 6% with a minimal loss (1%) in design performance. In addition, it is shown that the availability of multiple embedded memory block sizes in an FPGA reduces embedded memory dynamic power by an additional 9.6% by giving more choices to the computer-aided design algorithms
Russell Tessier, Vaughn Betz, David Neto, Aaron Egier, Thiagaraja Gopalsamy
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2006 Power-aware RAM mapping for FPGA embedded memory blocks
abstract
Embedded memory blocks are important resources in contemporary FPGA devices. When targeting FPGAs, application designers often specify high-level memory functions which exhibit a range of sizes and control structures. These logical memories must be mapped to FPGA embedded memory resources such that physical design objectives are met. In this work a set of power-aware logical-to-physical RAM mapping algorithms are described which convert user-defined memory specifications to on-chip FPGA memory block resources. These algorithms minimize RAM dynamic power by evaluating a range of possible embedded memory block mappings and selecting the most power-efficient choice. Our automated approach has been integrated into a commercial FPGA compiler and tested with 40 large FPGA benchmarks. Through experimentation, we show that, on average, embedded memory dynamic power can be reduced by 21% and overall core dynamic power can be reduced by 7% with a minimal loss (1%) in design performance.
Russell Tessier, Vaughn Betz, David Neto, Thiagaraja Gopalsamy
FPGA2
2005 The Stratix II logic and routing architecture
abstract
This paper describes the Altera Stratix II™ logic and routing architecture. This architecture features a novel adaptive logic module (ALM) that is based on a 6-LUT, but can be partitioned into two smaller LUTs to efficiently implement circuits containing a range of LUT sizes that arises in conventional synthesis flows. This provides a performance increase of 15% in the Stratix II architecture while reducing area by 2%. The ALM also includes a more powerful arithmetic structure that can perform two bits of arithmetic per ALM, and perform a sum of up to three inputs. The routing fabric adds a new set of fast inputs to the routing multiplexers for another 3% improvement in performance, while other improvements in routing efficiency cause another 6% reduction in area. These changes in combination with other circuit and architecture changes in Stratix II contribute 27% of an overall 51% performance improvement (including architecture and process improvement). The architecture changes reduce area by 10% in the same process, and by 50% after including process migration.
David M. Lewis, Elias Ahmed, Gregg Baeckler, Vaughn Betz, Mark Bourgeault, David Cashman, David R. Galloway, Mike Hutton, Christopher Lane, Andy Lee, Paul Leventis, Sandy Marquardt, Cameron McClintock, Ketan Padalia, Bruce Pedersen, Giles Powell, Boris Ratchev, Srinivas Reddy, Jay Schleicher, Kevin Stevens, Richard Yuan, Richard Cliff, Jonathan Rose
FPGA4
2004 Simultaneous short-path and long-path timing optimization for FPGAs
abstract
This work presents the routing cost valleys (RCV) algorithm - the first published algorithm that simultaneously optimizes all short- and long-path timing constraints in a field-programmable gate array (FPGA). RCV is comprised of a new slack allocation algorithm that produces both minimum and maximum delay budgets for each circuit connection, and a new router that strives to meet and, if possible, surpass these connection delay constraints. RCV achieves excellent results. On a set of 100 large circuits, RCV improves both long-path and short-path timing slack significantly vs. an earlier computer-aided design (CAD) system that focuses solely on long-path timing. Even with no short-path timing constraints, RCV improves the clock speed of circuits by 3.9% on average. Finally, RCV is able to meet timing on all 72 peripheral component interconnect (PCI) cores tested, while an earlier algorithm fails to achieve timing on all 72 cores.
Ryan Fung, Vaughn Betz, William Chow
ICCAD2
2003 The StratixTM routing and logic architecture
abstract
This paper describes the Altera Stratix logic and routing architecture. The primary goals of the architecture were to achieve high performance and logic density. We give an overview of the entire device, and then focus on the logic and routing architecture. The Stratix logic architecture is based on a cluster of ten 4-input LUTs and its routing consists of staggered routing lines. We describe the development of the routing architecture, including its directional bias, its direct-drive routing which reduces both area and delay. The logic array block and logic cell design is also described, and new routing structures with in the logic array block, and logic element features are described.
David M. Lewis, Vaughn Betz, David Jefferson, Andy Lee, Christopher Lane, Paul Leventis, Sandy Marquardt, Cameron McClintock, Bruce Pedersen, Giles Powell, Srinivas Reddy, Chris Wysocki, Richard Cliff, Jonathan Rose
FPGA2
2000 Automatic generation of FPGA routing architectures from high-level descriptions
abstract
In this paper we present a “high-level” FPGA architecture description language which lets FPGA architects succinctly and quickly describe an FPGA routing architecture. We then present an “architecture generator” built into the VPR CAD tool [1, 2] that converts this high-level architecture description into a detailed and completely specified flat FPGA architecture. This flat architecture is the representation with which CAD optimization and visualization modules typically work. By allowing FPGA researchers to specify an architecture at a high-level, an architecture generator enables quick and easy “what-if” experimentation with a wide range of FPGA architectures. The net effect is a more fully optimized final FPGA architecture. In contrast, when FPGA architects are forced to use more traditional methods of describing an FPGA (such as the manual specification of every switch in the basic file of the FPGA), far less experimentation can be performed in the same time, and the architectures experimented upon are likely to be highly similar, leaving important parts of the design space completely unexplored.
Vaughn Betz, Jonathan Rose
FPGA1
2000 Timing-driven placement for FPGAs
abstract
In this paper we introduce a new Simulated Annealing-based timing-driven placement algorithm for FPGAs. This paper has three main contributions. First, our algorithm employs a novel method of determining source-sink connection delays during placement. Second, we introduce a new cost function that trades off between wire-use and critical path delay, resulting in significant reductions in critical path delay without significant increases in wire-use. Finally, we combine connection-based and path-based timing-analysis to obtain an algorithm that has the low time-complexity of connection-based timing-driven placement, while obtaining the quality of path-based timing-driven placement.
Alexander Marquardt, Vaughn Betz, Jonathan Rose
FPGA2
2000 Speed and area tradeoffs in cluster-based FPGA architectures
abstract
One way to reduce the delay and area of field-programmable gate arrays (FPGAs) is to employ logic-cluster-based architectures, where a logic cluster is a group of logic elements connected with high-speed local interconnections. In this paper, we empirically evaluate FPGA architectures with logic clusters ranging in size from 1 to 20, and show that compared to architectures with size 1 clusters, architectures with size 8 clusters have 23% less delay (30% faster clock speed) and require 14% less area. We also show that FPGA architectures with large cluster sizes can significantly reduce design compile time-an increasingly important concern as the logic capacity of FPGA's rises. For example, an architecture that uses size 20 clusters requires seven times less compile time than an architecture with size 1 clusters.
Alexander Marquardt, Vaughn Betz, Jonathan Rose
IEEE Trans. Very Large Scale Integr. Syst.2
1999 FPGA Routing Architecture: Segmentation and Buffering to Optimize Speed and Density
abstract
In this work we investigate the routing architecture of FPGAs, focusing primarily on determining the best distribution of routing segment lengths and the best mix of pass transistor and tri-state buffer routing switches. While most commercial FPGAs contain many length 1 wires (wires that span only one logic block) we find that wires this short lead to FPGAs that are inferior in terms of both delay and routing area. Our results show instead that it is best for FPGA routing segments to have lengths of 4 to 8 logic blocks. We also show that 50% to 80% of the routing switches in an FPGA should be pass transistors, with the remainder being tri-state buffers. Architectures that employ the best segmentation distributions and the best mixes of pass transistor and tri-state buffer switches found in this paper are not only 11% to 18% faster than a routing architecture very similar to that of the Xilinx XC4000X but also considerably simpler. These results are obtained using an architecture investigation infrastructure that contains a fully timing-driven router and detailed area and delay models.
Vaughn Betz, Jonathan Rose
FPGA1
1999 Using Cluster-Based Logic Blocks and Timing-Driven Packing to Improve FPGA Speed and Density
abstract
In this papel; we investigate the speed and area-eficiency of FPGAs employing "logic clusters" containing multiple LUTs and registers as their logic block.We introduce a new, timing-driven tool (T-VPack) to "pack" LUTs and registers into these logic clusters, and we show that this algorithm is superior to an existing packing algorithm.Then, using a realistic routing architecture and sophisticated delay and area models, we empirically evaluate FPGAs composed of clusters ranging in size from one to twenty LUTs, and show that clusters of size seven through ten provide the best area-delay trade-o@ Compared to circuits implemented in an FPGA composed of size one clusters, circuits implemented in an FPGA with size seven clusters have 30% less delay (a 43% increase in speed) and require 8% less area, and circuits implemented in an FPGA with size ten clusters have 34% less delay (a 52% increase in speed), and require no additional area.
Alexander Marquardt, Vaughn Betz, Jonathan Rose
FPGA2
1998 A Fast Routability-Driven Router for FPGAs
abstract
Three factors are driving the demand for rapid FPGA compilation. First, as FPGAs have grown in logic capacity, the compile computation has grown more quickly than the compute power of the available computers. Second, there exists a subset of users who are willing to pay for very high speed compile with a decrease in quality of result, and accordingly being required to use a larger FPGA or use more real-estate on a given FPGA than is otherwise necessary. Third, very high speed compile has been a long-standing desire of those using FPGA-based custom computing machines, as they want compile times at least closer to those of regular computers. This paper focuses on the routing phase of the compile process, and in particular on routability-driven routing (as opposed to timing-driven routing). We present a routing algorithm and routing tool that has three unique capabilities relating to very high-speed compile: 1. For a “low stress ” routing problem (which we define as the case where the track supply is at least 10 % greater than the minimum number of tracks per channel actually needed to route a circuit) the routing time is very fast. For example, the routing phase (after the netlist is parsed and the routing graph is constructed) for a 20,000 LUT/FF pair circuit with 30 % extra tracks is only 23 seconds on a 300 MHz Sparcstation. 2. For low-stress routing problems the routing time is nearlinear in the size of the circuit, and the linearity constant is very small: 1.1 ms per LUT/FF pair, or roughly 55,000 LUT/FF pairs per minute. 3. For more difficult routing problems (where the track supply is close to the minimum needed) we provide a method that quickly identifies and subdivides this class into two sub-classes: (i) those circuits which are difficult (but possible) to route and will take significantly more time than low-stress problems, and (ii) those circuits which are impossible to route. In the first case the user can choose to continue or reduce the amount of logic; in the second case the user is forced to reduce the amount of logic or obtain a larger FPGA. 1.
Jordan S. Swartz, Vaughn Betz, Jonathan Rose
FPGA2
1998 Effect of the prefabricated routing track distribution on FPGA area-efficiency
abstract
In most commercial field programmable gate arrays (FPGA's) the number of wiring tracks in each channel is the same across the entire chip. A long-standing open question for both FPGA's and channeled gate arrays is whether or not some uneven distribution of routing tracks across the chip would lead to an area benefit. For example, many circuit designers intuitively believe that most congestion occurs near the center of a chip, and hence expect that having wider routing channels near the chip center would be beneficial. In this paper, we determine the relative area-efficiency of several different routing track distributions. We first investigate FPGA's in which horizontal and vertical channels contain different numbers of tracks in order to determine if such a directional bias provides a density advantage. Second, we examine routing track distributions in which the track capacities vary from channel to channel. We compare the area efficiency of these nonuniform routing architectures to that of an FPGA with uniform channel capacities across the entire chip. The main result is that the most area-efficient global routing architecture is one with uniform (or very nearly uniform) channel capacities across the entire chip in both the horizontal and vertical directions. This paper shows why this result, which is contrary to the intuition of many FPGA architects, is true. While a uniform routing architecture is the most area-efficient, several nonuniform and directionally biased architectures are fairly area-efficient provided that appropriate choices are made for the pin positions on the logic blocks and the logic block array aspect ratio.
Vaughn Betz, Jonathan Rose
IEEE Trans. Very Large Scale Integr. Syst.1
1996 Directional bias and non-uniformity in FPGA global routing architectures
abstract
We investigate the effect of the prefabricated routing track distribution on the area-efficiency of FPGAs. The first question we address is whether horizontal and vertical channels should contain the same number of tracks (capacity), or if there is a density advantage with a directional bias. Secondly, should the channels have a uniform capacity, or is there an advantage when capacities vary from channel to channel? The key result is that the most area-efficient global routing architecture is one with uniform (or very nearly uniform) channel capacities across the entire chip in both the horizontal and vertical directions. Several non-uniform and directionally-biased architectures, however are fairly area-efficient provided that appropriate choices are made for the pin positions on the logic blocks and the logic array aspect ratio.
Vaughn Betz, Jonathan Rose
ICCAD1
1995 Using Architectural "Families" to Increase FPGA Speed and Density
abstract
In order to narrow the speed and density gap between FPGAs and MPGAs we propose the development of “families” of FPGAs. Each FPGA family is targeted at a single maximum logic capacity, and consists of several “siblings”, or FPGAs of different yet complementary architectures. Any given application circuit is implemented in the sibling with the most appropriate architecture. With properly chosen siblings, one can develop a family of FPGAs which will have better speed and density than any single FPGA. We apply this concept to create two different FPGA families, one composed of architectures with different types of hard-wired logic blocks and the other created from architectures with different types of heterogeneous logic blocks. We found that a family composed of eight chips with different hard-wired logic block architectures simultaneously improves density by 12 to 14% and speed by 18 to 20% over the best single hard-wired FPGA.
Vaughn Betz, Jonathan Rose
FPGA1