EDBT 2026 Demo / reviewers in the wild / expert
Jonathan Rose
dblp:r/JonathanRose
· DBLP profile ↗
116ranked-venue papers
14as first author
1since 2021 · last 2024
0000-0002-3551-2175ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 110 · 13 first-authorArtificial intelligence and machine learning · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
61 papers |
Electronic design automation · 55% Reconfigurable computing and FPGAs · 36% Processor architecture and microarchitecture · 4% | |
| Artificial intelligence
2 papers |
3D vision · 60% Face, body and person analysis · 20% Image recognition and object detection · 20% |
Topics — the 30 heaviest of 87, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation
physical design |
1.1 | 28 | 2014 | Towards interconnect-adaptive packing for FPGAs · FPGA 2014 The VTR project: architecture and CAD for FPGAs from verilog to routing · FPGA 2012 Architecture description and packing for logic blocks with hierarchy, modes and complex interconnect · FPGA 2011 |
Reconfigurable computing and FPGAs
FPGA architecture |
0.7 | 18 | 2014 | Towards interconnect-adaptive packing for FPGAs · FPGA 2014 Architecture description and packing for logic blocks with hierarchy, modes and complex interconnect · FPGA 2011 VPR 5.0: FPGA cad and architecture exploration tools with single-driver routing, heterogeneity and process scaling · FPGA 2009 |
Electronic design automation › physical design
interconnect synthesis |
0.5 | 2 | 2017 | Synchronization Constraints for Interconnect Synthesis · FPGA 2017 Fine-Grained Interconnect Synthesis · FPGA 2015 |
Reconfigurable computing and FPGAs
FPGA routing architecture |
0.4 | 10 | 2017 | Synchronization Constraints for Interconnect Synthesis · FPGA 2017 Fine-Grained Interconnect Synthesis · FPGA 2015 Using bus-based connections to improve field-programmable gate array density for implementing datapath circuits · FPGA 2005 |
Electronic design automation › physical design
packing |
0.3 | 3 | 2014 | Towards interconnect-adaptive packing for FPGAs · FPGA 2014 Architecture description and packing for logic blocks with hierarchy, modes and complex interconnect · FPGA 2011 Using Cluster-Based Logic Blocks and Timing-Driven Packing to Improve FPGA Speed and Density · FPGA 1999 |
Electronic design automation
logic synthesis |
0.3 | 8 | 2012 | The VTR project: architecture and CAD for FPGAs from verilog to routing · FPGA 2012 A synthesis oriented omniscient manual editor · FPGA 2004 Automatic transistor and physical design of FPGA tiles from an architectural specification · FPGA 2003 |
Electronic design automation
synchronization constraints |
0.3 | 1 | 2017 | Synchronization Constraints for Interconnect Synthesis · FPGA 2017 |
Reconfigurable computing and FPGAs › FPGA architecture
FPGA architecture exploration |
0.2 | 3 | 2012 | The VTR project: architecture and CAD for FPGAs from verilog to routing · FPGA 2012 Automated transistor sizing for FPGA architecture exploration · DAC 2008 TEMPT: Technology Mapping for the Exploration of FPGA Architectures with Hard-Wired Connections · DAC 1992 |
Reconfigurable computing and FPGAs
FPGA vs ASIC comparison |
0.2 | 3 | 2007 | Measuring the Gap Between FPGAs and ASICs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 Measuring the gap between FPGAs and ASICs · FPGA 2006 Panel: (When) Will FPGAs Kill ASICs? · DAC 2001 |
Reconfigurable computing and FPGAs › FPGA architecture
heterogeneous logic blocks |
0.1 | 2 | 2011 | Architecture description and packing for logic blocks with hierarchy, modes and complex interconnect · FPGA 2011 Using Architectural "Families" to Increase FPGA Speed and Density · FPGA 1995 |
Reconfigurable computing and FPGAs › FPGA architecture
FPGA architecture design |
0.1 | 2 | 2008 | Modeling routing demand for early-stage FPGA architecture development · FPGA 2008 Using bus-based connections to improve field-programmable gate array density for implementing datapath circuits · FPGA 2005 |
Performance modeling and evaluation
benchmarking |
0.1 | 2 | 2007 | Measuring the Gap Between FPGAs and ASICs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 Measuring the gap between FPGAs and ASICs · FPGA 2006 |
Reconfigurable computing and FPGAs › FPGA-based processor implementation
soft-core processor |
0.1 | 2 | 2007 | Exploration and Customization of FPGA-Based Soft Processors · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 Application-specific customization of soft processor microarchitecture · FPGA 2006 |
Electronic design automation
design space exploration |
0.1 | 2 | 2009 | VPR 5.0: FPGA cad and architecture exploration tools with single-driver routing, heterogeneity and process scaling · FPGA 2009 Design, layout and verification of an FPGA using automated tools · FPGA 2005 |
Electronic design automation
high-level synthesis |
0.1 | 2 | 2015 | Fine-Grained Interconnect Synthesis · FPGA 2015 The VTR project: architecture and CAD for FPGAs from verilog to routing · FPGA 2012 |
Electronic design automation › design automation tools › FPGA CAD
FPGA design tools |
0.1 | 1 | 2009 | VPR 5.0: FPGA cad and architecture exploration tools with single-driver routing, heterogeneity and process scaling · FPGA 2009 |
Reconfigurable computing and FPGAs › reconfigurable architecture › reconfigurable processor
soft vector processor |
0.1 | 1 | 2009 | Soft vector processors vs FPGA custom hardware: measuring and reducing the gap · FPGA 2009 |
Processor architecture and microarchitecture
vector processing |
0.1 | 1 | 2009 | Soft vector processors vs FPGA custom hardware: measuring and reducing the gap · FPGA 2009 |
Electronic design automation › benchmark generation
synthetic circuit generation |
0.1 | 2 | 2004 | Synthetic circuit generation using clustering and iteration · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004 Synthetic circuit generation using clustering and iteration · FPGA 2003 |
Electronic design automation › multi-objective optimization
area-time tradeoff |
0.1 | 1 | 2008 | Area and delay trade-offs in the circuit and architecture design of FPGAs · FPGA 2008 |
Integrated circuit design
circuit design |
0.1 | 1 | 2008 | Area and delay trade-offs in the circuit and architecture design of FPGAs · FPGA 2008 |
Electronic design automation
interconnect modeling |
0.1 | 1 | 2008 | Modeling routing demand for early-stage FPGA architecture development · FPGA 2008 |
Electronic design automation › circuit sizing
transistor sizing |
0.1 | 1 | 2008 | Automated transistor sizing for FPGA architecture exploration · DAC 2008 |
Reconfigurable computing and FPGAs
FPGA design flow |
0.1 | 2 | 2014 | Towards interconnect-adaptive packing for FPGAs · FPGA 2014 Constraints from Hell: How to Tell Makes a Good FPGA (Panel) · FPGA 1998 |
Electronic design automation
benchmark generation |
0.1 | 3 | 2002 | Automatic generation of synthetic sequential benchmark circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2002 Characterization and parameterized generation of synthetic combinational benchmark circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1998 Generation of Synthetic Sequential Benchmark Circuits · FPGA 1997 |
Electronic design automation › design methodology
latency-insensitive design |
0.1 | 1 | 2015 | Fine-Grained Interconnect Synthesis · FPGA 2015 |
Electronic design automation › hardware verification and test
hardware verification |
0.1 | 2 | 2005 | Design, layout and verification of an FPGA using automated tools · FPGA 2005 Logic Emulation: A Niche or a Future Standard for Design Verification? (Panel Abstract) · DAC 1993 |
Electronic design automation › benchmark generation
benchmark circuit generation |
0.1 | 2 | 2004 | Synthetic circuit generation using clustering and iteration · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004 Synthetic circuit generation using clustering and iteration · FPGA 2003 |
Reconfigurable computing and FPGAs › FPGA architecture
adaptive logic module |
0.1 | 1 | 2005 | The Stratix II logic and routing architecture · FPGA 2005 |
Electronic design automation › hardware verification and test › functional verification
FPGA verification |
0.1 | 1 | 2005 | Design, layout and verification of an FPGA using automated tools · FPGA 2005 |
Methods — techniques the papers use, named apart from their topics
deterministic latency synthesis · 0.2arbitration avoidance · 0.2speculative packing · 0.2pre-packing · 0.2interconnect-aware pin counting · 0.2timing-driven placement and routing · 0.1hard-block synthesis · 0.1automatic soft processor generation · 0.1circuit delay and area comparison · 0.1area-driven packing · 0.1multiresolution · 0.0multi-orientation · 0.0local weighted phase-correlation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Generation, Distillation and Evaluation of Motivational Interviewing-Style Reflections with a Foundational Language ModelabstractAndrew Brown, Jiading Zhu, Mohamed Abdelwahab, Alec Dong, Cindy Wang, Jonathan Rose. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Jiading Zhu, Mohamed Abdelwahab, Alec Dong, Cindy Wang, Jonathan Rose |
EACL (1) | 6 |
| 2020 | Detection and Correspondence Matching of Corneal Reflections for Eye Tracking Using Deep LearningabstractEye tracking systems that estimate the point-of-gaze are essential in extended reality (XR) systems as they enable new interaction paradigms and technological improvements. It is important for these systems to maintain accuracy when the headset moves relative to the head (known as device slippage) due to head movements or user adjustment. One of the most accurate eye tracking techniques, which is also insensitive to shifts of the system relative to the head, uses two or more infrared (IR) light emitting diodes to illuminate the eye and an IR camera to capture images of the eye. An essential step in estimating the point-of-gaze in these systems is the precise determination of the location of two or more corneal reflections (virtual images of the IR-LEDs that illuminate the eye) in images of the eye. Eye trackers tend to have multiple light sources to ensure at least one pair of reflections for each gaze position. The use of multiple light sources introduces a difficult problem: the need to match the corneal reflections with the corresponding light source over the range of expected eye movements. Corneal reflection detection and matching often fail in XR systems due to the proximity of camera and steep illumination angles of light sources with respect to the eye. The failures are caused by corneal reflections having varying shape and intensity levels or disappearance due to rotation of the eye, or the presence of spurious reflections. We have developed a fully convolutional neural network, based on the UNET architecture, that solves the detection and matching problem in the presence of spurious and missing reflections. Eye images of 25 people were collected in a virtual reality headset using a binocular eye tracking module consisting of five infrared light sources per eye. A set of 4,000 eye images were manually labelled for each of the corneal reflections, and data augmentation was used to generate a dataset of 40,000 images. The network is able to correctly identify and match 91% of corneal reflections present in the test set. This is comparable to a state-of-the-art deep learning system, but our approach requires 33 times less memory and executes 10 times faster. The proposed algorithm, when used in an eye tracker in a VR system, achieved an average mean absolute gaze error of 1°. This is a significant improvement over the state-of-the-art learning-based XR eye tracking systems that have reported gaze errors of 2-3°. Soumil Chugh, Braiden Brousseau, Jonathan Rose, Moshe Eizenman |
ICPR | 3 |
| 2020 | Optimizing FPGA Logic Block Architectures for ArithmeticabstractHardened adder and carry logic is widely used in commercial field-programmable gate arrays (FPGAs) to improve the efficiency of arithmetic functions. There are many design choices and complexities associated with such hardening, including circuit design, FPGA architectural choices, and the computer-aided design (CAD) flow. However, these choices have not been studied much and hence we explore a number of possibilities. We also highlight front-end elaboration optimization that helps ameliorate the restrictions placed on logic synthesis by hardened arithmetic. We show that hard adders and carry chains increase the performance of simple adders by a factor of 4 or more, but on larger benchmark designs that contain arithmetic improve the overall performance by 15%. Our results also show that for complete application circuits simple hardened ripple-carry adders perform as well as more complex carry-lookahead adders. Our best non-fracturable lookup table (non-fLUT) architecture with hardened arithmetic yields 12% better area-delay product than architectures without hardened arithmetic. We also investigate the impact of fLUTs and their interaction with hardened arithmetic. We find that fLUTs offer significant (12%-15%) area reduction, which is complementary to the delay reduction of hardened arithmetic. Therefore, our best fLUT architectures which use two bits of hardened arithmetic achieve 25% better area-delay product than non-fLUT architectures without hardened arithmetic. Kevin E. Murray, Jason Luu, Matthew J. P. Walker, Conor McCullough, Safeen Huda, Charles Chiasson, Kenneth B. Kent, Jason Helge Anderson, Jonathan Rose, Vaughn Betz |
IEEE Trans. Very Large Scale Integr. Syst. | 11 |
| 2018 | Automatic Topology Optimization for FPGA Interconnect SynthesisabstractThe goal of FPGA interconnect synthesis is to generate a physical network that connects user-supplied functional modules according to a logical specification of the desired connectivity. In this paper, we augment an existing FPGA interconnect synthesis flow with the ability to automatically design the topology of the generated network while reducing its area subject to user-supplied performance specifications. The key specification is a per-transmission importance value representing the designer's willingness to have a transmission contend with other transmissions. The designer may also optionally specify that certain transmissions will never temporally overlap. We present an iterative algorithm that generates a topology which respects these specifications, with the goal of reducing area. Optimization decisions are guided by pre-characterized area models of interconnect primitives and an analytical worst-case traffic contention model. We apply our approach to a case study of an FPGA-based linear algebra application, where we successfully optimize the topologies of two of its sub-networks resulting in area savings of 60% and 75% with no overall performance degredation. Alex Rodionov, Jonathan Rose |
FPL | 2 |
| 2018 | High-Performance Instruction Scheduling Circuits for Superscalar Out-of-Order Soft ProcessorsabstractSoft processors have a role to play in simplifying field-programmable gate array (FPGA) application design as they can be deployed only when needed, and it is easier to write and debug single-threaded software code than create hardware. The breadth of this second role increases when the performance of the soft processor increases, yet the sophisticated out-of-order superscalar approaches that arrived in the mid-1990s are not employed, despite their area cost now being easily tolerable. In this article, we take an important step toward out-of-order execution in soft processors by exploring instruction scheduling in an FPGA substrate. This differs from the hard-processor design problem because the logic substrate is restricted to LUTs, whereas hard processor scheduling circuits employ CAM and wired-OR structures to great benefit. We discuss both circuit and microarchitectural trade-offs and compare three circuit structures for the scheduler, including a new structure called a fused-logic matrix scheduler . Using our optimized circuits, we show that four-issue distributed schedulers with up to 54 entries can be built with the same cycle time as the commercial Nios II/f soft processor (240MHz). This careful design has the potential to significantly increase both the IPC and raw compute performance of a soft processor, compared to current commercial soft processors. Henry Wong, Vaughn Betz, Jonathan Rose |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2017 | Synchronization Constraints for Interconnect Synthesis
Alex Rodionov, Jonathan Rose |
FPGA | 2 |
| 2016 | High Performance Instruction Scheduling Circuits for Out-of-Order Soft ProcessorsabstractSoft processors have a role to play in easing the difficulty of designing applications into FPGAs for two reasons: first, they can be deployed only when needed, unlike permanent on-die hard processors. Second, for the portions of an application that can function sufficiently fast on a soft processor, it is far easier to write and debug single-threaded software code than to create hardware. The breadth of this second role increases when the performance of the soft processor increases, yet there has been little progress in the performance of soft processors since their commercial inception -- in particular, the sophisticated out-of-order superscalar approaches that arrived in the mid 1990s are not employed, despite the fact that their area cost is now easily tolerable. In this paper we take an important step towards out-of-order execution in soft processors by exploring instruction scheduling in an FPGA substrate. This differs from the hard-processor design problem because the logic substrate is restricted to LUTs, whereas hard processor scheduling circuits employ CAM and wired-OR structures to great benefit. We discuss both circuit and microarchitectural trade-offs, and compare three circuit structures for the scheduler, including a new structure called a fused-logic matrix scheduler. With this circuit, large schedulers up to 40 entries can be built with the same cycle time as the commercial Nios II/f soft processor (240~MHz). This careful design has the potential to significantly increase both the IPC and raw compute performance of a soft processor, compared to current commercial soft processors. Henry Wong, Vaughn Betz, Jonathan Rose |
FCCM | 3 |
| 2016 | Fine-Grained Interconnect SynthesisabstractOne of the key challenges for the FPGA industry going forward is to make the task of designing hardware easier. A significant portion of that design task is the creation of the interconnect pathways between functional structures. We present a synthesis tool that automates this process and focuses on the interconnect needs in the fine-grained (sub-IP-block) design space. Here there are several issues that prior research and tools do not address well: the need to have fixed, deterministic latency between communicating units (to enable high-performance local communication without the area overheads of latency insensitivity), and the ability to avoid generating unnecessary arbitration hardware when the application design can avoid it. Using a design example, our tool generates interconnect that requires 69% fewer lines of specification code than a handwritten Verilog implementation, which is a 32% overall reduction for the entire application. The resulting system, while requiring 6% more total functional and interconnect area, achieves the same performance. We also show a quantitative and qualitative advantages against an existing commercial interconnect synthesis tool, over which we achieve a 25% performance advantage and 15%/57% logic/memory area savings. Alex Rodionov, David Biancolin, Jonathan Rose |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2016 | Microarchitecture and Circuits for a 200 MHz Out-of-Order Soft Processor Memory SystemabstractAlthough FPGAs have grown in capacity, FPGA-based soft processors have grown very little because of the difficulty of achieving higher performance in exchange for area. Superscalar out-of-order processors promise large performance gains, and the memory subsystem is a key part of such a processor that must help supply increased performance. In this article, we describe and explore microarchitectural and circuit-level tradeoffs in the design of such a memory system. We show the significant instructions-per-cycle wins for providing various levels of out-of-order memory access and memory dependence speculation (1.32 × SPECint2000) and for the addition of a second-level cache (another 1.60 × ). With careful microarchitecture and circuit design, we also achieve a L1 translation lookaside buffers and cache lookup with 29% less logic delay than the simpler Nios II/f memory system. Henry Wong, Vaughn Betz, Jonathan Rose |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2015 | Fine-Grained Interconnect SynthesisabstractOne of the key challenges for the FPGA industry going forward is to make the task of designing hardware easier. A significant portion of that design task is the creation of the interconnect pathways between functional structures. We present a synthesis tool that automates this process and focuses on the interconnect needs in the fine-grained (sub-IP-block) design space. Here there are several issues that prior research and tools do not address well: the need to have fixed, deterministic latency between communicating units (to enable high-performance local communication without the area overheads of latency-insensitivity), and the ability to avoid generating un-necessary arbitration hardware when the application design can avoid it. Using a design example, our tool generates interconnect that requires 72% fewer lines of specification code than a hand-written Verilog implementation, which is a 33% overall reduction for the entire application. The resulting system, while requiring 4% more total functional and interconnect area, achieves the same performance. We also show a quantitative and qualitative advantages against an existing commercial interconnect synthesis tool, over which we achieve a 25% performance advantage and 17%/57% logic/memory area savings. Alex Rodionov, David Biancolin, Jonathan Rose |
FPGA | 3 |
| 2015 | Automatic FPGA system and interconnect construction with multicast and customizable topologyabstractModern FPGA system integration tools, such as Qsys and Vivado, seek to help designers easily instantiate and connect IP cores. These tools require the cores to ascribe to a specific interconnect abstraction, which is either a high-level memory-mapped or low-level point-to-point streaming protocol. We present a system-level construction tool that gives the designer more control over the nature of the interconnect - specifically permitting multicast communication and the ability to easily construct custom network topologies. Our tool requires less user input than Qsys and yields a 5% area savings and 35% reduction in simulated execution time for a large and complex linear algebra application. The primary goal of this work is to make the designer's job easier by automating the design of interconnect, including exploration of alternative structures and communication patterns. Alex Rodionov, Jonathan Rose |
FPT | 2 |
| 2014 | On Hard Adders and Carry Chains in FPGAsabstractHardened adder and carry logic is widely used in commercial FPGAs to improve the efficiency of arithmetic functions. There are many design choices and complexities associated with such hardening, including circuit design, FPGA architectural choices, and the CAD flow. There has been very little study, however, on these choices and hence we explore a number of possibilities for hard adder design. We also highlight optimizations during front-end elaboration that help ameliorate the restrictions placed on logic synthesis by hardened arithmetic. We show that hard adders and carry chains, when used for simple adders, increase performance by a factor of four or more, but on larger benchmark designs that contain arithmetic, improve overall performance by roughly 15%. We measure an average area increase of 5% for architectures with carry chains but believe that better logic synthesis should reduce this penalty. Interestingly, we show that adding dedicated inter-logic-block carry links or fast carry look-ahead hardened adders result in only minor delay improvements for complete designs. Jason Luu, Conor McCullough, Safeen Huda, Charles Chiasson, Kenneth B. Kent, Jason Helge Anderson, Jonathan Rose, Vaughn Betz |
FCCM | 9 |
| 2014 | Towards interconnect-adaptive packing for FPGAsabstractIn order to investigate new FPGA logic blocks, FPGA architects have traditionally needed to customize CAD tools to make use of the new features and characteristics of those blocks. The software development effort necessary to create such CAD tools can be a time-consuming process that can significantly limit the number and variety of architectures explored. Thus, architects want flexible CAD tools that can, with few or no software modifications, explore a diverse space. Existing flexible CAD tools suffer from impractically long runtimes and/or fail to efficiently make use of the important new features of the logic blocks being investigated. This work is a step towards addressing these concerns by enhancing the packing stage of the open-source VTR CAD flow [17] to efficiently deal with common interconnect structures that are used to create many kinds of useful novel blocks. These structures include crossbars, carry chains, dedicated signals, and others. To accomplish this, we employ three techniques in this work: speculative packing, pre-packing, and interconnect-aware pin counting. We show that these techniques, along with three minor modifications, result in improvements to runtime and quality of results across a spectrum of architectures, while simultaneously expanding the scope of architectures that can be explored. Compared with VTR 1.0 [17], we show an average 12-fold speedup in packing for fracturable LUT architectures with 20% lower minimum channel width and 6% lower critical path delay. We obtain a 6 to 7-fold speedup for architectures with non-fracturable LUTs and architectures with depopulated crossbars. In addition, we demonstrate packing support for logic blocks with carry chains. Jason Luu, Jonathan Rose, Jason Helge Anderson |
FPGA | 2 |
| 2014 | VTR 7.0: Next Generation Architecture and CAD System for FPGAsabstractExploring architectures for large, modern FPGAs requires sophisticated software that can model and target hypothetical devices. Furthermore, research into new CAD algorithms often requires a complete and open source baseline CAD flow. This article describes recent advances in the open source Verilog-to-Routing (VTR) CAD flow that enable further research in these areas. VTR now supports designs with multiple clocks in both timing analysis and optimization. Hard adder/carry logic can be included in an architecture in various ways and significantly improves the performance of arithmetic circuits. The flow now models energy consumption, an increasingly important concern. The speed and quality of the packing algorithms have been significantly improved. VTR can now generate a netlist of the final post-routed circuit which enables detailed simulation of a design for a variety of purposes. We also release new FPGA architecture files and models that are much closer to modern commercial architectures, enabling more realistic experiments. Finally, we show that while this version of VTR supports new and complex features, it has a 1.5× compile time speed-up for simple architectures and a 6× speed-up for complex architectures compared to the previous release, with no degradation to timing or wire-length quality. Jason Luu, Jeffrey B. Goeders, Michael Wainberg, Andrew Somerville, Thien Yu, Konstantin Nasartschuk, Miad Nasr, Tim Liu, Nooruddin Ahmed, Kenneth B. Kent, Jason Helge Anderson, Jonathan Rose, Vaughn Betz |
ACM Trans. Reconfigurable Technol. Syst. | 13 |
| 2014 | Quantifying the Gap Between FPGA and Custom CMOS to Aid Microarchitectural DesignabstractThis paper compares the delay and area of a comprehensive set of processor building block circuits when implemented on custom CMOS and FPGA substrates, then uses these results to show how soft processor microarchitectures should be different from those of hard processors. We find that the ratios of the custom CMOS versus FPGA area for different building blocks varies considerably more than the speed ratios, thus, area ratios have more impact on microarchitecture choices. Complete processor cores on an FPGA use 17-27 × more area (“area ratio”) than the same design implemented in custom CMOS. Building blocks with dedicated hardware support on FPGAs such as SRAMs, adders, and multipliers are particularly area-efficient (2-7×), while multiplexers and content-addressable memories (CAM) are particularly area-inefficient (>100×). Applying these results, we find out-of-order soft processors should use physical register file organizations to minimize CAM size. Henry Wong, Vaughn Betz, Jonathan Rose |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | Efficient methods for out-of-order load/store execution for high-performance soft processorsabstractAs FPGAs continue to increase in size, it becomes increasingly feasible and desirable to build higher performance soft processors. Preserving the familiar single-threaded programming model can be done with an out of order processor. The ability to execute memory loads and stores out of order has a large impact on performance, but this is difficult to do because the dependencies between stores and loads are not known until addresses are computed. Out of order memory disambiguation is traditionally done with CAMs in the load queue and store queue, but large CAMs are inefficient on FPGAs. Store Queue Index Prediction (SQIP) and NoSQ propose to replace CAMs with store-load forwarding prediction and load re-execution. We implement four memory disambiguation schemes (in-order, CAM, SQIP, NoSQ) on a Stratix IV FPGA and evaluate the area and delay trade-offs. We find that CAM area and delay degrade quickly with load/store queue size, while SQIP and NoSQ have little degradation with queue size but have area overhead for prediction and predictor training hardware. SQIP and NoSQ use less area than CAMs beyond 32 and 16 load/store queue entries, respectively, and have higher maximum frequency beyond 4 entries. Henry Wong, Vaughn Betz, Jonathan Rose |
FPT | 3 |
| 2012 | The VTR project: architecture and CAD for FPGAs from verilog to routingabstractTo facilitate the development of future FPGA architectures and CAD tools -- both embedded programmable fabrics and pure-play FPGAs -- there is a need for a large scale, publicly available software suite that can synthesize circuits into easily-described hypothetical FPGA architectures. These circuits should be captured at the HDL level, or higher, and pass through logical and physical synthesis. Such a tool must provide detailed modelling of area, performance and energy to enable architecture exploration. As software flows themselves evolve to permit design capture at ever higher levels of abstraction, this downstream full-implementation flow will always be required. This paper describes the current status and new release of an ongoing effort to create such a flow - the 'Verilog to Routing' (VTR) project, which is a broad collaboration of researchers. There are three core tools: ODIN II for Verilog Elaboration and front-end hard-block synthesis, ABC for logic synthesis, and VPR for physical synthesis and analysis. ODIN II now has a simulation capability to help verify that its output is correct, as well as specialized synthesis at the elaboration step for multipliers and memories. ABC is used to optimize the 'soft' logic of the FPGA. The VPR-based packing, placement and routing is now fully timing-driven (the previous release was not) and includes new capability to target complex logic blocks. In addition we have added a set of four large benchmark circuits to a suite of previously-released Verilog HDL circuits. Finally, we illustrate the use of the new flow by using it to help architect a floating-point unit in an FPGA, and contrast it with a prior, much longer effort that was required to do the same thing. Jonathan Rose, Jason Luu, Chi Wai Yu, Opal Densmore, Jeffrey B. Goeders, Andrew Somerville, Kenneth B. Kent, Peter Jamieson, Jason Helge Anderson |
FPGA | 1 |
| 2012 | On the difficulty of pin-to-wire routing in FPGAsabstractWhile FPGA programmable routing networks are designed to connect logic block output pins to input pins, FPGA users and architects sometimes become motivated to create connections between pins and specific wires in an FPGA. We call these pin-to-wire connections, and they are motivated by several reasons: first, a desire to employ routing-by-abutment, as commonly done in custom VLSI, to build modular, pre-laid out systems. Second, partial reconfiguration of FPGAs often requires that circuits in the FPGA connect by abutment. Third, pin-to-wire routing is required to make use of resources that reside within the routing network itself, such as the plentiful multiplexers in the network, or even the configuration bits themselves. In this paper we attempt to understand and measure how difficult it is to form such pin-to-wire connections. We show, for example, under an experimental scenario close to routing-by-abutment, that the total routed wirelength compared to a flat placement of the complete system increases by about 6%, that the critical path delay increases by 15% and the router effort goes up by a factor of 3.5. To achieve this result, it is important to be careful in selecting the specific target wires. Overall we demonstrate that while pin-to-wire connections definitely impose increased stress on the routing architecture and router, it is possible to route some reasonable number of them, and so they can be used under some circumstances. Niyati Shah, Jonathan Rose |
FPL | 2 |
| 2012 | An energy-efficient, fast FPGA hardware architecture for OpenCV-Compatible object detectionabstractThe presence of cameras and powerful computers on modern mobile devices gives rise to the hope that they can perform computer vision tasks as we walk around. However, the computational demand and energy consumption of computer vision tasks such as object detection, recognition and tracking make this challenging. At the same time, a fixed vision hard core on the SoC contained in a mobile chip may not have the flexibility needed to adapt to new situations, or evolve as new algorithms are discovered. This may mean that computer vision on a mobile device is the killer application for FPGAs, and could motivate the inclusion of FPGAs, in some form, within modern smartphones. In this paper we present a novel hardware architecture for object detection, that is bit-for-bit compatible with the object classifiers in the widely-used open source OpenCV computer vision software. The architecture is novel, compared to prior work in this area, in two ways: its memory architecture, and its particular SIMD-type of processing. The implementation, which consists of the full system, not simply the kernel, outperforms a same-generation technology mobile processor by a factor of 59 times, and is 13.5 times more energy-efficient. Braiden Brousseau, Jonathan Rose |
FPT | 2 |
| 2012 | Portable and scalable FPGA-based acceleration of a direct linear system solverabstractFPGAs have the potential to serve as a platform for accelerating many computations including scientific applications. However, the large development cost and short life span for FPGA designs have limited their adoption by the scientific computing community. FPGA-based scientific computing and many kinds of embedded computing could become more practical if there were hardware libraries that were portable to any FPGA-based system with performance that scaled with the size of the FPGA. To illustrate this idea we have implemented one common super-computing library function: the LU factorization method for solving systems of linear equations. This paper describes a method for making the design both portable and scalable that should be illustrative if such libraries are to be built in the future. The design is a software-based generator that leverages both the flexibility of a software programming language and the parameters inherent in an hardware description language. The generator accepts parameters that describe the FPGA capacity and external memory capabilities. We compare the performance of our engine executing on the largest FPGA available at the time of this work (an Altera Stratix III 3S340) to a single processor core fabricated in the same 65nm IC process running a highly optimized software implementation from the processor vendor. For single precision matrices on the order of 10,000 × 10,000 elements, the FPGA implementation is 2.2 times faster and the energy dissipated per useful GFLOP operation is a factor of 5 times less. For double precision, the FPGA implementation is 1.7 times faster and 3.5 times more energy efficient. Wei Zhang 0222, Vaughn Betz, Jonathan Rose |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2012 | Portable, Flexible, and Scalable Soft Vector ProcessorsabstractField-programmable gate arrays (FPGAs) are increasingly used to implement embedded digital systems, however, the hardware design necessary to do so is time-consuming and tedious. The amount of hardware design can be reduced by employing a microprocessor for less-critical computation in the system. Often this microprocessor is implemented using the FPGA reprogrammable fabric as a soft processor which presently have simple architectures and moderate performance. Our goal is to scale the performance of existing soft processors hence expanding their suitability to more critical computation. To this end we propose extending soft processors with vector extensions to exploit the abundant data parallelism found in many embedded kernels. Such a soft vector processor can execute these kernels much faster than a single-core hence reducing the need for hardware implementations. We observe this improved execution speed through experimentation with vector extended soft processor architecture (VESPA) which is designed, implemented, and evaluated on real FPGA hardware. VESPA is shown to effectively scale performance up to 32 lanes, while providing substantial architectural flexibility to create a fine-grained design space. With these characteristics, and portability across FPGA devices, soft vector processors can provide exact-fit architectures which can efficiently and more easily implement data parallel workloads over custom FPGA hardware design. Peter Yiannacouras, J. Gregory Steffan, Jonathan Rose |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | Architecture description and packing for logic blocks with hierarchy, modes and complex interconnectabstractThe development of future FPGA fabrics with more sophisticated and complex logic blocks requires a new CAD flow that permits the expression of that complexity and the ability to synthesize to it. In this paper, we present a new logic block description language that can depict complex intra-block interconnect, hierarchy and modes of operation. These features are necessary to support modern and future FPGA complex soft logic blocks, memory and hard blocks. The key part of the CAD flow associated with this complexity is the packer, which takes the logical atomic pieces of the complex blocks and groups them into whole physical entities. We present an area-driven generic packing tool that can pack the logical atoms into any heterogeneous FPGA described in the new language, including many different kinds of soft and hard logic blocks. We gauge its area quality by comparing the results achieved with a lower bound on the number of blocks required, and then illustrate its explorative capability in two ways: on fracturable LUT soft logic architectures, and on hard block memory architectures. The new infrastructure attaches to a flow that begins with a Verilog front-end, permitting the use of benchmarks that are significantly larger than the usual ones, and can target heterogenous FPGAs. Jason Luu, Jason Helge Anderson, Jonathan Rose |
FPGA | 3 |
| 2011 | The role of FPGAs in a converged future with heterogeneous programmable processors: pre-conference workshopabstractThe battle of fixed function devices vs. programmable devices has been won by the programmables. The question facing us now is to determine what kinds of programmability to place on next generation systems/devices. Research and development on many applications has shown that different kinds of hardware and software programmability succeed for different application classes: powerful, singlethread-optimized CPUs continue to do very well for many applications; the General Purpose GPU is carving a niche in high throughput, parallel floating point codes in addition to its home turf of graphics; the FPGA is particularly good at variable bit-size computations and data steering, as well as parallel distributed control of networks. Future systems may well need all three types of these types of engines, and perhaps interesting mixtures of them. This is particularly true when we deal with the combined goals of optimizing cost, performance and energy. Jonathan Rose, Guy Lemieux |
FPGA | 1 |
| 2011 | Comparing FPGA vs. custom cmos and the impact on processor microarchitectureabstractAs soft processors are increasingly used in diverse applications, there is a need to evolve their microarchitectures in a way that suits the FPGA implementation substrate. This paper compares the delay and area of a comprehensive set of processor building block circuits when implemented on custom CMOS and FPGA substrates. We then use the results of these comparisons to infer how the microarchitecture of soft processors on FPGAs should be different from hard processors on custom CMOS. Henry Wong, Vaughn Betz, Jonathan Rose |
FPGA | 3 |
| 2011 | VPR 5.0: FPGA CAD and architecture exploration tools with single-driver routing, heterogeneity and process scalingabstractThe VPR toolset has been widely used in FPGA architecture and CAD research, but has not evolved over the past decade. This article describes and illustrates the use of a new version of the toolset that includes four new features: first, it supports a broad range of single-driver routing architectures, which have superior architectural and electrical properties over the prior multidriver approach (and which is now employed in the majority of FPGAs sold). Second, it can now model, for placement and routing a heterogeneous selection of hard logic blocks. This is a key (but not final) step toward the incluion of blocks such as memory and multipliers. Third, we provide optimized electrical models for a wide range of architectures in different process technologies, including a range of area-delay trade-offs for each single architecture. Finally, to maintain robustness and support future development the release includes a set of regression tests for the software. To illustrate the use of the new features, we explore several architectural issues: the FPGA area efficiency versus logic block granularity, the effect of single-driver routing, and a simple use of the heterogeneity to explore the impact of hard multipliers on wiring track count. Jason Luu, Ian Kuon, Peter Jamieson, Ted Campbell, Andy Gean Ye, Wei Mark Fang, Kenneth B. Kent, Jonathan Rose |
ACM Trans. Reconfigurable Technol. Syst. | 8 |
| 2011 | Exploring Area and Delay Tradeoffs in FPGAs With Architecture and Automated Transistor DesignabstractField-programmable gate arrays (FPGAs) are used in a variety of markets that have differing cost, performance and power consumption requirements. While it would be ideal to serve all these markets with a single FPGA family, the diversity in the needs of these markets means that generally more than one family is appropriate. Consequently, FPGA vendors have moved to provide a diverse set of families that sit at different points in the area-speed-power design space. This paper aims to understand the circuit and architectural design attributes of FPGAs that enable tradeoffs between area and speed, and to determine the magnitude of the possible tradeoffs. This will be useful for architects seeking to determine the number of device families in a suite of offerings, as well as the changes to make between families. We explore a broad range of architectures and circuit designs and developed a transistor sizing tool that automatically optimizes each design. In this paper, we describe this tool and demonstrate that it achieves results that are comparable to past work but with vastly less effort. We then use the designs produced by the tool to explore the range of tradeoffs possible. We find that through architecture and transistor sizing changes it is possible to usefully vary the area of an FPGA by a factor of 2.0 and the performance of an FPGA by a factor of 2.1. We also observe that the range of area and delay tradeoffs possible by varying only the transistor sizing of a single architecture is larger than the ranges observed in past architectural experiments. In addition to transistor size, we note that LUT size is one of the most useful parameters for trading off area and delay. Ian Kuon, Jonathan Rose |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | Enhancing the Area Efficiency of FPGAs With Hard Circuits Using Shadow ClustersabstractThere is a dramatic logic density gap between field-programmable gate arrays (FPGAs) and application-specific integrated circuits, and this gap is the main reason FPGAs are not cost-effective in high-volume applications. Modern FPGAs narrow this gap by including “hard” circuits such as memories and multipliers, which are very efficient when they are used. However, if these hard circuits are not used, they go wasted (including the very expensive programmable routing that surrounds the logic), and have a negative impact on logic density. In this paper, we present an architectural concept, called shadow clusters, which seeks to mitigate this loss. A shadow cluster is a standard FPGA logic “cluster” (typically consisting of a group of lookup tables and flip-flops) that is placed “behind” every hard circuit, and can programmably, through simple, small multiplexers, replace the hard circuit in the event it is not needed. A shadow cluster is effective because the largest area cost, by far, in an FPGA is for the programmable routing that connects the logic. The shadow cluster area cost is small, and yet it enables more consistent employment of the programmable routing across applications with varying demand for hard circuits. We introduce new terminology to describe the economics of hard circuits on FGPAs, and provide a scientific way to measure the area effectiveness. We measure the area efficiency of FPGAs with and without shadow clusters, and show that a modern commercial architecture (with a fixed ratio of multipliers to soft logic) would gain 4.7% in area efficiency by employing shadow clusters. Indeed, every architecture we studied under “reasonable” conditions never showed a loss of area efficiency. Furthermore, we show that most area-efficient architecture that employs the shadow cluster concept is 12.5% better than the most area-efficient architecture without shadow clusters. Peter Jamieson, Jonathan Rose |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Fine-grain performance scaling of soft vector processorsabstractEmbedded systems are often implemented on FPGA devices and 25% of the time include a soft processor--a processor built using the FPGA reprogrammable fabric. Because of their prevalence and flexibility, soft processors are compelling targets for customization--although current soft processors provide few architectural variations. Recent work has proposed augmenting soft processors with customizable vector processing support, enabling designers to easily scale performance by exploiting the data parallelism available in an application. However this approach provides only coarse-grain scaling, by successively doubling the number of vector datapaths for less than double the performance. Peter Yiannacouras, J. Gregory Steffan, Jonathan Rose |
CASES | 3 |
| 2009 | FPGA-based Monte Carlo Computation of Light Absorption for Photodynamic Cancer TherapyabstractPhotodynamic therapy (PDT) is a method of treating cancer that combines light and light-sensitive drugs to selectively destroy cancerous tumours without harming the healthy tissue. The success of PDT depends on the accurate computation of light dose distribution. Monte Carlo (MC) simulations can provide an accurate solution for light dose distribution, but have high computation time that prevents them from being used in treatment planning. To alleviate this problem, a hardware design of an MC simulation based on the gold standard software in biophotonics was implemented on a large modern FPGA. This implementation achieved a 28-fold speedup and 716-fold lower power-delay product compared to the gold standard software executed on a 3 GHz Intel Xeon 5160 processor. The accuracy of the hardware was compared to the gold standard using a realistic skin model. An experiment using 100 million photon packets yielded a light dose distribution that diverged by less than 0.1 mm. We also describe our development methodology, which employs an intermediate hardware description in SystemC prior to Verilog coding that led to significant design effort efficiency. Jason Luu, Keith Redmond, William Lo, Paul Chow, Lothar Lilge, Jonathan Rose |
FCCM | 6 |
| 2009 | VPR 5.0: FPGA cad and architecture exploration tools with single-driver routing, heterogeneity and process scalingabstractThe VPR toolset [6, 7] has been widely used to perform FPGA architecture and CAD research, but has not evolved over the past decade to include many architectural features now present in modern FPGAs. This paper describes a new version of the toolset that includes four significant features: first, it now supports a broad range of single-driver routing architectures [29, 4, 16]. Single-driver routing has significantly different architectural and electrical properties from the multi-driver approach previously modelled, and is now employed in the majority of FPGAs sold. Second, the new release can now model a heterogeneous selection of hard logic blocks, which could include the hard memory and multipliers that are now ubiquitous in FPGAs. Third, we provide optimized electrical models of a wide range of architectures in different process technologies, including a range of area-delay tradeoffs for each single architecture. Prior releases of VPR did not publish even one architecture file with accurate resistance and capacitance parameters. Finally, to maintain robustness and to support future development the release includes a set of regression tests to check functionality and quality of result of the output of the tools. Jason Luu, Ian Kuon, Peter Jamieson, Ted Campbell, Andy Gean Ye, Wei Mark Fang, Jonathan Rose |
FPGA | 7 |
| 2009 | Soft vector processors vs FPGA custom hardware: measuring and reducing the gapabstractSoft processors are often used in FPGA-based systems because of their ease-of-use, but for a given computation there is a significant gap in area/performance between a C code implementation executing on a soft processor and a custom FPGA hardware implementation. Recent research has demonstrated that soft processors augmented with support for vector instructions provide significant improvements in performance and scalability for data-parallel workloads. In this work, using an FPGA platform equipped with DDR memory executing data-parallel benchmarks from the industry-standard EEMBC suite, we measure the area/performance gaps between (i) C programs executing on a scalar soft processor, (ii) hand-vectorized programs executing on a soft vector processor, and (iii) custom FPGA hardware. We demonstrate that the wall clock performance gap between scalar executed C and custom hardware can be drastically reduced using our improved soft vector processors, even though they are still clocked 3x slower than custom hardware. We identify loop overhead, data delivery, and exact resource usage as three key advantages of custom hardware that we propose to mitigate in our soft vector processor respectively by decoupling pipelines, tuning cache design, supporting prefetching, and automatically eliminating unused instructions and datapath width. We show that together these improvements increase performance by 3x and reduce the area of the fastest soft vector processor by 2x, significantly reducing the need for designers to resort to more challenging custom hardware implementations. Peter Yiannacouras, J. Gregory Steffan, Jonathan Rose |
FPGA | 3 |
| 2009 | The evolution of architecture exploration of programmable devicesabstractAs integrated circuit fabrication processes continue to provide exponential increases in density of transistors with each generation, the question of what to do with those transistors becomes ever more interesting. The most fundamental part of that question is the global organization of the structures created from the transistors, most commonly referred to as the *architecture* of the device. Most IC architecture exploration that is done is quite empirical, with example uses driving through tools to experimentally test new ideas for structures and organizations. This method is used in both programmable logic hardware such as FPGAs, and in programmable instruction set processors. As the processor world now seeks to gain performance through parallelism, its architecture questions have begun to look more similar to those in the FPGA domain. In this talk the author discuss the evolution of the architecture exploration processes that we have worked on at the University of Toronto, and of the new levels that we are currently trying to build. The current effort focusses on HDL-level circuits as "example uses" and this turns out to be rather intricate in the face of the kinds of architecture questions that could be posed. In the future, it may well be that some form of software is the input "example use," thus bringing the processor and FPGA world closer together. For this to work, there needs to be an effective CAD/compiler flow from software to the HDL level. The author give perspective on the state of this art, and discuss what kind of commonality might evolve in architecture exploration tools for FPGAs and processors. Jonathan Rose |
FPL | 1 |
| 2009 | Data parallel FPGA workloads: Software versus hardwareabstractCommercial soft processors are unable to effectively exploit the data parallelism present in many embedded systems workloads, requiring FPGA designers to exploit it (laboriously) with manual hardware design. Recent research has demonstrated that soft processors augmented with support for vector instructions provide significant improvements in performance and scalability for data parallel workloads. These soft vector processors provide a software environment for quickly encoding data parallel computation, but their competitiveness with manual hardware design in terms of area and performance remains unknown. In this work, using an FPGA platform equipped with DDR memory executing data-parallel EEMBC embedded benchmarks, we measure the area/performance gaps between (i) a scalar soft processor, (ii) our improved soft vector processor, and (iii) custom FPGA hardware. We demonstrate that the 432times wall clock performance gap between scalar executed C and custom hardware can be reduced significantly to 17times using our improved soft vector processor, while silicon-efficiency is improved by 3times in terms of area delay product. We modified the architecture to mitigate three key advantages we observed in custom hardware: loop overhead, data delivery, and exact resource usage. Combined these improvements increase performance by 3times and reduce area by almost half, significantly reducing the need for designers to resort to more challenging custom hardware implementations. Peter Yiannacouras, J. Gregory Steffan, Jonathan Rose |
FPL | 3 |
| 2008 | VESPA: portable, scalable, and flexible FPGA-based vector processorsabstractWhile soft processors are increasingly common in FPGA-based embedded systems, it remains a challenge to scale their performance. We propose extending soft processor instruction sets to include support for vector processing. The resulting system of vectorized software and soft vector processor hardware is (i) portable to any FPGA architecture and vector processor configuration, (ii) scalable to larger yet higher-performance designs, and (iii) flexible, allowing the underlying vector processor to be customized to match the needs of each application. Using our robust and verified parameterized vector processor design and industry-standard EEMBC benchmarks, we evaluate the performance and area trade-offs for different soft vector processor configurations using an FPGA development platform with DDR SDRAM. We find that on average we can scale performance from 1.8x up to 6.3x for a vector processor design that saturates the capacity of our platform's Stratix 1S80 FPGA. We also automatically generate application-specific vector processors with reduced datapath width and instruction set support which combined reduce the area by up to 70% (61% on average) without affecting performance. Peter Yiannacouras, J. Gregory Steffan, Jonathan Rose |
CASES | 3 |
| 2008 | Automated transistor sizing for FPGA architecture explorationabstractThe creation of an FPGA requires extensive transistor-level design. This is necessary for both the final design, and during architecture exploration, when many different logic and routing architectures are considered. For such explorations, it is not feasible to spend significant amounts of time on transistor-level design. This paper presents an automated transistor sizing tool for FPGA architecture exploration that uses a two-phased approach - a coarse rapid phase with simple modeling followed by refinement with much more accurate models. The output of the system is a design optimized towards a specific area-delay criterion. We compare the quality of our results to prior manual and partially automated approaches. Also, our tool has been used to produce hundreds of candidate architectures which we are releasing to support future high quality explorations. Ian Kuon, Jonathan Rose |
DAC | 2 |
| 2008 | Modeling routing demand for early-stage FPGA architecture developmentabstractArchitecture development for FPGAs has typically been a very empirical discipline, requiring the synthesis of benchmark circuits into candidate architectures. This is difficult to do in the early stages of architecture development, however, because there is no complete architecture to synthesize circuits into. The effort required to create prototype tools for nascent architectures is far too great for every new logic block or routing architecture idea, and so it would be extremely helpful to have a simple and intuitive FPGA interconnect model to guide the architect Wei Mark Fang, Jonathan Rose |
FPGA | 2 |
| 2008 | Area and delay trade-offs in the circuit and architecture design of FPGAsabstractField-programmable gate arrays (FPGAs) are used in a wide range of markets that have differing cost, performance and power consumption requirements. It would be advantageous if a single device family could serve these varied needs but the economics of catering to this wide distribution of market demands suggest more than one family is appropriate. Consequently, FPGA vendors have moved to provide a more diverse set of families that sit at different points in the area-speed-power design space. Ian Kuon, Jonathan Rose |
FPGA | 2 |
| 2008 | Portable and scalable FPGA-based acceleration of a direct linear system solverabstractFPGAs are becoming an attractive platform for accelerating many computations including scientific applications. However, their adoption has been limited by the large development cost and short life span of FPGA designs. We believe that FPGA-based scientific computation would become far more practical if there were hardware libraries that were portable to any FPGA with performance that could scale with the resources of the FPGA. To illustrate this idea we have implemented one common supercomputing library function: the LU factorization method for solving linear systems. This paper discusses issues in making the design both portable and scalable. The design is automatically generated to match the FPGA’s capabilities and external memory through the use of parameters. We compared the performance of the design on the FPGA to a single processor core and found that it performs 2.2 times faster, and that the energy dissipated per computation is a factor 5 times less. Wei Zhang 0222, Vaughn Betz, Jonathan Rose |
FPT | 3 |
| 2007 | Architecting Hard Crossbars on FPGAs and Increasing their Area Efficiency with Shadow ClustersabstractWe explore the architecture of on-chip hard crossbars in FPGAs and show that the area efficiency of such FPGAs can be improved when combined with shadow clusters (which are soft-logic LUT-based clusters that are architected to sit "behind" the multiplier), as an exemplar of an application circuit that appears less commonly in the designs targeting FPGAs. The metric that we seek to improve is the "frequency" that the need for hard crossbars must appear in the FPGA's target application suite for the inclusion of the hard crossbar to appear to be area-neutral. For example, we show that this break-even point for a hard 32 full-way crossbar changes from 32% of benchmarks needing to require crossbars to 9% for FPGAs with shadow clusters. Peter Jamieson, Jonathan Rose |
FPT | 2 |
| 2007 | Measuring the Gap Between FPGAs and ASICsabstractThis paper presents experimental measurements of the differences between a 90-nm CMOS field programmable gate array (FPGA) and 90-nm CMOS standard-cell application-specific integrated circuits (ASICs) in terms of logic density, circuit speed, and power consumption for core logic. We are motivated to make these measurements to enable system designers to make better informed choices between these two media and to give insight to FPGA makers on the deficiencies to attack and, thereby, improve FPGAs. We describe the methodology by which the measurements were obtained and show that, for circuits containing only look-up table-based logic and flip-flops, the ratio of silicon area required to implement them in FPGAs and ASICs is on average 35. Modern FPGAs also contain "hard" blocks such as multiplier/accumulators and block memories. We find that these blocks reduce this average area gap significantly to as little as 18 for our benchmarks, and we estimate that extensive use of these hard blocks could potentially lower the gap to below five. The ratio of critical-path delay, from FPGA to ASIC, is roughly three to four with less influence from block memory and hard multipliers. The dynamic power consumption ratio is approximately 14 times and, with hard blocks, this gap generally becomes smaller Ian Kuon, Jonathan Rose |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | Exploration and Customization of FPGA-Based Soft ProcessorsabstractAs embedded systems designers increasingly use field-programmable gate arrays (FPGAs) while pursuing single-chip designs, they are motivated to have their designs also include soft processors, processors built using FPGA programmable logic. In this paper, we provide: 1) an exploration of the microarchitectural tradeoffs for soft processors and 2) a set of customization techniques that capitalizes on these tradeoffs to improve the efficiency of soft processors for specific applications. Using our infrastructure for automatically generating soft-processor implementations (which span a large area/speed design space while remaining competitive with Altera's Nios II variations), we quantify tradeoffs within soft-processor microarchitecture and explore the impact of tuning the microarchitecture to the application. In addition, we apply a technique of subsetting the instruction set to use only the portion utilized by the application. Through these two techniques, we can improve the performance-per-area of a soft processor for a specific application by an average of 25% Peter Yiannacouras, J. Gregory Steffan, Jonathan Rose |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2006 | Measuring the gap between FPGAs and ASICsabstractThis paper presents experimental measurements of the differences between a 90nm CMOS FPGA and 90nm CMOS Standard Cell ASICs in terms of logic density, circuit speed and power consumption. We are motivated to make these measurements to enable system designers to make better informed hoices between these two media and to give insight to FPGA makers on the deficiencies to attack and thereby improve FPGAs. In the paper, we describe the methodology by which the measurements were obtained and we show that, for circuits containing only combinational logic and flip-flops, the ratio of silicon area required to implement them in FPGAs and ASICs is on average 40. Modern FPGAs also contain "hard" blocks such as multiplier/accumulators and block memories and we find that these blocks reduce this average area gap significantly to as little as 21. The ratio of critical path delay, from FPGA to ASIC, is roughly 3 to 4, with less influence from block memory and hard multipliers. The dynamic power onsumption ratio is approximately 12 times and, with hard blocks, this gap generally becomes smaller. Ian Kuon, Jonathan Rose |
FPGA | 2 |
| 2006 | Application-specific customization of soft processor microarchitectureabstractA key advantage of soft processors (processors built on an FPGA programmable fabric) over hard processors is that they can be customized to suit an application program's specific software. This notion has been exploited in the past principally through the use of application-specific instructions. While commercial soft processors are now widely deployed, they are available in only a few microarchitectural variations. In this work we explore the advantage of tuning the processor's microarchitecture to specific software applications, and show that there are significant advantages in doing so.Using an infrastructure for automatically generating soft processors that span the area/speed design space (while remaining competitive with Altera's Nios II variations), we explore the impact of tuning several aspects of microarchitecture including: (i) hardware vs software multiplication support; (ii) shifter implementation; and (iii) pipeline depth, organization, and forwarding. We find that the processor design that is fastest overall (on average across our embedded benchmark applications) is often also the fastest design for an individual application. However, in terms of area efficiency (i.e., performance-per-area), we demonstrate that a tuned microarchitecture can offer up to 30% improvement for three of the benchmarks and on average 11.4% improvement over the fastest-on-average design. We also show that our benchmark applications use only 50% of the available instructions on average, and that a processor customized to support only that subset of the ISA for a specific application can on average offer 25% savings in both area and energy. Finally, when both techniques for customization are combined we obtain an average improvement in performance-per-area of 25%. Peter Yiannacouras, J. Gregory Steffan, Jonathan Rose |
FPGA | 3 |
| 2006 | Enhancing the area-efficiency of FPGAs with hard circuits using shadow clustersabstractThere is a dramatic logic density gap between FPGAs and ASICs, and this gap is the main reason FPGAs are not cost-effective in high volume applications. Modern FPGAs narrow this gap by including "hard" circuits such as memories and multipliers, which are very efficient when they are used. However, if these hard circuits are not used, they go wasted (including the very expensive programmable routing that surrounds the logic) and have a negative impact on logic density. In this paper we propose a new architectural concept, called shadow clusters, that seeks to mitigate this loss. A shadow cluster is a standard FPGA logic "cluster" that is placed "behind" every hard circuit and can programmably, through simple, small multiplexers, replace the hard circuit in the event it isn't needed. The authors measure the area-efficiency of FPGAs with and without shadow clusters and show that a modern commercial architecture (with a fixed ratio of multipliers to soft logic) would gain 4.7% in area-efficiency by employing shadow clusters. Indeed, every architecture we studied under "reasonable" conditions never showed a loss of area-efficiency. Furthermore, we show that most area-efficient architecture that employs the shadow cluster concept is 12.5 % better than the most area-efficient architecture without shadow clusters Peter Jamieson, Jonathan Rose |
FPT | 2 |
| 2006 | Invited Keynote 1: Closing the gap between FPGAs and ASICs
Jonathan Rose |
FPT | 1 |
| 2006 | Reconfigurable hardware implementation of a phase-correlation stereoalgorithm
Ahmad Darabiha, W. James MacLean, Jonathan Rose |
Mach. Vis. Appl. | 3 |
| 2006 | Using Bus-Based Connections to Improve Field-Programmable Gate-Array Density for Implementing Datapath CircuitsabstractAs the logic capacity of field-programmable gate arrays (FPGAs) increases, they are increasingly being used to implement large arithmetic-intensive applications, which often contain a large proportion of datapath circuits. Since datapath circuits usually consist of regularly structured components (called bit-slices) which are connected together by regularly structured signals (called buses), it is possible to utilize datapath regularity in order to achieve significant area savings through FPGA architectural innovations. This paper describes such an FPGA routing architecture, called the multibit routing architecture, which employs bus-based connections in order to exploit datapath regularity. It is experimentally shown that, compared to conventional FPGA routing architectures, the multibit routing architecture can achieve 14% routing area reduction for implementing datapath circuits, which represents an overall FPGA area savings of 10%. This paper also empirically determines the best values of several important architectural parameters for the new routing architecture including the most area efficient granularity values and the most area efficient proportion of bus-based connections. Andy Gean Ye, Jonathan Rose |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2005 | The microarchitecture of FPGA-based soft processorsabstractAs more embedded systems are built using FPGA platforms, there is an increasing need to support processors in FPGAs. One option is the soft processor, a programmable instruction processor implemented in the reconfigurable logic of the FPGA. Commercial soft processors have been widely deployed, and hence we are motivated to understand their microarchitecture. We must re-evaluate microarchiteture in the soft processor context because an FPGA platform is significantly different than an ASIC platform---for example, the relative speed of memory and logic is quite different in the two platforms, as is the area cost. In this paper we present an infrastructure for rapidly generating RTL models of soft processors, as well as a methodology for measuring their area, performance, and power. Using our automatically-generated soft processors we explore the microarchitecture trade-off space including: (i) hardware vs software multiplication support; (ii) shifter implementations; and (iii) pipeline depth, organization, and forwarding. For example, we find that a 3-stage pipeline has better wall-clock-time performance than deeper pipelines, despite lower clock frequency. We also compare our designs to Altera's NiosII commercial soft processor variations and find that our automatically generated designs span the design space while remaining very competitive. Peter Yiannacouras, Jonathan Rose, J. Gregory Steffan |
CASES | 2 |
| 2005 | Design, layout and verification of an FPGA using automated toolsabstractCreating a new FPGA is a challenging undertaking because of the significant effort that must be spent on circuit design, layout and verification. It currently takes approximately 50 to 200 person years from architecture definition to tape-out for a new FPGA family. Such a lengthy development time is necessary because the process is primarily done manually. Simplifying and shortening the design process would be advantageous since it could reduce the time to market for new FPGAs while also enhancing architecture explorations. One way to accomplish this is through automation and, in this paper, we describe our efforts to automate the entire process by making use of a previously developed set of tools that assist in the creation of the repeatable FPGA tile [25]. Our aim is to demonstrate the feasibility of a CAD flow that uses an input FPGA architecture description to generate a layout that can be sent for fabrication. We prove the feasibility of this proposition by actually designing and fabricating a complete FPGA. Initial functional testing of the FPGA appears promising but is inconclusive at this time. Through this architecture to layout process, we investigate the issues that are faced in the architecture selection, circuit design, layout and verification of such an automatically produced FPGA. We found that there are significant savings in design time. As well, we demonstrate that we can produce a layout using automated tools that is only 36% larger than a commercial FPGA device layout. Given the significant time savings and the relatively minor area penalty, we feel that this work demonstrates that automated layout of FPGAs is practical and advantageous. Ian Kuon, Aaron Egier, Jonathan Rose |
FPGA | 3 |
| 2005 | The Stratix II logic and routing architectureabstractThis paper describes the Altera Stratix II™ logic and routing architecture. This architecture features a novel adaptive logic module (ALM) that is based on a 6-LUT, but can be partitioned into two smaller LUTs to efficiently implement circuits containing a range of LUT sizes that arises in conventional synthesis flows. This provides a performance increase of 15% in the Stratix II architecture while reducing area by 2%. The ALM also includes a more powerful arithmetic structure that can perform two bits of arithmetic per ALM, and perform a sum of up to three inputs. The routing fabric adds a new set of fast inputs to the routing multiplexers for another 3% improvement in performance, while other improvements in routing efficiency cause another 6% reduction in area. These changes in combination with other circuit and architecture changes in Stratix II contribute 27% of an overall 51% performance improvement (including architecture and process improvement). The architecture changes reduce area by 10% in the same process, and by 50% after including process migration. David M. Lewis, Elias Ahmed, Gregg Baeckler, Vaughn Betz, Mark Bourgeault, David Cashman, David R. Galloway, Mike Hutton, Christopher Lane, Andy Lee, Paul Leventis, Sandy Marquardt, Cameron McClintock, Ketan Padalia, Bruce Pedersen, Giles Powell, Boris Ratchev, Srinivas Reddy, Jay Schleicher, Kevin Stevens, Richard Yuan, Richard Cliff, Jonathan Rose |
FPGA | 23 |
| 2005 | Using bus-based connections to improve field-programmable gate array density for implementing datapath circuitsabstractAbstract—As the logic capacity of field-programmable gate arrays (FPGAs) increases, they are increasingly being used to implement large arithmetic-intensive applications, which often contain a large proportion of datapath circuits. Since datapath circuits usually consist of regularly structured components (called bitslices) which are connected together by regularly structured signals (called buses), it is possible to utilize datapath regularity in order to achieve significant area savings through FPGA architectural innovations. This paper describes such an FPGA routing architecture, called the multibit routing architecture, which employs busbased connections in order to exploit datapath regularity. It is experimentally shown that, compared to conventional FPGA routing architectures, the multibit routing architecture can achieve 14% routing area reduction for implementing datapath circuits, which represents an overall FPGA area savings of 10%. This paper also empirically determines the best values of several important architectural parameters for the new routing architecture including the most area efficient granularity values and the most area efficient proportion of bus-based connections. Index Terms—Area efficiency, datapath regularity, field-programmable gate arrays (FPGAs), reconfigurable fabric, routing architecture. I. Andy Gean Ye, Jonathan Rose |
FPGA | 2 |
| 2005 | A Verilog RTL Synthesis Tool for Heterogeneous FPGAsabstractModern heterogeneous FPGAs contain "hard" specific-purpose structures such as blocks of memory and multipliers in addition to the completely flexible "soft" programmable logic and routing. These hard structures provide major benefits, yet raise interesting questions in FPGA CAD and architecture. To develop high-quality CAD mapping algorithms for these structures, and indeed to measure the quality of proposed new structures in the architectural domain, it is essential to have a flexible tool at the RTL synthesis level that permits heterogeneous FPGA CAD and architecture experimentation. In this paper we present a synthesis tool, called Odin, and an algorithm that permits flexible targeting of hard structures in FPGAs. Odin maps Verilog designs to two different FPGA CAD flows: Altera's Quartus, and the academic VPR CAD flow. We have expended significant effort to make the quality of this tool comparable to an industrial front-end synthesis tool, and we present mapping results for our benchmarks that show the quality of our results. Peter Jamieson, Jonathan Rose |
FPL | 2 |
| 2005 | Measuring and Utilizing the Correlation Between Signal Connectivity and Signal Positioning for FPGAs Containing Multi-Bit Building BlocksabstractAs the logic capacity of FPGA increases, there has been a corresponding increase in the variety of FPGA building blocks. From a mere collection of the conventional logic blocks, FPGAs now can include digital signal processors, multipliers, multi-bit addressable memory cells, and even processor cores; and one of the common characteristics of these new building blocks is their multi-bit design, where each block is designed specifically to process several bits of data at a time. This multi-bit processing paradigm is significantly different from the single-bit processing design of the conventional FPGA logic blocks; and it creates differentiation in signals through its bussed structures. Consequently, this paper examines the correlation between the positions of the signals in buses and the connectivity of these signals. Based on the correlation measurements, a multi-bit routing architecture is then proposed along with its routing tool. It is experimentally shown that, comparing to the conventional routing architectures, the multi-bit architecture requires 12% less area to implement; and in particular, it needs 27% less routing switches to connect its multi-bit blocks to their routing tracks, and 18% less configuration memory to store the configuration information. Andy Gean Ye, Jonathan Rose |
FPL | 2 |
| 2005 | The Transmogrifier-4: An FPGA-Based Hardware Development System with Multi-Gigabyte Memory Capacity and High Host and Memory Bandwidth
Joshua Fender, Jonathan Rose, David R. Galloway |
FPT | 2 |
| 2004 | A synthesis oriented omniscient manual editorabstractThe cost functions used to evaluate logic synthesis transformations for FPGAs are far removed from the final speed and routability determined after placement, routing and timing analysis. This distance has given rise to the field of physical synthesis, which attempts to improve logic synthesis by employing cost functions that contain placement, routing and/or timing analysis information.In this work we take this notion to an extreme that we call omniscience, in which post-routing timing analysis is provided in the context of a manual editor in which the user selects logical and physical transformations. After each incremental circuit modification, the user is informed of the circuit performance after routing and timing analysis. Since the computations involved in providing this level of information are large, we restrict the application to relatively small circuits, no larger than 1000 logic elements.Using this approach on a commercial FPGA, we propose a set of logic transformations specific to the logic and routing architecture of the Xilinx Virtex-E device. On a set of 10 circuits we have achieved an average performance improvement of 10% when both logical and physical changes are used. Another value of the editor is that it reveals new types of automatable physical-synthesis transformations and optimization strategies that arise from architectural properties of the target device. Tomasz S. Czajkowski, Jonathan Rose |
FPGA | 2 |
| 2004 | Transistor grouping and metal layer trade-offs in automatic tile layout of FPGAsabstractThe physical layout of modern commercial FPGAs is one of the last bastions of manual VLSI layout. Our recent work has automated the FPGA layout process from architectural description to mask-level layout of the repeated FPGA tile. Here we improve on that work using two approaches: 1) by making better choices for the grouping of circuitry into the cells used for the layout and 2) through better allocation of metal. The new groupings improve the FPGA tile area by between 10% and 14%. That, together with the superior metal layer allocation allows us to automatically lay out a very accurate capture of a Xilinx Virtex-E tile that is only 54% to 92% larger than the real thing. With two additional metal layers, our tile area is only 13% larger. In addition, we show that a standard cell implementation is 102% larger than the real Xilinx Virtex-E. Ian Kuon, Aaron Egier, Jonathan Rose |
FPGA | 3 |
| 2004 | Hardware Accelerated Novel Protein Identification
Anish Alex, Jonathan Rose, Ruth Isserlin-Weinberger, Christopher W. V. Hogue |
FPL | 2 |
| 2004 | Using multi-bit logic blocks and automated packing to improve field-programmable gate array density for implementing datapath circuitsabstractAs the logic capacity of field-programmable gate arrays (FPGAs) increases, they are being increasingly used to implement large arithmetic-intensive applications, which often contain a large proportion of datapath circuits. Since datapath circuits usually consist of regularly structured components, called bit-slices, it is possible to utilize datapath regularity in order to achieve significant area savings through FPGA architectural innovations. This work describes such an FPGA logic block architecture, called a multi-bit logic block, which employs configuration memory sharing to exploit datapath regularity. It is experimentally shown that, comparing to conventional FPGA logic blocks, the multi-bit logic blocks can achieve 18% to 26% logic block area reduction for implementing datapath circuits, which represents an overall FPGA area saving of 5% to 13%. A packing algorithm for the multi-bit logic block architecture is also proposed in this paper; and it is used to empirically find the best values for several important architectural parameters of the new architecture, including the most area efficient granularity values and the most area efficient amount of configuration memory sharing. Andy Gean Ye, Jonathan Rose |
FPT | 2 |
| 2004 | Synthetic circuit generation using clustering and iterationabstractThe development of next-generation computer-aided design tools and field programmable gate array architectures require benchmark circuits to experiment with new algorithms and architectures. There has always been a shortage of good public benchmarks for these purposes, and even companies that have access to proprietary customer designs could benefit from designs that meet size and other particular specifications. In this paper, we present a new method of generating realistic synthetic benchmark circuits to help alleviate this shortage. The method significantly improves the quality of previous work by imposing a hierarchy of circuits through clustering and by using a simpler method of characterizing the nature of sequential circuits. Also, in contrast to current constructive generation methods (Hutton et al., 1998), (Hutton et al., 2002), (Darnauer and Dai, 1996), (Iwama and Hino, 1994), (Iwama et al., 1997), (Harlow and Brglez, 1997), (Ghosh et al., 1998), (http://www.cbl.ncsu.edu//spl bsol/-publications //spl bsol/-/spl bsol/#2000-TR@CBL-01-Ghosh), (Pistorius et al., 2000), (Stroobandt et al., 2000), (Verplaetse et al., 2002), we employ new iterative techniques in the generation that provide better control over the generated circuit's characteristics. As in previous work, we assess the realism of the generated circuits by comparing properties of real circuits and generated "clones" of the real circuit after placement and routing. On average, the real and clone circuits' total detailed wirelength differ by only 14%, a major improvement over previous results. In addition, the minimum track count is within 14% and the critical-path delay is within 10%. Paul D. Kundarewich, Jonathan Rose |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2004 | The effect of LUT and cluster size on deep-submicron FPGA performance and densityabstractIn this paper, we revisit the field-programmable gate-array (FPGA) architectural issue of the effect of logic block functionality on FPGA performance and density. In particular, in the context of lookup table, cluster-based island-style FPGAs (Betz et al. 1997) we look at the effect of lookup table (LUT) size and cluster size (number of LUTs per cluster) on the speed and logic density of an FPGA. We use a fully timing-driven experimental flow (Betz et al. 1997), (Marquardt, 1999) in which a set of benchmark circuits are synthesized into different cluster-based (Betz and Rose, 1997, 1998) and (Marquardt, 1999) logic block architectures, which contain groups of LUTs and flip-flops. Across all architectures with LUT sizes in the range of 2 to 7 inputs, and cluster size from 1 to 10 LUTs, we have experimentally determined the relationship between the number of inputs required for a cluster as a function of the LUT size (K) and cluster size (N). Second, contrary to previous results, we have shown that clustering small LUTs (sizes 2 and 3) produces better area results than what was presented in the past. However, our results also show that the performance of FPGAs with these small LUT sizes is significantly worse (by almost a factor of 2) than larger LUTs. Hence, as measured by area-delay product, or by performance, these would be a bad choice. Also, we have discovered that LUT sizes of 5 and 6 produce much better area results than were previously believed. Finally, our results show that a LUT size of 4 to 6 and cluster size of between 3-10 provides the best area-delay product for an FPGA. Elias Ahmed, Jonathan Rose |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | Video-Rate Stereo Depth Measurement on Programmable HardwareabstractThis paper describes the implementation of a stereo depth measurement algorithm in hardware on field programmable gate arrays (FPGAs). This system generates 8 bit sub-pixel disparities on 256 by 360 pixel images at video rate (30 frames/sec). The algorithm implemented is a multi-resolution, multi-orientation phase-based technique called local weighted phase-correlation (Fleet, 1994). Hardware implementation speeds up the performance more than 300 times that of the same algorithm running in software. In this paper, we describe the programmable hardware platform, the base stereo vision algorithm and the design of the hardware. We include various trade-offs required to make the hardware small enough to fit on our system and fast enough to work at video rate. We also show sample outputs from the functioning hardware. Although this paper is specifically focused on phase-based stereo vision FPGA realizations, most of the design issues are common to other DSP and vision applications. Ahmad Darabiha, Jonathan Rose, W. James MacLean |
CVPR (1) | 2 |
| 2003 | Synthetic circuit generation using clustering and iterationabstractThe development of next-generation CAD tools and FPGA architectures requires benchmark circuits to experiment with new algorithms and architectures. There has always been a shortage of good public benchmarks for these purposes, and even companies that have access to proprietary customer designs could benefit from designs that meet size and other particular specifications. In this paper, we present a new method of generating realistic synthetic benchmark circuits to help alleviate this shortage. The method significantly improves the quality of previous work by imposing the natural hierarchy of circuits through clustering and by using a simpler method of characterizing the nature of sequential circuits. Also, in contrast to current constructive generation methods, we employ new iterative techniques in the generation that provide better control over the generated circuit s characteristics. As in previous work, we assess the realism of the generated circuits by comparing properties of real circuits and generated "clones" of the real circuit after placement and routing. On average, the real and clone circuits' total detailed wirelength differed by only 14%, a major improvement over previous results. In addition, the minimum track count was within 14% and the critical path delay was within 10%. Paul D. Kundarewich, Jonathan Rose |
FPGA | 2 |
| 2003 | The StratixTM routing and logic architectureabstractThis paper describes the Altera Stratix logic and routing architecture. The primary goals of the architecture were to achieve high performance and logic density. We give an overview of the entire device, and then focus on the logic and routing architecture. The Stratix logic architecture is based on a cluster of ten 4-input LUTs and its routing consists of staggered routing lines. We describe the development of the routing architecture, including its directional bias, its direct-drive routing which reduces both area and delay. The logic array block and logic cell design is also described, and new routing structures with in the logic array block, and logic element features are described. David M. Lewis, Vaughn Betz, David Jefferson, Andy Lee, Christopher Lane, Paul Leventis, Sandy Marquardt, Cameron McClintock, Bruce Pedersen, Giles Powell, Srinivas Reddy, Chris Wysocki, Richard Cliff, Jonathan Rose |
FPGA | 14 |
| 2003 | Automatic transistor and physical design of FPGA tiles from an architectural specificationabstractOne of the most difficult and time-consuming steps in the creation of an FPGA is its transistor-level design and physical layout. Modern commercial FPGAs typically consume anywhere from 50 to 200 man-years simply in the layout step. To date, automated tools have only been employed in small parts of the periphery and programming circuitry. The core tiles, which are repeated many times, are subject to painstaking manual design and layout. In this paper we present a new system (called GILES, for Good Instant Layout of Erasable Semiconductors) that automatically generates a transistor-level schematic from a high-level architectural specification of an FPGA. It also generates a cell-level netlist that is placed and routed automatically. The architectural specification is the one used as input to the VPR [3] architectural exploration tool. The output is the mask-level layout of a single tile that can be replicated to form an FPGA array. We describe a new placement tool that simultaneously places and compacts the layout to minimize white space and wiring demand, and a special-purpose router built for this task.GILES can place and route a tile consisting of four 4-input LUT logic cells and all of its programmable wires in a 0.18μm CMOS process using 8 layers of metal and 25983μm2 of area. When we generate the layout of an architecture similar to the Xilinx Virtex-E FPGA (built in a 0.18μm process) GILES requires only 47% more area than the original. The layout area of an architecture similar to the Altera Apex 20K400E (also built in a 0.18µm process) constructed by GILES requires 97% more area than the original. Ketan Padalia, Ryan Fung, Mark Bourgeault, Aaron Egier, Jonathan Rose |
FPGA | 5 |
| 2003 | A high-speed ray tracing engine built on a field-programmable systemabstractRay tracing is a method of rendering high-quality images and video by calculating what happens to virtual light rays in a 3-dimensional scene. It is capable of creating for more realism than traditional Z-buffering methods. This paper describes the design of a hardware ray tracing system implemented on a multi-FPGA Xilinx Virtex-E prototyping system. The result is a hardware ray tracer that is capable of out-performing a 2.4GHz Pentium 4, running a well-known high performance software ray tracing algorithm, by up to a factor of thirty. When these results are projected forward into a next generation FPGA system, consisting of a single large Virtex 2 Pro FPGA, it is found that the system should be able to out perform the same Pentium 4 by up to two orders of magnitude, and the fastest known hardware implementation, the AR350, by up to a factor of three. Joshua Fender, Jonathan Rose |
FPT | 2 |
| 2003 | A parameterized automatic cache generator for FPGAsabstractCaches in FPGAs can improve the performance of soft processors and other applications beset by slow storage components. In this paper we present a cache generator which can produce caches with a variety of associativities, latencies, and dimensions. This tool allows system designers to effortlessly create, and investigate different caches in order to better meet the needs of their target system. The effect of these three parameters on the area and speed of the caches is also examined and we show that the designs can meet a wide range of specifications and are in general fast and compact. Peter Yiannacouras, Jonathan Rose |
FPT | 2 |
| 2002 | EVE: a CAD tool for manual placement and pipelining assistance of FPGA circuitsabstractAs FPGAs push ever deeper into mainstream digital design, there is an increasing desire for high-performance circuits. This paper describes a manual editor, called EVE, which can assist a designer to perform manual packing, placement and pipelining of commercial FPGA circuits to achieve a meaningful increase in performance. This effort is inspired by Von Herzen's paper [15] [16], which proposed the notion of an "Event Horizon" - a high-speed circuit design approach in which complete knowledge of the timing effect of every synthesis change is used. It is very laborious to implement circuits using this approach; therefore we try to augment manual design tools in order to make this Event Horizon methodology easier to perform. This paper describes a first step in that direction, which focuses on placement, packing and pipelining. EVE provides an interactive environment that immediately reroutes and timing analyzes after each user circuit modification, giving an exact value for critical path delay. It can also suggest good placement positions and provide flip-flop insertion assist during pipelining. Compared to a state-of-the-art Synthesis and place and route flow, we used EVE to achieve an average of 12.7% higher operating frequency on a set of eight Xilinx Virtex-E circuits of 250 or fewer LUTs. William Chow, Jonathan Rose |
FPGA | 2 |
| 2002 | Synthesizing datapath circuits for FPGAs with emphasis on area minimizationabstractLarge circuits, whether they are arithmetic, digital signal processing, switching, or processors, typically contain a greater portion of highly regular datapath logic. Datapath synthesis algorithms preserve these regular structures, so they can be exploited by packing, placement, and routing tools for speed or density. Typical datapath synthesis algorithms, however, sacrifice area to gain regularity. Current algorithms can have as much as 30% to 40% area inflation when compared with traditional flat synthesis algorithms. This paper describes a datapath synthesis algorithm with very low area overhead, which is an enhancement to the module compaction algorithm. We propose two word-level optimizations - multiplexer tree collapsing and operation reordering. They reduce the area inflation to 3%-8% as compared with flat synthesis. Our synthesis results also retain significant amount of regularity from the original designs. Andy Gean Ye, Jonathan Rose, David M. Lewis |
FPT | 2 |
| 2002 | Automatic generation of synthetic sequential benchmark circuitsabstractThe design of programmable logic architectures and supporting computer-aided design tools fundamentally requires both a good understanding of the combinatorial nature of netlist graphs and sufficient quantities of realistic examples to evaluate or benchmark the results. In this paper, the authors investigate these two issues. They introduce an abstract model for describing sequential circuits and a collection of statistical parameters for better understanding the nature of circuits. Based upon this model they introduce and formally define the signature of a circuit netlist and the signature equivalence of netlists. They give an algorithm (GEN) for generating sequential benchmark netlists, significantly expanding previous work (Hutton et al, 1998) which generated purely combinational circuits. By comparing synthetic circuits to existing benchmarks and random graphs they show that GEN circuits are significantly more realistic than random graphs. The authors further illustrate the viabilty of the methodology by applying GEN to a case study comparing two partitioning algorithms. Mike Hutton, Jonathan Rose, Derek G. Corneil |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2001 | Panel: (When) Will FPGAs Kill ASICs?abstractThere was a time - in the dim historical past - when foundries actually made ASICs with only 5000 to 50,000 logic gates. But FPGAs and CPLDs conquered those markets and pushed ASIC silicon toward opportunities with more logic, volume, and speed. Today's largest FPGAs approach the few-million-gate size of a typical ASIC design, and continue to sprout embedded cores, such as CPUs, memories, and interfaces. And given the risks of nonworking nanometer silicon, FPGA costs and time-to-market are looking awfully attractive. So, will FPGAs kill ASICs? ASIC technologists certainly think not. ASICs are themselves sprouting patches of programmable FPGA fabric, and pushing new realms of size and especially speed. New tools claim to have tamed the convergence problems of older ASIC flows. Is the future to be found in a market full of FPGAs with ASIC-like cores? ASICs with FPGA cores? Other exotic hybrids? Our panelists will share their disagreements on these prognostications. Rob A. Rutenbar, Max Baron, Thomas Daniel, Rajeev Jayaraman, Zvi Or-Bach, Jonathan Rose, Carl Sechen |
DAC | 6 |
| 2001 | Mixing buffers and pass transistors in FPGA routing architecturesabstractThe routing architecture of an FPGA consists of the length of the wires, the type of switch used to connect wires (buffered, unbuffered, fast or slow) and the topology of the interconnection of the switches and wires. FPGA routing architecture has a major influence on the logic density and speed of FPGA devices. Previ?ous work [] based on a 0.35um CMOS process has suggested that an architecture consisting of length 4 wires (where the length of a wire is measured in terms of the number of logic blocks it passes before being switched) and half of the programmable switches are active buffers, and half are pass transistors. In that work, however, the topology of the routing architecture prevented buffered tracks from connecting to pass-transistor tracks. This restriction prevents the creation of interconnection trees for high fanout nets that have a mixture of buffers and pass transistors. Electrical simulations sug?gest that connections closer to the leaves on interconnection trees are faster using pass transistors, but it is essential to buffer closer to the source. This latter effect is well known in regular ASIC routing [2]. Mike Sheng, Jonathan Rose |
FPGA | 2 |
| 2001 | Structural analysis and generation of synthetic digital circuits with memoryabstractOne of the most difficult aspects of experimental reconfigurable architecture or computer-aided design (CAD) tool research is obtaining sufficiently large benchmark circuits. One approach to obtaining such circuits is to generate them stochastically. Current circuit generators construct combinational and sequential logic circuits. Many of today's devices, however, are being used to implement entire systems, and often these systems contain on-chip storage. This paper describes a circuit generator that constructs circuits containing significant amounts of memory. To ensure the circuits are realistic, we have performed a detailed structural analysis of such circuits; this analysis is also described in this paper. Steve Wilton, Jonathan Rose, Zvonko G. Vranesic |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2000 | The effect of LUT and cluster size on deep-submicron FPGA performance and densityabstractWe use a fully timing-driven experimental flow [4] [15] in which a set of benchmark circuits are synthesized into different cluster-based [2] [3] [15] logic block architectures, which contain groups of LUTs and flip-flops. We look across all architectures with LUT sizes in the range of 2 inputs to 7 inputs, and cluster size from 1 to 10 LUTs. In order to judge the quality of the architecture we do both detailed circuit level design and measure the demand of routing resources for every circuit in each architecture. Elias Ahmed, Jonathan Rose |
FPGA | 2 |
| 2000 | Automatic generation of FPGA routing architectures from high-level descriptionsabstractIn this paper we present a “high-level” FPGA architecture description language which lets FPGA architects succinctly and quickly describe an FPGA routing architecture. We then present an “architecture generator” built into the VPR CAD tool [1, 2] that converts this high-level architecture description into a detailed and completely specified flat FPGA architecture. This flat architecture is the representation with which CAD optimization and visualization modules typically work. By allowing FPGA researchers to specify an architecture at a high-level, an architecture generator enables quick and easy “what-if” experimentation with a wide range of FPGA architectures. The net effect is a more fully optimized final FPGA architecture. In contrast, when FPGA architects are forced to use more traditional methods of describing an FPGA (such as the manual specification of every switch in the basic file of the FPGA), far less experimentation can be performed in the same time, and the architectures experimented upon are likely to be highly similar, leaving important parts of the design space completely unexplored. Vaughn Betz, Jonathan Rose |
FPGA | 2 |
| 2000 | Timing-driven placement for FPGAsabstractIn this paper we introduce a new Simulated Annealing-based timing-driven placement algorithm for FPGAs. This paper has three main contributions. First, our algorithm employs a novel method of determining source-sink connection delays during placement. Second, we introduce a new cost function that trades off between wire-use and critical path delay, resulting in significant reductions in critical path delay without significant increases in wire-use. Finally, we combine connection-based and path-based timing-analysis to obtain an algorithm that has the low time-complexity of connection-based timing-driven placement, while obtaining the quality of path-based timing-driven placement. Alexander Marquardt, Vaughn Betz, Jonathan Rose |
FPGA | 3 |
| 2000 | Real-time, frame-rate face detection on a configurable hardware system (poster abstract)
Rob McCready, Jonathan Rose |
FPGA | 2 |
| 2000 | A novel and efficient routing architecture for multi-FPGA systemsabstractMulti-FPGA systems (MFSs) are used as custom computing machines, logic emulators and rapid prototyping vehicles. A key aspect of these systems is their programmable routing architecture which is the manner in which wires, FPGAs and field-programmable interconnect devices (FPIDs) are connected. Several routing architectures for MFSs have been proposed, and previous research has shown that the partial crossbar is one of the best existing architectures. In this paper, we propose a new routing architecture, called the hybrid complete-graph and partial-crossbar (HCGP) which has superior speed and cost compared to a partial crossbar. The new architecture uses both hard-wired and programmable connections between the FPGAs. We compare the performance and cost of the HCGP and partial crossbar architectures experimentally, by mapping a set of 15 large benchmark circuits into each architecture. A customized set of partitioning and interchip routing tools were developed, with particular attention paid to architecture-appropriate interchip routing algorithms. We show that the cost of the partial crossbar (as measured by the number of pins on all FPGAs and FPIDs required to fit a design), is on average 20% more than the new HCGP architecture and as much as 25% more. Furthermore, the critical path delay for designs implemented on the partial crossbar were on average 20% more than the HCGP architecture and up to 43% more. Using our experimental approach, we also explore a key architecture parameter associated with the HCGP architecture-the proportion of hard-wired connections versus programmable connections-to determine its best value. Mohammed A. S. Khalid, Jonathan Rose |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2000 | Speed and area tradeoffs in cluster-based FPGA architecturesabstractOne way to reduce the delay and area of field-programmable gate arrays (FPGAs) is to employ logic-cluster-based architectures, where a logic cluster is a group of logic elements connected with high-speed local interconnections. In this paper, we empirically evaluate FPGA architectures with logic clusters ranging in size from 1 to 20, and show that compared to architectures with size 1 clusters, architectures with size 8 clusters have 23% less delay (30% faster clock speed) and require 14% less area. We also show that FPGA architectures with large cluster sizes can significantly reduce design compile time-an increasingly important concern as the logic capacity of FPGA's rises. For example, an architecture that uses size 20 clusters requires seven times less compile time than an architecture with size 1 clusters. Alexander Marquardt, Vaughn Betz, Jonathan Rose |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 1999 | FPGA Routing Architecture: Segmentation and Buffering to Optimize Speed and DensityabstractIn this work we investigate the routing architecture of FPGAs, focusing primarily on determining the best distribution of routing segment lengths and the best mix of pass transistor and tri-state buffer routing switches. While most commercial FPGAs contain many length 1 wires (wires that span only one logic block) we find that wires this short lead to FPGAs that are inferior in terms of both delay and routing area. Our results show instead that it is best for FPGA routing segments to have lengths of 4 to 8 logic blocks. We also show that 50% to 80% of the routing switches in an FPGA should be pass transistors, with the remainder being tri-state buffers. Architectures that employ the best segmentation distributions and the best mixes of pass transistor and tri-state buffer switches found in this paper are not only 11% to 18% faster than a routing architecture very similar to that of the Xilinx XC4000X but also considerably simpler. These results are obtained using an architecture investigation infrastructure that contains a fully timing-driven router and detailed area and delay models. Vaughn Betz, Jonathan Rose |
FPGA | 2 |
| 1999 | Using Cluster-Based Logic Blocks and Timing-Driven Packing to Improve FPGA Speed and DensityabstractIn this papel; we investigate the speed and area-eficiency of FPGAs employing "logic clusters" containing multiple LUTs and registers as their logic block.We introduce a new, timing-driven tool (T-VPack) to "pack" LUTs and registers into these logic clusters, and we show that this algorithm is superior to an existing packing algorithm.Then, using a realistic routing architecture and sophisticated delay and area models, we empirically evaluate FPGAs composed of clusters ranging in size from one to twenty LUTs, and show that clusters of size seven through ten provide the best area-delay trade-o@ Compared to circuits implemented in an FPGA composed of size one clusters, circuits implemented in an FPGA with size seven clusters have 30% less delay (a 43% increase in speed) and require 8% less area, and circuits implemented in an FPGA with size ten clusters have 34% less delay (a 52% increase in speed), and require no additional area. Alexander Marquardt, Vaughn Betz, Jonathan Rose |
FPGA | 3 |
| 1999 | Trading Quality for Compile Time: Ultra-Fast Placement for FPGAsabstractArticle Trading quality for compile time: ultra-fast placement for FPGAs Share on Authors: Yaska Sankar Department of Electrical and Computer Engineering, University of Toronto, Toronto, ON, Canada M5S 3G4 Department of Electrical and Computer Engineering, University of Toronto, Toronto, ON, Canada M5S 3G4View Profile , Jonathan Rose Department of Electrical and Computer Engineering, University of Toronto, Toronto, ON, Canada M5S 3G4 Department of Electrical and Computer Engineering, University of Toronto, Toronto, ON, Canada M5S 3G4View Profile Authors Info & Claims FPGA '99: Proceedings of the 1999 ACM/SIGDA seventh international symposium on Field programmable gate arraysFebruary 1999 Pages 157–166https://doi.org/10.1145/296399.296449Online:01 February 1999Publication History 68citation515DownloadsMetricsTotal Citations68Total Downloads515Last 12 Months7Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Yaska Sankar, Jonathan Rose |
FPGA | 2 |
| 1999 | The design of an SRAM-based field-programmable gate array. I. ArchitectureabstractField-programmable gate arrays (FPGAs) are now widely used for the implementation of digital systems, and many commercial architectures are available. Although the literature and data books contain detailed descriptions of these architectures, there is very little information on how the high-level architecture was chosen, and no information on the circuit-level or physical design of the devices. This paper describes the high-level architectural design of a static-random-access memory programmable FPGA. A forthcoming Part II will address the circuit design issues through to the physical layout. The logic block and routing architecture of the FPGA was determined through experimentation with benchmark circuits and custom-built computer-aided design tools. The resulting logic block is an asymmetric tree of four-input lookup tables that are hard-wired together and a segmented routing architecture with a carefully chosen segment length distribution. Paul Chow, Soon Ong Seo, Jonathan Rose, Kevin Chung, Gerard Páez-Monzón, Immanuel Rahardja |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 1999 | The design of a SRAM-based field-programmable gate array-Part II: Circuit design and layoutabstractFor Pt.I see ibid., vol.7, pp.191-7 (1999). Field-programmable gate arrays (FPGA's) are now widely used for the implementation of digital systems, and many commercial architectures are available. Although the literature and data books contain detailed descriptions of these architectures, there is very little information on how the high-level architecture was chosen and no information on the circuit-level or physical design of the devices. In Part I of this paper, we described the high-level architectural design of a static random-access memory programmable FPGA. This paper will address the circuit-design issues through to the physical layout. We address area-speed tradeoffs in the design of the logic block circuits and in the connections between the logic and the routing structure. All commercial FPGA designs are done using full-custom hand layout to obtain absolute minimum die sizes. This is both labor and time intensive. We propose a design style with a minitile that contains a portion of all the components in the logic tile, resulting in less full-custom effort. The minitile is replicated in a 4/spl times/4 array to create a macro tile. The minitile is optimized for layout density and speed, and is customized in the array by adding appropriate vias. This technique also permits easy changing of the hard-wired connections in the logic block architecture and the segmentation length distribution in the routing architecture. Paul Chow, Soon Ong Seo, Jonathan Rose, Kevin Chung, Gerard Páez-Monzón, Immanuel Rahardja |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 1999 | The memory/logic interface in FPGAs with large embedded memory arraysabstractAs the capacities of field-programmable gate arrays (FPGAs) grow, they will be used to implement much larger circuits than ever before. These larger circuits often require significant amounts of storage. In order to address these storage requirements, FPGAs with large embedded memory arrays are now being developed by several vendors. One of the crucial components of an FPGA with on-chip memory is the routing structure between the memory arrays and logic resources. If this memory/logic interface is not flexible enough, many circuits will be unroutable, while if it is too flexible, it will be slower and consume more chip area than is necessary. In this paper, we show that an interconnect in which each memory pin can connect to between four and seven logic routing tracks is best in terms of both area and speed. We also show that by adding switches to support nets that connect multiple memory arrays, we can reduce the memory access time by up to 25% and improve the routability slightly. Steve Wilton, Jonathan Rose, Zvonko G. Vranesic |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1998 | A Hybrid Complete-Graph Partial-Crossbar Routing Architecture for Multi-FPGA SystemsabstractMulti-FPGA systems (MFSs) are used as custom computing machines, logic emulators and rapid prototyping vehicles. A key aspect of these systems is their programmable routing architecture; the manner in which wires, FPGAs and Field-Programmable Interconnect Devices (FPIDs) are connected. Several routing architectures for MFSs have been proposed [Arno92] [Butt92] [Hauc94] [Apti96] [Vuil96] and previous research has shown that the partial crossbar is one of the best existing architectures [Kim96] [Khal97]. In this paper we propose a new routing architecture, called the Hybrid Complete-Graph and Partial-Crossbar (HCGP) which has superior speed and cost compared to a partial crossbar. The new architecture uses both hard-wired and programmable connections between the FPGAs.We compare the performance and cost of the HCGP and partial crossbar architectures experimentally, by mapping a set of 15 large benchmark circuits into each architecture. A customized set of partitioning and inter-chip routing tools were developed, with particular attention paid to architecture-appropriate inter-chip routing algorithms. We show that the cost of the partial crossbar (as measured by the number of pins on all FPGAs and FPIDs required to fit a design), is on average 20% more than the new HCGP architecture and as much as 35% more. Furthermore, the critical path delay for designs implemented on the partial crossbar increased, and were on average 9% more than the HCGP architecture and up to 26% more.Using our experimental approach, we also explore a key architecture parameter associated with the HCGP architecture: the proportion of hard-wired connections versus programmable connections, to determine its best value. Mohammed A. S. Khalid, Jonathan Rose |
FPGA | 2 |
| 1998 | Constraints from Hell: How to Tell Makes a Good FPGA (Panel)abstractThe FPGA development is an extraordinarily complex task, involving many people working in architecture, chip design, software, marketing and production. Each tends to focus on their immediate task, and have their own measures of goodness for what they do, as well as very particular constraints relevant to their discipline. In this panel we will discuss these measures and constraints in an attempt to see how they succeed in transcending the artificial borders in each discipline. Jonathan Rose, Sinan Kaptanoglu, Clive McCarthy, Rob Smith, Sandip Vij |
FPGA | 1 |
| 1998 | A Fast Routability-Driven Router for FPGAsabstractThree factors are driving the demand for rapid FPGA compilation. First, as FPGAs have grown in logic capacity, the compile computation has grown more quickly than the compute power of the available computers. Second, there exists a subset of users who are willing to pay for very high speed compile with a decrease in quality of result, and accordingly being required to use a larger FPGA or use more real-estate on a given FPGA than is otherwise necessary. Third, very high speed compile has been a long-standing desire of those using FPGA-based custom computing machines, as they want compile times at least closer to those of regular computers. This paper focuses on the routing phase of the compile process, and in particular on routability-driven routing (as opposed to timing-driven routing). We present a routing algorithm and routing tool that has three unique capabilities relating to very high-speed compile: 1. For a “low stress ” routing problem (which we define as the case where the track supply is at least 10 % greater than the minimum number of tracks per channel actually needed to route a circuit) the routing time is very fast. For example, the routing phase (after the netlist is parsed and the routing graph is constructed) for a 20,000 LUT/FF pair circuit with 30 % extra tracks is only 23 seconds on a 300 MHz Sparcstation. 2. For low-stress routing problems the routing time is nearlinear in the size of the circuit, and the linearity constant is very small: 1.1 ms per LUT/FF pair, or roughly 55,000 LUT/FF pairs per minute. 3. For more difficult routing problems (where the track supply is close to the minimum needed) we provide a method that quickly identifies and subdivides this class into two sub-classes: (i) those circuits which are difficult (but possible) to route and will take significantly more time than low-stress problems, and (ii) those circuits which are impossible to route. In the first case the user can choose to continue or reduce the amount of logic; in the second case the user is forced to reduce the amount of logic or obtain a larger FPGA. 1. Jordan S. Swartz, Vaughn Betz, Jonathan Rose |
FPGA | 3 |
| 1998 | Characterization and parameterized generation of synthetic combinational benchmark circuitsabstractThe development of new field-programmed, mask-programmed, and laser-programmed gate-array architectures is hampered by the lack of realistic test circuits that exercise both the architectures and their automatic placement and routing algorithms. In this paper, we present a method and a tool for generating parameterized and realistic synthetic circuits. To obtain the realism, we propose a set of graph-theoretic characteristics that describe a physical netlist, and have built a tool that can measure these characteristics on existing circuits. The generation tool uses the characteristics as constraints in the synthetic circuit generation. To validate the quality of the generated netlists, parameters that are not specified in the generation are compared with those of real circuits and with those of more "random" graphs. Mike Hutton, Jonathan Rose, Jerry P. Grossman, Derek G. Corneil |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1998 | Effect of the prefabricated routing track distribution on FPGA area-efficiencyabstractIn most commercial field programmable gate arrays (FPGA's) the number of wiring tracks in each channel is the same across the entire chip. A long-standing open question for both FPGA's and channeled gate arrays is whether or not some uneven distribution of routing tracks across the chip would lead to an area benefit. For example, many circuit designers intuitively believe that most congestion occurs near the center of a chip, and hence expect that having wider routing channels near the chip center would be beneficial. In this paper, we determine the relative area-efficiency of several different routing track distributions. We first investigate FPGA's in which horizontal and vertical channels contain different numbers of tracks in order to determine if such a directional bias provides a density advantage. Second, we examine routing track distributions in which the track capacities vary from channel to channel. We compare the area efficiency of these nonuniform routing architectures to that of an FPGA with uniform channel capacities across the entire chip. The main result is that the most area-efficient global routing architecture is one with uniform (or very nearly uniform) channel capacities across the entire chip in both the horizontal and vertical directions. This paper shows why this result, which is contrary to the intuition of many FPGA architects, is true. While a uniform routing architecture is the most area-efficient, several nonuniform and directionally biased architectures are fairly area-efficient provided that appropriate choices are made for the pin positions on the logic blocks and the logic block array aspect ratio. Vaughn Betz, Jonathan Rose |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1998 | The Transmogrifier-2: a 1 million gate rapid-prototyping systemabstractThis paper describes the Transmogrifier-2 (TM-2), a second-generation multifield programmable gate array (FPGA) rapid-prototyping system. The largest version of the system will comprise 16 boards that each contain two Altera 10K50 FPGA's, four I-Cube interconnect chips, and up to 8 Mbytes of memory. The inter-FPGA routing architecture of the TM-2 uses a novel interconnect structure, a nonuniform partial crossbar, that provides a constant delay between any two FPGA's in the system. The TM-2 architecture is modular and scalable, meaning that systems of various sizes can be constructed from copies of the same board, while maintaining routability and the constant delay feature. Other features include a system-level programmable clock that allows single-cycle access to off-chip memory, and programmable clock waveforms with edge resolution of 10 ns. The first Transmogrifier-2 boards have been manufactured and are functional. They have recently been used successfully in some simple graphics acceleration applications. David M. Lewis, David R. Galloway, Marcus van Ierssel, Jonathan Rose, Paul Chow |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 1997 | Generation of Synthetic Sequential Benchmark CircuitsabstractAbstract—The design of programmable logic architectures and supporting computer-aided design tools fundamentally requires both a good understanding of the combinatorial nature of netlist graphs and sufficient quantities of realistic examples to evaluate or benchmark the results. In this paper, the authors investigate these two issues. They introduce an abstract model for describing sequential circuits and a collection of statistical parameters for better understanding the nature of circuits. Based upon this model they introduce and formally define the signature of a circuit netlist and the signature equivalence of netlists. They give an algorithm (GEN) for generating sequential benchmark netlists, significantly expanding previous work (Hutton et al., 1998) which generated purely combinational circuits. By comparing synthetic circuits to existing benchmarks and random graphs they show that GEN circuits are significantly more realistic than random graphs. The authors further illustrate the viabilty of the methodology by applying GEN to a case study comparing two partitioning algorithms. Index Terms—Benchmark, digital circuits, placement. I. Mike Hutton, Jonathan Rose, Derek G. Corneil |
FPGA | 2 |
| 1997 | The Transmogrifier-2: A 1 Million Gate Rapid Prototyping SystemabstractThis paper describes the Transmogrifier-2, a second generation multi-FPGA system. The largest version of the system will comprise 16 boards that each contain two Altera 10K50 FPGAs, four I-cube interconnect chips, and up to 8 Mbytes of memory. The inter-FPGA routing architecture of the TM-2 uses a novel interconnect structure, a non-uniform partial crossbar, that provides a constant delay between any two FPGAs in the system. The TM-2 architecture is modular and scalable, meaning that various sized systems can be constructed from the same board, while maintaining routability and the constant delay feature. Other features include a system-level programmable clock that allows single-cycle access to off-chip memory, and programmable clock waveforms with resolution to 10ns. The first Transmogrifier-2 boards have been manufactured and are functional. They have recently been used successfully in some simple graphics acceleration applications. David M. Lewis, David R. Galloway, Marcus van Ierssel, Jonathan Rose, Paul Chow |
FPGA | 4 |
| 1997 | Architectural and Physical Design Challenges for One-Million Gate FPGAs and BeyondabstractProcess technology advances tell us that the one-million gate Field-Programmable Gate Array (FPGA) will soon be here, and larger devices shortly after that. We feel that current architectures will not extend directly to this scale because: they do not handle routing delays effectively; they require excessive compile/place/route times; and because they do not exploit new opportunities presented by the increase in available transistors and wiring. In this paper we describe several challenges that will need to be solved for these large-scale FPGAs to realize their full potential. 1. Jonathan Rose, Dwight D. Hill |
FPGA | 1 |
| 1997 | Memory-to-Memory Connection Structures in FPGAs with Embedded Memory ArraysabstractThis paper shows that the speed of FPGAs with large embedded memory arrays can be improved by adding direct programmable connections between the memories. Nets that connect to multiple memory arrays are often difficult to route, and are often part of the critical path of circuit implementations. The memory-to-memory connection structure proposed in this paper allows for the efficient implementation of these nets, resulting in a reduction in memory access time of up to 25% and a slight improvement in routability. 1 Introduction As FPGAs become larger, they will be used to implement entire systems, rather than small logic subcircuits. One of the key differences between these large systems and the smaller logic subcircuits is that the systems often contain memory. Architectural support for the efficient implementation of memory in nextgeneration FPGAs, therefore, is crucial. Several vendors offer FPGAs with architectural support for memory [1, 2, 3, 4, 5, 6, 7]. The memory resources in ... Steve Wilton, Jonathan Rose, Zvonko G. Vranesic |
FPGA | 2 |
| 1996 | Characterization and Parameterized Random Generation of Digital CircuitsabstractThe development of new Field-Programmed, Mask-Programmed and Laser-Programmed Gate Array architectures is hampered by the lack of realistic test circuits that exercise both the architectures and their automatic placement and routing algorithms.In this paper, we present a method and a tool for generating parameterized and realistic random circuits.To obtain the realism, we propose a set of graph-theoretic characteristics that describe a physical netlist, and have built a tool that can measure these characteristics on existing circuits.The generation tool uses the characteristics as constraints in the random circuit generation.To validate the quality of the generated netlists, parameters that are not speci ed in the generation are c ompared with those of real circuits, and with those of \random" graphs.rameter which i s a c haracteristic of the circuit in question. Circuit Characterization Circuit Generation ValidationCharacteristics and Measured Parameters (n, nPI, nPO, delay, shape, edge-length, fanout dist'n) 2 Circuit Characterization This section describes some of the statistical and structural characteristics of circuits which w e h a v e identi ed.For the purposes of this paper we focus on combinational circuits only, and have used the MCNC benchmark circuits to form the Mike Hutton, Jerry P. Grossman, Jonathan Rose, Derek G. Corneil |
DAC | 3 |
| 1996 | Directional bias and non-uniformity in FPGA global routing architecturesabstractWe investigate the effect of the prefabricated routing track distribution on the area-efficiency of FPGAs. The first question we address is whether horizontal and vertical channels should contain the same number of tracks (capacity), or if there is a density advantage with a directional bias. Secondly, should the channels have a uniform capacity, or is there an advantage when capacities vary from channel to channel? The key result is that the most area-efficient global routing architecture is one with uniform (or very nearly uniform) channel capacities across the entire chip in both the horizontal and vertical directions. Several non-uniform and directionally-biased architectures, however are fairly area-efficient provided that appropriate choices are made for the pin positions on the logic blocks and the logic array aspect ratio. Vaughn Betz, Jonathan Rose |
ICCAD | 2 |
| 1995 | Using Architectural "Families" to Increase FPGA Speed and DensityabstractIn order to narrow the speed and density gap between FPGAs and MPGAs we propose the development of “families” of FPGAs. Each FPGA family is targeted at a single maximum logic capacity, and consists of several “siblings”, or FPGAs of different yet complementary architectures. Any given application circuit is implemented in the sibling with the most appropriate architecture. With properly chosen siblings, one can develop a family of FPGAs which will have better speed and density than any single FPGA. We apply this concept to create two different FPGA families, one composed of architectures with different types of hard-wired logic blocks and the other created from architectures with different types of heterogeneous logic blocks. We found that a family composed of eight chips with different hard-wired logic block architectures simultaneously improves density by 12 to 14% and speed by 18 to 20% over the best single hard-wired FPGA. Vaughn Betz, Jonathan Rose |
FPGA | 2 |
| 1995 | Architecture of Centralized Field-Configurable MemoryabstractAs the capacities of FPGAs grow, it becomes feasible to implement the memory portions of systems directly on an FPGA together with logic. We believe that such an FPGA must contain specialized architectural support in order to implement memories efficiently. The key feature of such architectural support is that it must be flexible enough to accommodate many different memory shapes (widths and depths) as well as allowing different numbers of independently-addressed memory blocks. This paper describes a family of centralized Field-Configurable Memory architectures which consist of a number of memory arrays and dedicated mapping blocks to combine these arrays. We also present a method for comparing these architectures, and use this method to examine the tradeoffs involved in choosing the array size and mapping block capabilities. Steve Wilton, Jonathan Rose, Zvonko G. Vranesic |
FPGA | 2 |
| 1994 | Definition and solution of the memory packing problem for field-programmable systems
David Karchmer, Jonathan Rose |
ICCAD | 2 |
| 1993 | Logic Emulation: A Niche or a Future Standard for Design Verification? (Panel Abstract)abstractLogic Emulations systems consist of boards of re-programmable Field-Programmable Gate Arrays (FPGAs) programmable interconnect, and large software synthesis and analysis packages. They allow ASIC and system-level designs to be represented in hardware that performs at a significant fraction of the real device speeds. They can be used in the end-product for system-level debugging and analysis. As a result, they present an opportunity to perform faster and more accurate design verification. This panel will focus on the role of logic emulation in design verification and will consider the following questions: - Do the benefits of time-to-market and better end product quality justify the investment in logic emulation? - How do logic emulation systems compare with other quick prototyping methods, such as: 1)Build-your-own FPGA prototype on a printed-circuit board, and 2)Build-your-own FPGA prototype using Field-Programmable Interconnect Chips/Device - How does logic emulation compare with traditional simulation in speed, capacity and convenience, including the newer compiled-code simulators? - What is the relationship with special-purpose hardware simulation accelerators? - Will logic emulation be used to verify designs at architectural level? - Does logic emulation only apply to certain types of applications? - Will the emulation speed become a critical barrier to market acceptance? - What will be the effect of declining emulation cost? - What role does logic emulation play in software/hardware co-design? Jonathan Rose |
DAC | 1 |
| 1993 | Architecture of field-programmable gate arraysabstractA survey of field-programmable gate array (FPGA) architectures and the programming technologies used to customize them is presented. Programming technologies are compared on the basis of their volatility, size parasitic capacitance, resistance, and process technology complexity. FPGA architectures are divided into two constituents: logic block architectures and routing architectures. A classification of logic blocks based on their granularity is proposed, and several logic blocks used in commercially available FPGAs are described. A brief review of recent results on the effect of logic block granularity on logic density and performance of an FPGA is then presented. Several commercial routing architectures are described in the context of a general routing architecture model. Finally, recent results on the tradeoff between the flexibility of an FPGA routing architecture, its routability, and its density are reviewed.> Jonathan Rose, Abbas El Gamal, Alberto L. Sangiovanni-Vincentelli |
Proc. IEEE | 1 |
| 1993 | Synthesis method for field programmable gate arraysabstractLogic synthesis algorithms and methods for field-programmable gate arrays (FPGAs) are reviewed. The three most popular types of FPGA architectures are considered, namely, those using logic blocks based on lookup-tables, multiplexers, and wide AND/OR arrays, respectively. The emphasis is on tools that attempt to minimize the area of the combinational logic part of a design, since little work has been done on optimizing performance or routability, or on synthesis of the sequential part of a design. The different tools surveyed are compared using a suite of benchmark designs.> Alberto L. Sangiovanni-Vincentelli, Abbas El Gamal, Jonathan Rose |
Proc. IEEE | 3 |
| 1993 | A stochastic model to predict the routability of field-programmable gate arraysabstractOne area of particular importance is the design of an FPGA routing architecture, which houses the user-programmable switches and wires that are used to interconnect the FPGAs logic resources. Because the routing switches consume significant chip area and introduce propagation delays, the design of the routing architecture greatly influences both the area utilization and speed performance of an FPGA. FPGA routing architectures have already been studied using experimental techniques. This paper describes a stochastic model that facilitates exploration of a wide range of FPGA routing architectures using a theoretical approach. In the stochastic model an FPGA is represented as an N*N array of logic blocks separated by both horizontal and vertical routing channels, similar to a Xilinx FPGA. A circuit to be routed is represented by additional parameters that specify the total number of connections, and each connection's length and trajectory. The stochastic model gives an analytic expression for the routability of the circuit in the FPGA. Practically speaking, routability can be viewed as the likelihood that a circuit can be successfully routed in a given FPGA. The routability predictions from the model are validated by comparing them with the results of a previously published experimental study on FPGA routability.> Stephen Brown 0003, Jonathan Rose, Zvonko G. Vranesic |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1992 | TEMPT: Technology Mapping for the Exploration of FPGA Architectures with Hard-Wired Connections
Kevin Chung, Jonathan Rose |
DAC | 2 |
| 1992 | Improving FPGA Routing Architectures Using Architecture and CAD InteractionsabstractThe interactions between the CAD tools that are used to configure the routing resources of a field-programmable gate array (FPGA) and the design of the routing architecture itself are examined. Such an understanding is used to determine where to reduce the number of routing switches in the FPGA while maintaining routability. Experiments are used to study a switch block that was previously thought to have unacceptably low flexibility. It is shown that the performance of this switch block can be improved by adapting the global router to require less flexibility in the architecture, and by careful placement of physical pins on the logic blocks. It is demonstrated that the fewest routing switches are required when each logical pin appears on only one side of the logic cell rather than two or more.> Benjamin Tseng, Jonathan Rose, Stephen Brown 0003 |
ICCD | 2 |
| 1992 | A detailed router for field-programmable gate arraysabstractA detailed routing algorithm, called the coarse graph expander (CGE), that has been designed specifically for field-programmable gate arrays (FPGAs) is described. The algorithm approaches this problem in a general way, allowing it to be used over a wide range of different FPGA routing architectures. It addresses the issue of scarce routing resources by considering the side effects that the routing of one connection has on another, and also has the ability to optimize the routing delays of time-critical connections. CGE has been used to obtain excellent routing results for several industrial circuits implemented in FPGAs with various routing architectures. The results show that CGE can route relatively large FPGAs in very close to the minimum number of tracks as determined by global routing, and it can successfully optimize the routing delays of time-critical connections. CGE has a linear run time over circuit size.> Stephen Brown 0003, Jonathan Rose, Zvonko G. Vranesic |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1991 | Chortle-crf: Fast Technology Mapping for Lookup Table-Based FPGAsabstractA new technology mapping algorithm for lookup tablebased Field Programmable Gate Arrays (FPGA) is presented. The major innovation is a method for choosing gate-level decompositions based on bin packing. This approach is up to 28 times faster than a previous exhaustive approach. The algorithm also exploits reconvergent paths and replication of logic at fanout nodes to reduce the number of lookup tables in the circuit. The new algorithm is implemented in the Chortle-crf program. In an experimental comparison Chortle-crf requires 14 YO fewer lookup tables than Chortle [Fran90] and 10 ~o fewer lookup tables than mis-pga [Murg90a] to implement a set of benchmark networks. Chortle-crf can also implement a network as a circuit of Xilinx 3000 series Configurable Logic Blocks (CLBS). To implement the benchmark networks as circuits of CLBS Chortle-crf requires 12 70 fewer CLBS than mis-pga and 22 % fewer CLBS than XNFOPT [Xili89]. In these experiments Chortle-crf waa an average of 68 times faster than mis-pga and 30 times faster than XNFOPT. 1 Robert J. Francis, Jonathan Rose, Zvonko G. Vranesic |
DAC | 2 |
| 1991 | Will the Field-Programmable Gata Array Replace the Mask-Programmable Gate Array? (Panel Abstract)
Jonathan Rose |
DAC | 1 |
| 1991 | Technology Mapping on Lookup Table-Based FPGAs for PerformanceabstractA novel technology mapping algorithm that reduces the delay of combinational circuits implemented with lookup-table-based field-programmable gate arrays (FPGAs) is presented. The algorithm reduces the contribution of logic block delays to the critical path delay by reducing the number of lookup tables on the critical path. The key feature of the algorithm is the use of bin packing to determine the gate-level decomposition of every node in the network. In addition, reconvergent paths and the replication of logic at fanout nodes are exploited to further reduce the depth of the lookup table circuit. For fanout-free trees the algorithm will construct the optimal depth K-input table circuit when K is less than or equal to 6.> Robert J. Francis, Jonathan Rose, Zvonko G. Vranesic |
ICCAD | 2 |
| 1990 | Chortle: A Technology Mapping Program for Lookup Table-Based Field Programmable Gate ArraysabstractField Programmable Gate Arrays are new devices that combine the versatility of a Gate Array with the user-programmability of a PAL. This paper describes an algorithm for technology mapping of combinational logic into Field Programmable Gate Arrays that use lookup table memories to realize combinational functions. It is difficult to map into lookup tables using previous techniques because a single lookup table can perform a large number of logic functions, and prior approaches require each function to be instantiated separately in a library. The new algorithm, implemented in a program called Chortle uses the fact that a K-input lookup table can implement any Boolean function of K-inputs, and so does not require a library-based approach. Chortle takes advantage of this complete functionality to evaluate all possible decompositions of the input Boolean network nodes. It can determine the optimal (in area) mapping for fanout-free trees of combinational logic. In comparisons with the MIS II technology mapper, on MCNC-89 Logic Synthesis benchmarks Chortle achieves superior results in significantly less time. 1 Robert J. Francis, Jonathan Rose, Kevin Chung |
DAC | 2 |
| 1990 | A Detailed Router for Field-Programmable Gate ArraysabstractThe course graph expansion (CGE) detailed routing algorithm is presented for FPGAs (field-programmable gate arrays). The algorithm has the ability to resolve routing conflicts by considering the side-effects of one connection on another, and can be used over a wide range of FPGA interconnection architectures. CGE has been used to obtain excellent routing results for several industrial circuits with various FPGA routing architectures. The results show that CGE is able to route relatively large FPGAs in the absolute minimum number of tracks as determined by global routing, and that CGE has a linear run-time over circuit size.> Stephen Brown 0003, Jonathan Rose, Zvonko G. Vranesic |
ICCAD | 2 |
| 1990 | Parallel global routing for standard cellsabstractThe potential speedup of a standard cell global router using a general-purpose multiprocessor is investigated. LocusRoute, a global routing algorithm for standard cells, and its parallel implementation are presented. The uniprocessor speed and quality of LocusRoute is comparable to modern global routers. LocusRoute compares favorably with the TimberWolf 5.0 global router and a maze router that searches the same space more completely. Two successful methods of parallel decomposition of the router are presented. The first, in which multiple wires are routed in parallel, uses the notion of chaotic parallelism to achieve significant performance gains by relaxing data dependencies, at the cost of a minor loss in quality. Using iteration and careful assignment of wires to processors, this degradation is reduced. The approach achieves measured speedups from 5 to 14 using 15 processors. The second parallel decomposition technique is the evaluation of different routes for each wire on separate processors. It achieves speedups of up to 6 using 10 processors. It is demonstrated that when these two approaches are combined, the aggregate speedup is the product of the individual approaches' speedup, and, using an improved scheduling approach, it can be even greater. With a simple model based on these results, speedups of more than 75 using 150 processors are predicted.> Jonathan Rose |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1990 | Temperature measurement and equilibrium dynamics of simulated annealing placementsabstractOne way to alleviate the heavy computation required by simulated annealing placement algorithms is to replace a significant fraction of the higher or middle temperatures with a faster heuristic, and then follow it with simulated annealing. A crucial issue in this approach is the determination of the starting temperature for the simulated annealing phase-a temperature should be chosen that causes an appropriate amount of optimization to be done, but makes good use of the structure provided by the heuristic. A method for measuring the temperature of an existing placement is presented. The approach is based on the measurement of the probability distribution of the change in cost function, P( Delta C), and makes the assumption that the placement is in simulated annealing equilibrium at some temperature. The temperature of placements produced by both a simulated annealing and a min-cut placement algorithm are measured, and good agreement with known temperatures is obtained. The P( Delta C) distribution is also used to give an interesting view of the equilibrium dynamics of simulated annealing.> Jonathan Rose, Wolfgang Klebsch |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1988 | LocusRoute: A Parallel Global Router for Standard Cells
Jonathan Rose |
DAC | 1 |
| 1988 | Temperature measurement of simulated annealing placementsabstractOne way to reduce the computational requirements of simulated annealing placement algorithms is to use a faster heuristic to replace the early phase of simulated annealing. Such systems need to know the starting temperature for the annealing phase that makes the best use of the existing structure, yet provides an appropriate amount of improvement. A method for determining the temperature of an existing placement from an analysis of the probability distribution of the change in cost function is presented. Using this view, a novel definition of equilibrium is given and the equilibrium temperature of a placement is defined. Temperatures of placements produced both by a simulated annealing and a min-cut placement algorithm are measured.> Jonathan Rose, Wolfgang Klebsch |
ICCAD | 1 |
| 1988 | Parallel standard cell placement algorithms with quality equivalent to simulated annealingabstractAn algorithm called heuristic spanning creates parallelism by simultaneously investigating different areas of the plausible combinatorial search space. It is used to replace the high-temperature portion of simulated annealing. The low-temperature portion of simulated annealing is sped up by a technique called section annealing, in which placement is geographically divided and the pieces are assigned to separate processors. Each processor generates simulated-annealing-style moves for the cells in its area and communicates the moves to other processors as necessary. Heuristic spanning and section annealing are shown experimentally to converge to the same final cost function as regular simulated annealing. These approaches achieve significant speedup over uniprocessor simulated annealing, giving high-quality VLSI placement of standard cells in a short period of time.> Jonathan Rose, W. Martin Snelgrove, Zvonko G. Vranesic |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |