Steve Wilton

dblp:w/StevenJEWilton · also Steven J. E. Wilton · DBLP profile ↗
← Back
131ranked-venue papers
14as first author
21since 2021 · last 2025
0000-0002-1241-6690ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 128 · 14 first-author · 20 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2025 Chronbench: An Incremental HDL Benchmark Suite
abstract
FPGA CAD tools are often intended to compile an entire design from scratch to maximize the quality of results. Even with modern CAD algorithms this is a slow process. This paradigm limits designer productivity during incremental development, since every development iteration must endure the entire compilation process. Many vendor tools therefore offer ‘incremental modes’ that partially reuse compilation results to accelerate development at the HDL level of abstraction. Unfortunately, there is limited academic research into more sophisticated incremental HDL flows. We believe a key obstacle to research in this area is the lack of benchmarks which encapsulate realistic HDL development histories. As such we introduce Chronbench, a suite of HDL benchmarks which encapsulate development history as a chronological series of synthesizable commits in a git repository. In addition to five such benchmarks we present a tool for converting a public repository into a into a Chronbench benchmark. Further, we synthesize, place, and route 170 commits in order to fully characterize the suite. Finally, we analyze the characterization data to produce some key insights about the relative magnitude of HDL development changes and observe that approximately half of real development commits do not significantly impact device utilization, indicating significant potential for reuse during HDL development.
Zakary Nafziger, Steve Wilton
FCCM2
2025 Open-Source FPGA Routing Runtime Prediction for Improved Productivity Via Smart Route Termination
abstract
Field-Programmable Gate Array (FPGA) routing is computationally expensive, taking hours or days with no guarantee of success. While prior work has used machine learning (ML) to guide placement and routing or predict routing outcomes, the process remains challenging to model precisely. A recent work has proposed using ML to predict the number of iterations remaining in a negotiated congestion router while it runs, enabling early termination of routing runs unlikely to succeed. However, that approach has key limitations hindering its utility: (1) iteration count is poorly correlated with runtime, (2) it ignores prediction confidence when deciding whether to exit, and (3) it cannot assess whether extending a routing run past a predefined limit is worthwhile. This paper presents a new ML-based framework that addresses these limitations. We introduce a method for estimating router workload based on node traversals in the FPGA routing resource graph, which strongly correlates with runtime and enables more accurate early exit decisions. We also propose a tunable success-confidence threshold that allows users to trade off runtime against success rate and we design a ML mixture of experts architecture to enable this thresholding effectively. Finally, we show how our architecture can “look ahead” to determine whether a routing run is likely to succeed if allowed to delay termination and continue past its initial time limit. We implement our approach on top of the negotiated congestion routing algorithm and, in our experiments on very difficult-toroute circuits, we find that the number of circuits successfully routed within a fixed cumulative routing time budget increases by 215 % compared with the approach from prior work.
Andrew David Gunter, Steve Wilton
FPL2
2025 Versatile Place and Route with Continuous Routing Runtime Prediction and Smart Route Termination
abstract
Field-Programmable Gate Array (FPGA) routing is computationally expensive, taking hours or days with no guarantee of success. While prior work has used machine learning (ML) to guide placement and routing or predict routing outcomes, the process remains challenging to model precisely. This demo presents a new machine learning-based tool that addresses this problem.
Andrew David Gunter, Steve Wilton
FPL2
2025 Using Data to Reduce Uncertainty in FPGA Routing
abstract
The prefabricated resources in a field-programmable gate array (FPGA) can pose challenges resulting in unexpectedly long routing times. This sometimes leads to FPGA engineers prematurely terminating a viable routing run because they mistakenly believe the lengthy runtime indicates an unroutable design. In other cases, the FPGA engineer may wait long periods of time on a run that is doomed to never converge to a solution. As this leads to time wasted on runs which never reach design closure in either case, it is ideal to instead have an ML model decide whether or not to terminate routing. In this work, we introduce data-driven machine learning (ML) techniques for predicting both FPGA design routability and routing runtime continuously during routing.
Andrew David Gunter, Steve Wilton
FPL2
2025 QUTE: Quantifying Uncertainty in TinyML models with Early-exit-assisted ensembles for model-monitoring
abstract
Uncertainty quantification (UQ) provides a resource-efficient solution for on-device monitoring of tinyML models deployed remotely without access to true labels. However, existing UQ methods impose significant memory and compute demands, making them impractical for ultra-low-power, KB-sized tinyML devices. Prior work has attempted to reduce overhead by using early-exit ensembles to quantify uncertainty in a single forward pass, but these approaches still carry prohibitive costs. To address this, we propose QUTE, a novel resource-efficient early-exit-assisted ensemble architecture optimized for tinyML models. QUTE introduces additional output blocks at the final exit of the base network, distilling early-exit knowledge into these blocks to form a diverse yet lightweight ensemble. We show that QUTE delivers superior uncertainty quality on tiny models, achieving comparable performance on larger models with 59% smaller model sizes than the closest prior work. When deployed on a microcontroller, QUTE demonstrates a 31% reduction in latency on average. In addition, we show that QUTE excels at detecting accuracy-drop events, outperforming all prior works.
Nikhil Ghanathe, Steve Wilton
ICML2
2024 A Semi Black-Box Adversarial Bit- Flip Attack with Limited DNN Model Information
abstract
Despite the rising prevalence of deep neural networks (DNNs) in cyber-physical systems, their vulnerability to adversarial bit-flip attacks (BFAs) is a noteworthy concern. This paper proposes B3FA, a semi-black-box BFA-based parameter attack on DNNs, assuming the adversary has limited knowledge about the model. We consider practical scenarios often feature a more restricted threat model for real-world systems, contrasting with the typical BFA models that presuppose the adversary's full access to a network's inputs and parameters. The introduced bit-flip approach utilizes a magnitude-based ranking method and a statistical reconstruction technique to identify the vulnerable bits. We demonstrate the effectiveness of B3FA on several DNN models in a semi-black-box setting. For example, B3FA could drop the accuracy of a MobileNetV2 from 69.84% to 9% with only 20 bit-flips in a real-world setting.
Behnam Ghavami, Mani Sadati, Mohammad Shahidzadeh, Lesley Shannon, Steve Wilton
ICCD5
2024 ZOBNN: Zero-Overhead Dependable Design of Binary Neural Networks with Deliberately Quantized Parameters
abstract
Low-precision weights and activations in deep neural networks (DNNs) outperform their full-precision counterparts in terms of hardware efficiency. When implemented with low-precision operations, specifically in the extreme case where network parameters are binarized (i.e. BNNs), the two most frequently mentioned benefits of quantization are reduced memory consumption and a faster inference process. In this paper, we introduce a third advantage of very low-precision neural networks: improved fault-tolerance attribute. We investigate the impact of memory faults on state-of-the-art binary neural networks (BNNs) through comprehensive analysis. Despite the inclusion of floating-point parameters in BNN architectures to improve accuracy, our findings reveal that BNNs are highly sensitive to deviations in these parameters caused by memory faults. In light of this crucial finding, we propose a technique to improve BNN dependability by restricting the range of float parameters through a novel deliberately uniform quantization. The introduced quantization technique results in a reduction in the proportion of floating-point parameters utilized in the BNN, without incurring any additional computational overheads during the inference stage. The extensive experimental fault simulation on the proposed BNN architecture (i.e. ZOBNN) reveal a remarkable 5X enhancement in robustness compared to conventional floating-point DNN. Notably, this improvement is achieved without incurring any computational overhead. Crucially, this enhancement comes without computational overhead. ZOBNN excels in critical edge applications characterized by limited computational resources, prioritizing both dependability and real-time performance.
Behnam Ghavami, Mohammad Shahidzadeh, Lesley Shannon, Steve Wilton
IOLTS4
2024 Designing an IEEE-Compliant FPU that Supports Configurable Precision for Soft Processors
abstract
Field Programmable Gate Arrays (FPGAs) are commonly used to accelerate floating-point (FP) applications. Although researchers have extensively studied FPGA FP implementations, existing work has largely focused on standalone operators and frequency-optimized designs. These works are not suitable for FPGA soft processors which are more sensitive to latency, impose a lower frequency ceiling, and require IEEE FP standard compliance. We present an open-source floating-point unit (FPU) for FPGA RISC-V soft processors that is fully IEEE compliant with configurable levels of FP precision. Our design emphasizes runtime performance with 25% lower latency in the most common instructions compared to previous works while maintaining efficient resource utilization. Our FPU also allows users to explore various mantissa widths without having to rewrite or recompile their algorithms. We use this to investigate the scalability of our reduced-precision FPU across numerous microbenchmark functions as well as more complex case studies. Our experiments show that applications like the discrete cosine transformation and the Black-Scholes model can realize a speedup of more than 1.35x in conjunction with a 43% and 35% reduction in lookup table and flip-flop resources while experiencing less than a 0.025% average loss in numerical accuracy with a 16-bit mantissa width.
Chris Keilbart, Yuhui Gao, Martin Chua, Eric Matthews, Steve Wilton, Lesley Shannon
ACM Trans. Reconfigurable Technol. Syst.5
2023 T-RecX: Tiny-Resource Efficient Convolutional neural networks with early-eXit
abstract
Deploying Machine learning (ML) on milliwatt-scale edge devices (tinyML) is gaining popularity due to recent breakthroughs in ML and Internet of Things (IoT). Most tinyML research focuses on model compression techniques that trade accuracy (and model capacity) for compact models to fit into the KB-sized tiny-edge devices. In this paper, we show how such models can be enhanced by the addition of an early exit intermediate classifier. If the intermediate classifier exhibits sufficient confidence in its prediction, the network exits early thereby, resulting in considerable savings in time. Although early exit classifiers have been proposed in previous work, these previous proposals focus on large networks, making their techniques suboptimal/impractical for tinyML applications. Our technique is optimized specifically for tiny-CNN sized models. In addition, we present a method to alleviate the effect of network overthinking by leveraging the representations learned by the early exit. We evaluate T-RecX on three CNNs from the MLPerf tiny benchmark suite for image classification, keyword spotting and visual wake word detection tasks. Our results show that T-RecX 1) improves the accuracy of baseline network, 2) achieves 31.58% average reduction in FLOPS in exchange for one percent accuracy across all evaluated models. Furthermore, we show that our methods consistently outperform popular prior works on the tiny-CNNs we evaluate.
Nikhil Ghanathe, Steve Wilton
CF2
2023 A Machine Learning Approach for Predicting the Difficulty of FPGA Routing Problems
abstract
In this paper, we present a Machine Learning (ML) Mixture of Experts (MoE) technique to predict the number of iterations needed for a Pathfinder-based FPGA router to complete a routing problem. Given a placed circuit, our technique uses features gathered on each routing iteration to predict if the circuit is routable and how many more iterations will be required to successfully route the circuit. This enables early exit for routing problems which are unlikely to be completed in a target number of iterations. Such early exit may help to achieve a successful route within tractable time by allowing the user to quickly retry the circuit compilation with a different random seed, a modified circuit design, or a different FPGA. We demonstrate our predictor in the VTR 8 framework; compared to VTR's predictor, our ML predictor incurs lower prediction errors on the Koios Deep Learning and Titan23 benchmark suites. Based on our tests, equipping VTR with our ML predictor would reduce time wasted on unroutable designs by 31% while also allowing 28% more routable designs to be completed.
Andrew David Gunter, Steve Wilton
FCCM2
2023 Reformulating the FPGA Routability Prediction Problem with Machine Learning
abstract
Field-Programmable Custom Compute technology is now commonplace in important commercial settings. This has primarily been driven by improvements in Field-Programmable Gate Array (FPGA) technology. Commercial use cases typically demand the implementation of complex designs on large FPGAs, increasing typical FPGA design compile times. In particular, the routing compilation step can take days to complete. Amplifying the negative impact of these long compilations is that FPGA users have no guarantee that compilation will be successful.
Andrew David Gunter, Steve Wilton
FCCM2
2023 Designing a configurable IEEE-compliant FPU that supports variable precision for soft processors
abstract
FPGAs are an increasingly popular medium for many high-performance data center workloads and the rapidly-expanding artificial intelligence domain. These applications often make extensive use of floating-point (FP) numbers defined by the IEEE 754 standard [1]. Although researchers have extensively studied FPGA-based hardware FP implementations, existing work has largely focused on standalone and throughput-optimized data-path designs. Such designs optimize performance by increasing throughput with long pipelines and high frequencies. This approach is not suitable for soft processors, which are more sensitive to latency in order to reduce stalls due to data hazards. Additionally, the frequency ceiling imposed by other internal components of the soft processor necessarily limits the maximum operating frequency of the Floating-Point Unit (FPU).
Chris Keilbart, Yuhui Gao, Martin Chua, Eric Matthews, Steve Wilton, Lesley Shannon
FCCM5
2023 Towards a Machine Learning Approach to Predicting the Difficulty of FPGA Routing Problems
abstract
In this poster, we present a Machine Learning (ML) technique to predict the number of iterations needed for a Pathfinder-based FPGA router to complete a routing problem. Given a placed circuit, our technique uses features gathered on each routing iteration to predict if the circuit is routable and how many more iterations will be required to successfully route the circuit. This enables early exit for routing problems which are unlikely to be completed in a target number of iterations. Such early exit may help to achieve a successful route within tractable time by allowing the user to quickly retry the circuit compilation with a different random seed, a modified circuit design, or a different FPGA. We demonstrate our predictor in the VTR 8 framework; compared to VTR's predictor, our ML predictor incurs lower prediction errors on the Koios Deep Learning benchmark suite. This corresponds with an approximate time saving of 48% from early rejection of unroutable FPGA designs while also successfully completing 5% more routable designs and having a 93% shorter early exit latency.
Andrew David Gunter, Steve Wilton
FPGA2
2022 Boosting Domain-Specific Debug Through Inter-frame Compression
abstract
Acceleration of machine learning models is proving to be an important application for FPGAs. Unfortunately, debugging such models during training or inference is difficult. Software simulations of a machine learning system may be of insufficient detail to provide meaningful debug insight, or may require infeasibly long run-times. Thus, it is often desirable to debug the accelerated model while it is running on real hardware. Effective on-chip debug often requires instrumenting a design with additional circuitry to store run-time data, consuming valuable chip resources. Previous work has developed methods to perform lossy compression of signals by exploiting machine learning specific knowledge, thereby increasing the amount of debug context that can be stored in an on-chip trace buffer. However, all prior work compresses each successive element in a signal of interest independently. Since debug signals may have temporal similarity in many machine learning applications there is an opportunity to further increase trace buffer utilization. In this paper, we present an architecture to perform lossless temporal compression in addition to the existing lossy element-wise compression. We show that, when applied to a typical machine learning algorithm in realistic debug scenarios, we are able to store twice as much information in an on-chip buffer while increasing the total area of the debug instrument by approximately 25%. The impact is that, for a given instrumentation budget, a significantly larger trace window is available during debug, possibly allowing a designer to narrow down the root cause of a bug faster.
Zakary Nafziger, Martin Chua, Daniel H. Noronha, Steve Wilton
FPT4
2022 Positive-Phase Temperature Scaling for Quantum-Assisted Boltzmann Machine Training
abstract
Quantum-assisted sampling is a promising technique to enable training probabilistic ML models, which otherwise depend on slow-mixing classical sampling methods; such as, the use of Quantum Annealing Processors (QAP) to train Boltzmann Machines (BMs). Previous work has shown that QAPs can sample from a Boltzmann distribution, although, at an unknown instance-dependent temperature. Due to this distribution divergence, existing training algorithms have resorted to negative-phase temperature scaling. This method, although effective under arduous tuning, introduces unwanted noise to the sampleset due to the quantization errors caused by the underutilization of the QAP bias ranges; and is prone to bias overflow. We introduce a change in the training algorithm to allow positive-phase temperature scaling; an approach that reduces the impact of quantization noise, while still incorporating temperature scaling. As a result, we see an overall improvement in the convergence rate and testing accuracy, when compared to the state-of-the-art approach.
Jose P. Pinilla, Steve Wilton
SC2
2022 Adaptive Clock Management of HLS-generated Circuits on FPGAs
abstract
In this article, we present Syncopation , a performance-boosting fine-grained timing analysis and adaptive clock management technique for High-Level Synthesis-generated circuits implemented on Field-Programmable Gate Arrays. The key idea is to use the HLS scheduling information along with the placement and routing results to determine the worst-case timing path for individual clock cycles. By adjusting the clock period on a cycle-by-cycle basis, we can increase performance of an HLS-generated circuit. Our experiments show that Syncopation improves performance by 3.2% (geomean) across all benchmarks (up to 47%). In addition, by employing targeted synthesis techniques along with Syncopation, we can achieve 10.3% performance improvement (geomean) across all benchmarks (up to 50%). Syncopation instrumentation is implemented entirely in soft logic without requiring alterations to the HLS-synthesis toolchain or changes to the FPGA, and has been validated on real hardware.
Kahlan Gibson, Esther Roorda, Daniel H. Noronha, Steve Wilton
ACM Trans. Reconfigurable Technol. Syst.4
2022 Rethinking Embedded Blocks for Machine Learning Applications
abstract
The underlying goal of FPGA architecture research is to devise flexible substrates that implement a wide variety of circuits efficiently. Contemporary FPGA architectures have been optimized to support networking, signal processing, and image processing applications through high-precision digital signal processing (DSP) blocks. The recent emergence of machine learning has created a new set of demands characterized by: (1) higher computational density and (2) low precision arithmetic requirements. With the goal of exploring this new design space in a methodical manner, we first propose a problem formulation involving computing nested loops over multiply-accumulate (MAC) operations, which covers many basic linear algebra primitives and standard deep neural network (DNN) kernels. A quantitative methodology for deriving efficient coarse-grained compute block architectures from benchmarks is then proposed together with a family of new embedded blocks, called MLBlocks. An MLBlock instance includes several multiply-accumulate units connected via a flexible routing, where each configuration performs a few parallel dot-products in a systolic array fashion. This architecture is parameterized with support for different data movements, reuse, and precisions, utilizing a columnar arrangement that is compatible with existing FPGA architectures. On synthetic benchmarks, we demonstrate that for 8-bit arithmetic, MLBlocks offer 6× improved performance over the commercial Xilinx DSP48E2 architecture with smaller area and delay; and for time-multiplexed 16-bit arithmetic, achieves 2× higher performance per area with the same area and frequency. All source codes and data, along with documents to reproduce all the results in this article, are available at http://github.com/raminrasoulinezhad/MLBlocks .
Seyedramin Rasoulinezhad, Esther Roorda, Steve Wilton, Philip H. W. Leong, David Boland
ACM Trans. Reconfigurable Technol. Syst.3
2022 FPGA Architecture Exploration for DNN Acceleration
abstract
Recent years have seen an explosion of machine learning applications implemented on Field-Programmable Gate Arrays (FPGAs) . FPGA vendors and researchers have responded by updating their fabrics to more efficiently implement machine learning accelerators, including innovations such as enhanced Digital Signal Processing (DSP) blocks and hardened systolic arrays. Evaluating architectural proposals is difficult, however, due to the lack of publicly available benchmark circuits. This paper addresses this problem by presenting an open-source benchmark circuit generator that creates realistic DNN-oriented circuits for use in FPGA architecture studies. Unlike previous generators, which create circuits that are agnostic of the underlying FPGA, our circuits explicitly instantiate embedded blocks, allowing for meaningful comparison of recent architectural proposals without the need for a complete inference computer-aided design (CAD) flow. Our circuits are compatible with the VTR CAD suite, allowing for architecture studies that investigate routing congestion and other low-level architectural implications. In addition to addressing the lack of machine learning benchmark circuits, the architecture exploration flow that we propose allows for a more comprehensive evaluation of FPGA architectures than traditional static benchmark suites. We demonstrate this through three case studies which illustrate how realistic benchmark circuits can be generated to target different heterogeneous FPGAs.
Esther Roorda, Seyedramin Rasoulinezhad, Philip H. W. Leong, Steve Wilton
ACM Trans. Reconfigurable Technol. Syst.4
2021 Flexible Instrumentation for Live On-Chip Debug of Machine Learning Training on FPGAs
abstract
FPGAs have recently shown promise for accelerating machine learning training. This has led to research into the co-design of narrow-precision accelerator architectures and the investigation of novel machine learning models. Such research can be extremely expensive, as the steep cost of training a model can increase several-fold due to the need of performing hyper-parameter tuning and adjustments to the model to ensure acceptable convergence speed and accuracy. In this scenario, monitoring key data on-chip is essential to more quickly understand and diagnose problems, significantly reducing training costs.Previous work has proposed on-chip debug instrumentation to monitor key signals for both general-purpose circuits and inference algorithms. This instrumentation either performs limited on-chip compression, or is extremely restricted in the amount of run-time customization that may occur. We argue that for training applications, the extremely long and expensive training runs warrant significantly more flexibility in the on-chip instrumentation, even at the expense of some chip area.In this paper, we propose flexible debug instrumentation that allows for the live debugging of machine learning systems during training. Different from previous debug instrumentation, our instrumentation offers firmware programmability, allowing the researcher to gather data in a large variety of ways that would likely not be anticipated at compile time.
Daniel H. Noronha, Zhiqiang Que, Wayne Luk, Steve Wilton
FCCM4
2021 MAFIA: Machine Learning Acceleration on FPGAs for IoT Applications
abstract
Recent breakthroughs in ML have produced new classes of models that allow ML inference to run directly on milliwatt-powered IoT devices. On one hand, existing ML-to-FPGA compilers are designed for deep neural-networks on large FPGAs. On the other hand, general-purpose HLS tools fail to exploit properties specific to ML inference, thereby resulting in suboptimal performance. We propose MAFIA, a tool to compile ML inference on small form-factor FPGAs for IoT applications. MAFIA provides native support for linear algebra operations and can express a variety of ML algorithms, including state-of-the-art models. We show that MAFIA-generated programs outperform best-performing variant of a commercial HLS compiler by 2.5 × on average.
Nikhil Ghanathe, Vivek Seshadri, Rahul Sharma 0001, Steve Wilton, Aayan Kumar
FPL4
2021 In-circuit tuning of deep learning designs
Zhiqiang Que, Daniel H. Noronha, Ruizhe Zhao, Xinyu Niu, Steve Wilton, Wayne Luk
J. Syst. Archit.5
2020 Syncopation: Adaptive Clock Management for High-Level Synthesis Generated Circuits on FPGAs
abstract
High-level synthesis (HLS) tools improve hardware designer productivity by enabling software design techniques during hardware development. During HLS the delay of paths can only be estimated, so the resulting circuit may suffer from unbalanced computational path delays across clock cycles. Since the maximum operating frequency of circuits is determined statically using the worst-case timing path, unbalanced paths may lead to reduced performance compared to circuits designed at the hardware level. In this paper, we address this using Syncopation, a performance-boosting fine-grained timing analysis and adaptive clock management technique for HLS circuits. The key idea is to use the HLS scheduling information along with the results from placement and routing to determine the worst-case timing path for individual clock cycles. By then adjusting the clock period on a cycle-to-cycle basis, we can increase circuit performance. Our experiments show that Syncopation and fine-grained timing analysis can improve performance without altering the HLS-synthesis toolchain.
Kahlan Gibson, Esther Roorda, Daniel H. Noronha, Steve Wilton
FPL4
2020 Fast Turnaround HLS Debugging Using Dependency Analysis and Debug Overlays
abstract
High-level synthesis (HLS) has gained considerable traction over recent years, as it allows for faster development and verification of hardware accelerators than traditional RTL design. While HLS allows for most bugs to be caught during software verification, certain non-deterministic or data-dependent bugs still require debugging the actual hardware system during execution. Recent work has focused on techniques to allow designers to perform in-system debug of HLS circuits in the context of the original software code; however, like RTL debug, the user must still determine the root cause of a bug using small execution traces, with lengthy debug turns. In this work, we demonstrate techniques aimed at reducing the time HLS designers spend performing in-system debug. Our approaches consist of performing data dependency analysis to guide the user in selecting which variables are observed by the debug instrumentation, as well as an associated debug overlay that allows for rapid reconfiguration of the debug logic, enabling rapid switching of variable observation between debug iterations. In addition, our overlay provides additional debug capability, such as selective function tracing and conditional buffer freeze points. We explore the area overhead of these different overlay features, showing a basic overlay with only a 1.7% increase in area overhead from the baseline debug instrumentation, while a deluxe variant offers 2×--7× improvement in trace buffer memory utilization with conditional buffer freeze support.
Al-Shahna Jamal, Eli Cahill, Jeffrey B. Goeders, Steve Wilton
ACM Trans. Reconfigurable Technol. Syst.4
2019 On-chip FPGA Debug Instrumentation for Machine Learning Applications
abstract
FPGAs provide a promising implementation option for many machine learning applications. Although simulations or software models can be used to explore the design space of these applications, often the final behaviour can not be evaluated until the design is mapped to the FPGA and integrated into the target system. This may be because long run-times are required, or because the environment can not be adequately described using a software model. Once unexpected behaviour is observed, on-chip debug is notoriously difficult; typically a design is instrumented with on-chip trace buffers that record the run-time behaviour for later interrogation. In this paper, we describe instrumentation that can accelerate the process of debugging machine learning applications implemented on an FPGA. Unlike previous work, our instrumentation is optimized to take advantage of characteristics of this application domain. Our instruments gather useful domain-specific information about the observed variables instead of recording the raw values of those elements. Results show that the proposed instruments provide at least 17.8x longer visibility in the most conservative of our experiments at a low area and latency cost.
Daniel H. Noronha, Ruizhe Zhao, Jeffrey B. Goeders, Wayne Luk, Steve Wilton
FPGA5
2019 Towards In-Circuit Tuning of Deep Learning Designs
abstract
This paper presents InTune, a novel approach for in-circuit tuning of deep learning designs targeting implementations in field-programmable gate array technology. This approach combines two promising techniques: domain-specific adaptation and in-circuit tuning. Domain-specific adaptation exploits domain-specific information in adapting pre-trained models to specific application domains, replacing standard convolution layers with efficient convolution blocks; the effects of such adaptation are then assessed by in-circuit tuning instruments to provide information to application builders for tuning the design. This approach is illustrated by its deployment in tuning deep neural networks, and its potential for a new generation of domain-specific tools with tight integration of synthesis and in-circuit tuning is explored.
Zhiqiang Que, Daniel H. Noronha, Ruizhe Zhao, Steve Wilton, Wayne Luk
ICCAD4
2018 Extending post-silicon coverage measurement using time-multiplexed FPGA overlays
abstract
Test coverage has emerged as an essential metric for evaluating the effectiveness of both pre-silicon verification and post-silicon validation. Evaluating coverage post-silicon is difficult due to the lack of visibility into the internal operation of integrated circuits. Adding coverage monitors to a design consumes a significant amount of chip area. Field-Programmable Gate Arrays (FPGAs) are commonly deployed as a rapid prototyping platform to accelerate the validation of digital designs, and though they share same visibility challenges as post-silicon, recent work have proposed the use of overlays to improve debug effectiveness. In this paper, we describe how this emerging debug technology can be re-purposed to also implement coverage monitors in a time-multiplexed fashion to evaluate coverage at post-silicon.
Fatemeh Eslami, Eddie Hung, Steve Wilton
ETS3
2018 Architecture Exploration for HLS-Oriented FPGA Debug Overlays
abstract
High-Level Synthesis (HLS) promises improved designer productivity, but requires a debug ecosystem that allows designers to debug in the context of the original source code. Recent work has presented in-system debug frameworks where instrumentation added to the design collects trace data as the circuit runs, and a software tool that allows the user to replay the execution using the captured data. When searching for the root cause of a bug, the designer may need to modify the instrumentation to collect data from a new part of the design, requiring a lengthy recompile.
Al-Shahna Jamal, Jeffrey B. Goeders, Steve Wilton
FPGA3
2018 An FPGA Overlay Architecture Supporting Rapid Implementation of Functional Changes during On-Chip Debug
abstract
As Field-Programmable Gate Arrays become more complex, debugging designs implemented on these devices has become increasingly time-consuming. For many types of bugs, simulation is not sufficient, and the only way to uncover the root cause of unexpected behaviour is to run the design in hardware at speed. Many techniques that support on-chip debug have been described; typically, these techniques involve instrumenting the design to increase observability. In this paper, we describe instrumentation that not only increases observability, but that can also be used to control certain aspects of the design. Supported functional changes include applying small deviations in the control flow of the circuit, or the ability to override signal assignments to perform efficient "what if'" tests. Our approach uses a novel overlay architecture which allows these changes to be implemented during debug without recompiling the design. Changes can be made in seconds, dramatically reducing the time to perform a debug iteration. Our overlay is specifically optimized for designs created using a high-level synthesis (HLS) flow; by taking advantage of information from the HLS tool, the overhead of the overlay can be kept low.
Al-Shahna Jamal, Jeffrey B. Goeders, Steve Wilton
FPL3
2018 Simultaneous Inference and Training Using On-FPGA Weight Perturbation Techniques
abstract
We present an FPGA-optimized implementation of online neural network training based on weight perturbation (WP) techniques. When compared to the classic backpropagation (BP) algorithm, WP is capable of delivering competitive performance while occupying minimal area resources. Perturbation-based methods have been demonstrated as viable training techniques and are suitable for on-line learning applications which adapt to changing conditions. The viability of applying WP-based on-chip training for low-precision fixed-point hardware is demonstrated on two distinct MLP benchmarks: the Iris dataset classification network and an RF anomaly detector. When synthesized to a Xilinx Kintex-7 XC7K410T FPGA, WP offers a 3-10x area savings with <;1% degradation in accuracy compared with backpropagation. Compared with an inference-only implementation the overhead of introducing on-chip learning is approximately 30%.
Siddhartha 0003, Steve Wilton, David Boland, Barry Flower, Perry Blackmore, Philip H. W. Leong
FPT2
2018 LeFlow: Automatic Compilation of TensorFlow Machine Learning Applications to FPGAs
abstract
Acceleration of Machine Learning applications on Field-Programmable Gate Arrays (FPGAs) has shown to have advantages over other computing platforms in recent work. However, since machine learning code is often specified in a high-level software language such as Python, the manual translation of the algorithm to either C code for high-level synthesis or to Register Transfer Level (RTL) code for synthesis is time consuming and requires the designer to have expertise in designing hardware. In order to show how we can make FPGAs more accessible to software developers, we present a demonstration of LeFlow: an open-source tool which maps numerical computation models written in TensorFlow to synthesizable RTL. This demonstration includes two examples which begin with a model written in TensorFlow and show how a designer would use the LeFlow tool to generate Verilog, simulate the result, and synthesize the design to target FPGAs.
Daniel H. Noronha, Kahlan Gibson, Bahar Salehpour, Steve Wilton
FPT4
2018 Rapid Triggering Capability Using an Adaptive Overlay during FPGA Debug
abstract
Field Programmable Gate Array (FPGA) technology is rapidly gaining traction in a wide range of applications. Nonetheless, FPGAs still require long design and debug cycles. To debug hardware circuits, trace-based instrumentation is inserted into the design that enables capturing data during the circuit execution into on-chip memories for later offline analysis. Since on-chip memories are limited, a trigger circuitry is used to only record data related to specific events during the execution. However, during debugging, a circuit recompilation is required on modifying these instruments. This can be very slow, reducing debug productivity. In this article, we propose a non-intrusive and rapid triggering solution with a tailored overlay fabric and mapping algorithm that seeks to enable fast debug iterations without performing a recompilation. This overlay is specialized for small combinational and sequential circuits with a single output; such circuits are typical of common trigger functions. We present an adaptive strategy to construct the overlay fabric using spare FPGA resources at compile time. At debug time, our proposed trigger mapping algorithms adapt to this specialized overlay to rapidly implement combinational and sequential trigger circuits. Our results show that the overlay fabric can be reconfigured to map different triggering scenarios in less than 40s instead of recompiling the circuit during debug iterations, increasing debug productivity.
Fatemeh Eslami, Steve Wilton
ACM Trans. Design Autom. Electr. Syst.2
2018 Introduction to the Special Section on Deep Learning in FPGAs
abstract
International audience
Deming Chen, Andrew Putnam, Steve Wilton
ACM Trans. Reconfigurable Technol. Syst.3
2017 Accelerating in-system FPGA debug of high-level synthesis circuits using incremental compilation techniques
abstract
High-Level Synthesis has emerged as a promising technology for improving FPGA designer productivity, but will only be successful if it is accompanied by a debug ecosystem. Recent efforts have presented in-system debug techniques which allow a designer to debug an implementation, running on an FPGA, in the context of the original source code. These techniques typically store a history of all user variables on chip. To maximize the effectiveness of the on-chip memory, it is desirable to store only selected user variables. Unfortunately, this may lead to multiple debug runs. In existing frameworks, changing the variables to be stored between runs requires a full recompile. In this paper, we propose several flows that use incremental compilation to reduce the debug turn-around time. The first flow, in which the user circuit and instrumentation are co-optimized during compilation, gives the fastest debug clock speeds but suffers in user circuit performance once the debug instrumentation is removed. In the second flow, the optimization of the user circuit is sacrosanct. It is placed and routed first without having any constraints and the debug instrumentation is added later leading to the fastest user circuit clock speeds, but performance suffers slightly during debug. Using either flow, we achieve 40% reduction in debug turn-around times, on average.
Pavan Kumar Bussa, Jeffrey B. Goeders, Steve Wilton
FPL3
2017 Signal-Tracing Techniques for In-System FPGA Debugging of High-Level Synthesis Circuits
abstract
High-level synthesis (HLS) promises to increase designer productivity in the face of increasing field-programmable gate array sizes, and broaden the market of use, allowing software designers to reap the benefits of hardware implementation. One roadblock to HLS adoption is the lack of an in-system debugging infrastructure. Although designers can run their software code on a workstation, or simulate the register-transfer level, neither can reliably capture the behaviors, and therefore bugs, that may be present in the final system. Debugging hardware circuits in-system requires using signal-tracing to record circuit behavior for later offline analysis. In this paper, we present a debugging architecture, which automatically records key hardware signals, and relates them back to the original software source code. This architecture allows designers to debug HLS circuits in-system, in the context of the original source code. We present several signal-tracing techniques, tailored to HLS circuits, which allow a much longer execution trace to be captured. These techniques include signal compression, dynamically changing which signals are recorded cycle-by-cycle, and offline signal restoration. Compared to using an embedded logic analyzer to perform signal-tracing, our architecture increases the length of execution trace that can be recorded by 127X. For each 100 Kb of trace buffer memory, our architecture can record 15 369 executed lines of C code.
Jeffrey B. Goeders, Steve Wilton
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2016 Quantifying observability for in-system debug of high-level synthesis circuits
abstract
In recent years high-level synthesis (HLS) has seen considerable attention as it promises to increase designer productivity and make custom hardware implementation accessible to software developers. A challenge facing those developing HLS technologies is how to allow users to understand, debug and optimize their final hardware systems. Recently, several techniques have been developed to provide in-system debugging capabilities for HLS circuits. These techniques instrument the user's design with some debugging circuitry to provide observability into the circuit during execution. Due to resource constraints, it is usually infeasible to view all variable values for the entire circuit execution. Rather, instrumentation usually captures only some variable values and for only a portion of the circuit execution. In this paper we present a metric for measuring the observability into an executing HLS circuit. This metric reflects the portion of variable accesses that are available to the user, the duration of execution for which these values are available, as well as accommodating variations in importance between source code variables. This metric can be used to understand how different circuit observation networks can provide the user with different levels of observability into the HLS circuit execution. As a demonstration of the applicability of the metric, we first study differences between recent debugging approaches for HLS circuits, and quantify the level of observability provided by such architectures. We then explore different schemes to select which variables are accessible in the observation network, and measure impact on variable availability and length of captured execution trace.
Jeffrey B. Goeders, Steve Wilton
FPL2
2016 Enhanced source-level instrumentation for FPGA in-system debug of High-Level Synthesis designs
abstract
High-Level Synthesis (HLS) has emerged as a leading technology to reduce the design time and complexity that is associated with reconfigurable systems. In order to maintain the productivity promised by HLS, it is important that the designer can debug the system in the context of the high-level code. Currently, software simulations offer a quick and familiar method to target logic and syntax bugs, while software/hardware co-simulations are useful for synthesis verification. However, to analyze the behaviour of the circuit as it is running, the user is forced to understand waveforms from the synthesized design.
Jose P. Pinilla, Steve Wilton
FPT2
2016 An FPGA Architecture and CAD Flow Supporting Dynamically Controlled Power Gating
abstract
Leakage power is an important component of the total power consumption in field-programmable gate arrays (FPGAs) built using 90-nm and smaller technology nodes. Power gating was shown to be effective at reducing the leakage power. Previous techniques focus on turning OFF unused FPGA resources at configuration time; the benefit of this approach depends on resource utilization. In this paper, we present an FPGA architecture that enables dynamically controlled power gating, in which FPGA resources can be selectively powered down at run-time. This could lead to significant overall energy savings for applications having modules with long idle times. We also present a CAD flow that can be used to map applications to the proposed architecture. We study the area and power tradeoffs by varying the different FPGA architecture parameters and power gating granularity. The proposed CAD flow is used to map a set of benchmark circuits that have multiple power-gated modules to the proposed architecture. Power savings of up to 83% are achievable for these circuits. Finally, we study a control system of a robot that is used in endoscopy. Using the proposed architecture combined with clock gating results in up to 19% energy savings in this application.
Assem A. M. Bsoul, Steve Wilton, Kuen Hung Tsoi, Wayne Luk
IEEE Trans. Very Large Scale Integr. Syst.2
2015 Using Dynamic Signal-Tracing to Debug Compiler-Optimized HLS Circuits on FPGAs
abstract
High-level synthesis (HLS) for FPGA designs has received considerable attention in recent years. To make this design methodology mainstream, improved debugging technologies are essential. Ideally, a user should be able to debug their design using the original source code, without detailed knowledge of the underlying hardware, while the circuit executes in-situ. Although recent work has made progress toward this goal, existing solutions are unable to provide visibility into circuits that have been heavily optimized by the compiler. HLS compilers typically perform many optimizations, including moving variable values out of memories and into registers distributed throughout the design. Debugging such circuits typically requires either understanding the hardware and probing the appropriate RTL level registers, or ignoring these variables while debugging the design, neither of which is desirable. In this work we present a new signal-tracing technique, specifically designed for circuits that have been optimized by an HLS tool. Information is extracted from the HLS process to determine which signals are relevant to record each cycle. We automatically embed circuitry which dynamically selects the relevant signals, cycle-by-cycle, and records them into on-chip memories. In addition, we explore techniques to balance tracing between cycles to further improve memory efficiency. For each 100Kb of memory allocated to trace buffers, our technique can, on average, record and replay 4322 lines of source code, versus 141 lines using traditional tracing methods.
Jeffrey B. Goeders, Steve Wilton
FCCM2
2015 An adaptive virtual overlay for fast trigger insertion for FPGA debug
abstract
Field-programmable gate-array (FPGA) platforms are commonly used for prototyping complex designs, allowing designers to evaluate and validate the functionality at speeds that are orders of magnitude faster than simulation. To counter the limited observability of hardware, on-chip trace buffers are used to record the behaviour of a small subset of signals. To effectively use the limited capacity of these on-chip trace buffers, trigger circuitry is required to determine when to start and/or stop recording signal behaviour. Although it is possible to implement the trigger circuitry and add it to the user circuit at compile time, this would require recompiling a design every time the trigger circuit is modified, reducing debug productivity. In this paper, we present and evaluate an adaptive virtual overlay architecture for rapid trigger implementation. The overlay is built from logic and routing resources not used by the user circuit, reducing the overhead and impact on the user circuit. At debug time, the pre-synthesised overlay architecture can quickly be configured to implement the desired trigger functionality. We show that our overlay architecture provides flexibility required for mapping trigger circuitry with negligible impact on delay. We also show trigger mapping is significantly faster rather than recompile insertion, increasing debug productivity.
Fatemeh Eslami, Steve Wilton
FPT2
2015 Using Round-Robin Tracepoints to debug multithreaded HLS circuits on FPGAs
abstract
High-level synthesis (HLS) for FPGA designs has gained significant traction in recent years. A key component in its adoption is allowing users to debug their hardware systems in the context of the original source code. This is becoming even more challenging as modern HLS tools enable the user to provide multithreaded source code for synthesis to hardware. Although recent work has begun to tackle source-level debugging of HLS circuits, none have addressed doing this in multithreaded circuits. In such systems it may be necessary to observe the behaviour of multiple threads for long run times in order to locate obscure or non-deterministic bugs and performance issues. In this paper we present a trace-based debugging architecture which records values from user-selected tracepoints into on-chip memories during circuit execution. The recorded values can be provided to the user as a cycle-accurate timeline of events to aid them in debugging multithreaded HLS circuits. We present a novel technique to allow multiple hardware threads to share trace buffers, effectively increasing the execution trace that can be recorded. This is accomplished by analyzing the control and data flow graph to determine the maximum rates at which each thread can encounter tracepoints, using this information to select which threads can share trace buffers, and automatically generating round-robin circuitry to arbitrate access to the buffers. Using this technique we are able to obtain an average of 4X improvement in trace length for an 8 thread system. This provides users with a longer timeline of execution and greater visibility into the execution of multithreaded HLS circuits.
Jeffrey B. Goeders, Steve Wilton
FPT2
2014 High-level synthesis-based design methodology for Dynamic Power-Gated FPGAs
abstract
Static leakage power consumption is critical in modern FPGAs for many applications. Dynamic Power-Gating (DPG), in which parts of the FPGA in-use logic blocks are powered-down at run-time, is a promising technique to reduce the static power. Adoption of such emerging DPG enabled FPGA architectures remains challenging as the current tool-chains to program the FPGA does not support this type of power-gating. Moreover, manually identifying profitable power-gating opportunities in an application requires significant design expertise and is time consuming. In this paper, we propose a high-level synthesis-based design framework that exploits the dynamic power-gating feature of the FPGAs to minimize the static power dissipation. We use this framework on a set of CHStone benchmark suite and demonstrate that power-gating opportunities for hardware accelerators can be identified in an automatic way. Results show that up to 96% reduction in static energy is achieved for individual accelerators using dynamic power-gating technique.
Assem A. M. Bsoul, Steve Wilton, Peter Hallschmid, Richard Klukas
FPL3
2014 Incremental distributed trigger insertion for efficient FPGA debug
abstract
FPGA-based prototyping enables evaluating complex designs directly in hardware, at speeds orders of magnitude faster than simulation. However, this approach suffers from the lack of observability during debugging. To enhance observability, designers insert debug instrumentation; trace buffers are used to record a small subset of data. Since these buffers have limited capacity, trigger circuits are required to start and/or stop recording based on the values of selected signals in the circuit. Although it is possible to insert trigger circuits at compile time, changing the trigger behaviour requires re-compiling the design, increasing the cost of each debug iteration. In this paper, we propose inserting trigger circuits at run-time by distributing trigger logic over spare resources of a fully placed-and-routed design such that its mapping is completely preserved. We also propose CAD optimizations which improve routability of the trigger circuitry, and minimize the impact on circuit delay. We find that using our techniques to implement the trigger logic can be an order of magnitude faster than a full recompilation.
Fatemeh Eslami, Steve Wilton
FPL2
2014 Effective FPGA debug for high-level synthesis generated circuits
abstract
High-level synthesis (HLS) promises to increase designer productivity in the face of steadily increasing FPGA sizes, and broaden the market of use, allowing software designers to reap the benefits of hardware implementation. One roadblock to HLS adoption is the lack of a debugging infrastructure. To debug, designers can run their source code on a processor; however, this does not capture interactions with other system components. The alternative is to debug using the RTL, which is beyond the expertise of software designers, and impractical for hardware designers as the RTL may not resemble the original source code.
Jeffrey B. Goeders, Steve Wilton
FPL2
2014 Accelerating FPGA debug: Increasing visibility using a runtime reconfigurable observation and triggering network
abstract
FPGA technology is commonly used to prototype new digital designs before entering fabrication. Whilst these physical prototypes can operate many orders of magnitude faster than through a logic simulator, a fundamental limitation is their lack of on-chip visibility when debugging. To counter this, trace-buffer-based instrumentation can be installed into the prototype, allowing designers to capture a predetermined window of signal data during live operation for offline analysis. However, instead of requiring the designer to recompile their entire circuit every time the window is modified, this article proposes that an overlay network is constructed using only spare FPGA routing multiplexers to connect all circuit signals through to the trace instruments. Thus, during debugging, designers would only need to reconfigure this network instead of finding a new place-and-route solution. Furthermore, we describe how this network can deliver signals to both the trigger and trace units of these instruments, which are implemented simultaneously using dual-port RAMs. Our results show that new network configurations connecting any subset of signals to 80--90% of the available RAM capacity can be computed in less than 70 seconds, for a 100,000 LUT circuit, as many times as necessary. Our tool—QuickTrace—is available for download.
Eddie Hung, Steve Wilton
ACM Trans. Design Autom. Electr. Syst.2
2014 Incremental Trace-Buffer Insertion for FPGA Debug
abstract
As integrated circuits encapsulate more functionality and complexity, verifying that these devices operate correctly under all scenarios is an increasingly difficult task. Rather than using traditional verification techniques such as software simulation, more and more designers are taking advantage of the significantly higher clock speeds that can be achieved by using field-programmable gate-array (FPGA)-based prototypes. A key challenge to these prototypes is the lack of on-chip observability during debugging; one popular solution is to insert trace-buffers into the design to record a limited set of internal signals, but modifying this trace configuration often requires the entire circuit to be recompiled. In this paper, we propose that the original circuit mapping is fully preserved and incremental techniques are used to eliminate the need for a full recompilation, thereby accelerating the debugging process. By exploiting two opportunities available during trace-insertion: the ability to connect from any point of a signal to any trace-pin, and the internal symmetry of the FPGA architecture, we find that incremental trace-insertion can be 98 times faster than a full recompilation, return a routing solution with a shorter wirelength, and have a negligible effect on the critical-path delay of the original circuit when reclaiming 75% of the leftover memory capacity for tracing.
Eddie Hung, Steve Wilton
IEEE Trans. Very Large Scale Integr. Syst.2
2013 Escaping the Academic Sandbox: Realizing VPR Circuits on Xilinx Devices
abstract
This paper presents a new, open-source method for FPGA CAD researchers to realize their techniques on real Xilinx devices. Specifically, we extend the Verilog-To-Routing (VTR) suite, which includes the VPR place-and-route CAD tool on which many FPGA innovations have been based, to generate working Xilinx bitstreams via the Xilinx Design Language (XDL). Currently, we can faithfully translate VPR's heterogeneous packing and placement results into an exact Xilinx `map' netlist, which is then routed by its `par' tool. We showcase the utility of this new method with two compelling applications targeting a 40nm Virtex-6 device: a fair comparison of the area, delay, and CAD runtime of academia's state-of-the-art VTR How with a commercial, closed-source equivalent, along with a CAD experiment evaluated using physical measurements of on-chip power consumption and die temperature, over time. This extended How - VTR-to-Bitstream - is released to the community with the hope that it can enhance existing research projects as well as unlock new ones.
Eddie Hung, Fatemeh Eslami, Steve Wilton
FCCM3
2013 Towards simulator-like observability for FPGAs: a virtual overlay network for trace-buffers
abstract
The rising complexity of verification has led to an increase in the use of FPGA prototyping, which can run at significantly higher operating frequencies and achieve much higher coverage than logic simulations. However, a key challenge is observability into these devices, which can be solved by embedding trace-buffers to record on-chip signal values. Rather than connecting a predetermined subset of circuits signals to dedicated trace-buffer inputs at compile-time, in this work we propose that a virtual overlay network is built to multiplex all on-chip signals to all on-chip trace-buffers. Subsequently, at debug-time, the designer can choose a signal subset for observation. To minimize its overhead, we build this network out of unused routing multiplexers, and by using optimal bipartite graph matching techniques, we show that any subset of on-chip signals can be connected to 80-90% of the maximum trace-buffer capacity in less than 50 seconds.
Eddie Hung, Steve Wilton
FPGA2
2013 Maximum flow algorithms for maximum observability during FPGA debug
abstract
Due to the ever-increasing density and complexity of integrated circuits, FPGA prototyping has become a necessary part of the design process. To enhance observability into these devices, designers commonly insert trace-buffers to record and expose the values on a small subset of internal signals during live operation to help root-cause errors. For dense designs, routing congestion will restrict the number of signals that can be connected to these trace-buffers. In this work, we apply optimal network flow graph algorithms, a well studied technique, to the problem of transporting circuit signals to embedded trace-buffers for observation. Specifically, we apply a minimum cost maximum flow algorithm to gain maximum signal observability with minimum total wirelength. We showcase our techniques on both theoretical FPGA architectures using VPR, and with a Xilinx Virtex6 device, finding that for the latter, over 99.6% of all spare RAM inputs can be reclaimed for tracing across four large benchmarks.
Eddie Hung, Al-Shahna Jamal, Steve Wilton
FPT3
2013 Post-Silicon Code Coverage for Multiprocessor System-on-Chip Designs
abstract
Effective techniques for post-silicon validation are required to better evaluate functional correctness of increasingly complex multi and many-core SoCs. However, there is little data evaluating the coverage of post-silicon validation efforts on industrial-scale designs. In this paper, we address this knowledge gap by instrumenting a nontrivial SoC with on-chip coverage monitors to measure the coverage achieved by typical post-silicon validation tests, such as booting the operating system (OS). We compare coverage achieved pre and post-silicon, and also measure the area overhead required to monitor post-silicon coverage. Our results show that the typical test of booting the OS often achieves high coverage, well correlated to what is achieved by pre-silicon directed tests, but in some blocks the coverage can be low or markedly different between pre and post-silicon, highlighting the importance of post-silicon validation in general and post-silicon coverage measurement in particular.
Kyle Balston, Mehdi Karimibiuki, Alan J. Hu, André Ivanov, Steve Wilton
IEEE Trans. Computers5
2013 Towards development of an analytical model relating FPGA architecture parameters to routability
abstract
We present an analytical model relating FPGA architectural parameters to the routability of the FPGA. The inputs to the model include the channel width and the connection and the switch block flexibilities. The output is an estimate of the proportion of nets in a large circuit that can be expected to be successfully routed on the FPGA. We assume that the circuit is routed to the FPGA using a single-step combined global/detailed router. We show that the model correctly predicts routability trends. We also present an example application to demonstrate that this model may be a valuable tool for FPGA architects. When combined with the earlier works on analytical modeling, our model can be used to quickly predict the routability without going through any stage of an expensive CAD flow. We envisage that this model will benefit FPGA architecture designers and vendors to quickly evaluate FPGA routing fabrics.
Joydip Das, Steve Wilton
ACM Trans. Reconfigurable Technol. Syst.2
2013 Scalable Signal Selection for Post-Silicon Debug
abstract
As modern integrated circuits increase in size and complexity, more and more verification effort is necessary to ensure their error-free operation. This has motivated designers to apply post-silicon debugging techniques to their designs, such as by embedding trace instrumentation within. However, a key drawback to this approach is that only a small subset of a chip's internal signals can be traced, but selecting the most effective signals to observe must be determined before fabrication and before the nature of any errors is known. This paper explores the tradeoff between the scalability of automated signal selection algorithms, and the amount of circuit observability that they offer. Three selection methods are presented: a technique that optimizes for observability directly; a method based on the graph-centrality of the circuit's connectivity; and a hybrid technique that combines both algorithms through exploiting the circuit hierarchy. To quantify the observability of each technique, we define the debug difficulty metric to measure how accurately the traced data can be used to resolve a circuit's state behavior. Although we find that the graph-based method offers the least observability of the three algorithms, it was the only method that could be applied to our largest benchmark of over 50 000 flip-flops, computing a selection in less than 90 s. Last, we present a novel application that can only be enabled by these scalable algorithms-speculative debug insertion for field-programmable gate arrays.
Eddie Hung, Steve Wilton
IEEE Trans. Very Large Scale Integr. Syst.2
2012 A configurable architecture to limit wakeup current in dynamically-controlled power-gated FPGAs
abstract
A dynamically-controlled power-gated (DCPG) FPGA architecture has recently been proposed to reduce static energy dissipation during idle periods. During a power mode transition from an off state to on state, the wakeup current drawn from power supplies causes a voltage droop on the power distribution network of a device. If not handled appropriately, this current and the associated voltage droop could cause malfunction of the design and/or the device. In DCPG FPGAs, the amount of wakeup current is not known beforehand as the structures of power-gated modules are application dependent; thus, a configurable solution is required to handle wakeup current. In this paper we propose a programmable wakeup architecture for DCPG FPGAs. The proposed solution has two levels: a fixed intra-region level and a configurable inter-region level. The architecture ensures that a power-gated module can be turned on such that the wakeup current constraints are not violated. We study the area and power overheads of the proposed solution. Our results show that the area overhead of the proposed inrush current limiting architecture is less than 2% for a power gating region of size 3x3 or 4x4 tiles, and the leakage power saved is more than 85% in a region of size 4x4 tiles.
Assem A. M. Bsoul, Steve Wilton
FPGA2
2012 Limitations of incremental signal-tracing for FPGA debug
abstract
Developing state-of-the-art custom silicon can be a prohibitively expensive and risky undertaking, due in no small part to the need to perform thorough design verification. Field-Programmable Gate-Arrays offer a flexible platform for constructing prototypes to aid in their verification, but unlike software simulation, observability into these prototypes is a major challenge. Designers can choose to insert trace-instrumentation to enhance on-chip observability, but doing so often requires re-compiling the entire design for each new trace configuration. This work presents two contributions: to explore the limitations of incremental-synthesis for trace-buffer insertion, and to propose CAD optimizations exclusive to this application for improving runtime and routability. We find that 99.4% of all used cluster outputs (driving both combinational and sequential circuit signals) can be incrementally-traced to 75% of the free memory-capacity on an FPGA, an order of magnitude quicker than the original compilation and with a nominal impact on circuit delay, for a 20% minimum channel width (10% area) increase.
Eddie Hung, Steve Wilton
FPL2
2012 An FPGA with power-gated switch blocks
abstract
Static power consumption is an important component of the total power consumption in FPGAs built using 90nm and smaller technology nodes. A previous study proposed powering down regions of logic blocks in an FPGA when idle to reduce the static power dissipation. This previous work did not consider powering down the switch blocks (SBs). However, the static power of SBs constitute more than 50% of an FPGA's static power. In this paper, we present an architecture that enables selectively powering down SBs along with the logic blocks during their idle periods. The potential power savings from this architecture depends on the proportion of SBs that can be powered down. We present modifications to our CAD flow to maximize the number of such SBs, and we experimentally estimate their proportion using a set of synthetic benchmark circuits. Our estimation results show that 53% to 83% of the SBs can be powered down in a functional module of size 24×24 tiles and an architecture power gating regions of size 4×4 tiles, leading to overall static power reductions of 70% to 84% compared to an architecture that does not support power gating.
Assem A. M. Bsoul, Steve Wilton
FPT2
2012 VersaPower: Power estimation for diverse FPGA architectures
abstract
This paper presents VersaPower, a tool capable of modelling the power usage of many different field programmable gate array (FPGA) architectures.The latest release of the academic FPGA CAD tool, Versatile Place and Route 6.0 (VPR), supports new architecture features such as fracturable look-up tables and complex logic blocks. Past FPGA power models do not support these new features. VersaPower is designed to work closely with VPR to provide power estimation for any architecture supported by this new CAD flow. This allows researchers to investigate the effects on power usage of both new FPGA architectures, as well as new CAD algorithms. VersaPower is designed to operate with modern CMOS technologies, and is validated against SPICE using 22 nm, 45 nm and 130 nm technologies. Results show that for common architectures, roughly 60% HDL of power consumption is due to the routing fabric, 30% from logic blocks and 10% from the clock network. Architectures ODN supporting fracturable LUTs require 5-10% more power, as each CLB has additional I/O pins, increasing the sizes of local interconnect crossbars and connection boxes.
Jeffrey B. Goeders, Steve Wilton
FPT2
2012 Rapid RTL-based signal ranking for FPGA prototyping
abstract
As the capacity of integrated circuits increases, it is becoming increasingly difficult to ensure that a chip is free of design errors. Designers are increasingly turning to FPGA prototyping platforms to validate their designs much more extensively than is possible using simulation. A key challenge is one of visibility; signals can only be observed if they can be driven to pins of a chip. To enhance visibility during debug, designers regularly instrument their design with on-chip circuitry to record a small subset of signals at-speed for later off-chip analysis. The selection of which signals should be recorded critically affects the effectiveness of this approach. In this paper, we present an algorithm that ranks all signals in a design based on their predicted importance during validation. Compared to previous techniques, which analyze the circuit at the gate level, our algorithm works directly on the parse-tree representation of the circuit, and hence is orders of magnitude faster than these previous techniques. Our algorithm has been implemented as an integral part of Tektronix Certus, a commercial validation suite.
Steve Wilton, Bradley R. Quinton, Eddie Hung
FPT1
2012 Hierarchical Benchmark Circuit Generation for FPGA Architecture Evaluation
abstract
We describe a stochastic circuit generator that can be used to automatically create benchmark circuits for use in FPGA architecture studies. The circuits consist of a hierarchy of interconnected modules, reflecting the structure of circuits designed using a system-on-chip design flow. Within each level of hierarchy, modules can be connected in a bus, star, or dataflow configuration. Our circuit generator is calibrated based on a careful study of existing system-on-chip circuits. We show that our benchmark circuits lead to more realistic architectural conclusions than circuits generated using previous generators.
Cindy Mark, Scott Y. L. Chin, Lesley Shannon, Steve Wilton
ACM Trans. Embed. Comput. Syst.4
2012 Formal-Analysis-Based Trace Computation for Post-Silicon Debug
abstract
This paper presents a post-silicon debug methodology that provides a means to rewind, or backspace, a chip from a known crash state using a combination of on-chip real-time data collection and off-chip formal analysis methods. A complete debug flow is presented that considers practical considerations such as area, on-chip non-determinism and signal propagation delay. This flow, along with a low-overhead breakpoint circuit, allows for state-accurate breakpointing capabilities without the need to monitor the entire state of the chip. The flow and associated hardware was tested using a hardware prototype, which consists of an OpenRISC processor instrumented with the debug hardware connected to a PC running the formal verification algorithms. Traces hundreds of cycles long were obtained using the methodology presented in this paper.
Marcel Gort, Flavio M. de Paula, Johnny J. W. Kuan, Tor M. Aamodt, Alan J. Hu, Steve Wilton, Jin Yang 0006
IEEE Trans. Very Large Scale Integr. Syst.6
2012 Optimizing Floating Point Units in Hybrid FPGAs
abstract
This paper introduces a methodology to optimize coarse-grained floating point units (FPUs) in a hybrid field-programmable gate array (FPGA), where the FPU consists of a number of interconnected floating point adders/subtracters (FAs), multipliers (FMs), and wordblocks (WBs). The wordblocks include registers and lookup tables (LUTs) which can implement fixed point operations efficiently. We employ common subgraph extraction to determine the best mix of blocks within an FPU and study the area, speed and utilization tradeoff over a set of floating point benchmark circuits. We then explore the system impact of FPU density and flexibility in terms of area, speed, and routing resources. Finally, we derive an optimized coarse-grained FPU by considering both architectural and system-level issues. This proposed methodology can be used to evaluate a variety of FPU architecture optimizations. The results for the selected FPU architecture optimization show that although high density FPUs are slower, they have the advantages of improved area, area-delay product, and throughput.
Chi Wai Yu, Alastair M. Smith, Wayne Luk, Philip H. W. Leong, Steve Wilton
IEEE Trans. Very Large Scale Integr. Syst.5
2011 Towards scalable FPGA CAD through architecture
abstract
Long FPGA CAD runtime has emerged as a limitation to the future scaling of FPGA densities. Already, compile times on the order of a day are common, and the situation will only get worse as FPGAs get larger. Without a concerted effort to reduce compile times, further scaling of FPGAs will eventually become impractical.
Scott Y. L. Chin, Steve Wilton
FPGA2
2011 An analytical model relating FPGA architecture parameters to routability
abstract
We present an analytical model relating FPGA architectural parameters to the routability of the FPGA. The inputs to the model include the channel width and connection and switch block flexibilities, and the output is an estimate of the proportion of nets in a large circuit that can be expected to be routed on the FPGA. We assume that the circuit is routed to the FPGA using a single-step combined global/detailed router. Together with the earlier works on analytical modeling, our model can be used to predict the routability without going through an expensive CAD flow. We show that the model correctly predicts routability trends.
Joydip Das, Steve Wilton
FPGA2
2011 Speculative Debug Insertion for FPGAs
abstract
FPGA prototypes have become an increasingly important part of the overall integrated circuit design and verification flow, providing the ability to test an integrated circuit running at (near) speed with realistic inputs and outputs. When unexpected behaviour is observed in the prototype, it is necessary to determine the source of this behaviour, this usually requires observing signals that are internal to one of the devices in the prototype. Tools currently exist to enable FPGAs to be instrumented, but these are normally used in a reactive manner, that is, instrumentation is only added after incorrect behaviour has been observed. In this paper, we propose speculative debug insertion, in which a tool automatically predicts what signals will be useful during debug, and instruments the design during the first compilation. If done correctly, this can significantly accelerate the debug process, especially for large prototypes containing many FPGAs. However, it is important that this does not negatively affect the performance, capacity, power, or compilation time. We show that speculative debug insertion is possible, and experimentally evaluate the limits to speculative insertion.
Eddie Hung, Steve Wilton
FPL2
2011 Accelerated FPGA architecture design: Capabilities and limitations of analytical models
abstract
FPGA architects typically use experimental techniques to design new architectures. These techniques are time consuming, thus limiting the number of the architectures that can be investigated. Some previous works use analytical models to significantly accelerate the design of a new architecture. To properly capitalize on the benefits of the analytical models, the designers need to have an understanding of the capabilities and the limitations of the analytical models. In this paper, we use two representative architecture questions to provide such understanding. These two questions respectively investigate the optimization of a general-purpose FPGA architecture and the optimization of an application-specific FPGA architecture. For an optimized general purpose architecture, we show that the conclusions made by the analytical models are similar to the experimental techniques, with respect to three different design goals: area, delay and area-delay trade-off. This justifies the use of the analytical models in optimizing general-purpose FPGA architectures. We also find that the analytical models can not capture the behavior of `some' applications that contain `discrete effects'. We present this later finding and the related explanations to show that the analytical models can not optimize application-specific architectures in some cases.
Joydip Das, Steve Wilton
FPT2
2011 Performance and Cost Tradeoffs in Metal-Programmable Structured ASICs (MPSAs)
abstract
As process technology scales, the design effort and nonrecurring engineering (NRE) costs associated with the development of integrated circuits is becoming extremely high. Structured ASICs offer one solution to these problems. However, to realize their full potential, their performance and cost advantages, architectures, and CAD must be fully understood. We believe that this can lead to wider adoption of structured ASICs. In this paper, we take a step in this direction and investigate the area, delay, power, and cost tradeoffs in metal-programmable structured ASICs (MPSAs). In particular, we quantify the impact of the number of user-defined (custom) metal mask layers on these metrics. Results indicate that for lowest cost, the number of custom layers should be minimized, especially for small die sizes (e.g., less than 100${\hbox {mm}}^{2}$). Delay and power, however, can be improved by a few additional custom layers. With two custom metal layers, MPSAs can be 2$\times$–10$\times$cheaper than cell-based ICs (CBICs).
Usman Ahmed, Guy Lemieux, Steve Wilton
IEEE Trans. Very Large Scale Integr. Syst.3
2011 An Analytical Model Relating FPGA Architecture to Logic Density and Depth
abstract
This paper presents an analytical model that relates FPGA architectural parameters to the logic size and depth of an FPGA implementation. In particular, the model relates the lookup-table size, the cluster size, and the number of inputs per cluster to the amount of logic that can be packed into each lookup-table and cluster, the number of used inputs per cluster, and the depth of the circuit after technology mapping and clustering. Comparison to experimental results shows that our model has good accuracy. We illustrate how the model can be used in FPGA architectural investigations to complement the experimental approach. The model's accuracy, combined with the simple form of the equations, make them a powerful tool for FPGA architects to better understand and guide the development of future FPGA architectures.
Joydip Das, Andrew Lam, Steve Wilton, Philip H. W. Leong, Wayne Luk
IEEE Trans. Very Large Scale Integr. Syst.3
2010 The impact of interconnect architecture on via-programmed structured ASICs (VPSAs)
abstract
In this paper, we evaluate the performance of an FPGA-like interconnect fabric for structured ASICs which is based upon fixed metal and programmable vias. We call this type of device a via-programmed structured ASIC or VPSA. We look at two different types of VPSA routing fabrics: one uses jumper wiring and the other uses crossover wiring. The performance of these fabrics is compared against an ASIC-like interconnect fabric, otherwise known as a metal-programmed structured ASIC or MPSA, which can be configured by customizing metal and via layers. We study the impact of these routing fabrics on cost, area, power and delay metrics. The results for different fabrics span a wide range, suggesting the routing architecture plays a very important role in their overall performance and it should be thoroughly researched.
Usman Ahmed, Guy Lemieux, Steve Wilton
FPGA3
2010 An FPGA architecture supporting dynamically controlled power gating
abstract
Leakage power is an important component of the total power consumption in FPGAs built using 90 nm and smaller technology nodes. Power gating, in which regions of the chip can be powered down, has been shown to be effective at reducing leakage power. However, previous techniques focus on statically-controlled power gating. In this paper, we propose a modification to the fabric of an FPGA that enables dynamically-controlled power gating, in which logic clusters can be selectively powered-down at run-time. For applications containing blocks with large idle times, this could lead to significant leakage power savings. Our architecture utilizes the existing routing fabric and unused input pins of logic clusters to route the power control signals. No modifications to the existing routing algorithms are required to support the new architecture. We study the area and power tradeoffs by varying the basic architecture parameters of an FPGA, and by varying the size of the power gating regions. We also study the leakage energy savings using a model that characterizes an application in terms of its structure and behavior. We show less than 1% of area overhead for a power gating region size of 3X3 logic tiles. Using the application model, we show that up to 40% leakage energy reduction can be achieved using the proposed architecture for different application parameters, not including power dissipated by the power state controller.
Assem A. M. Bsoul, Steve Wilton
FPT2
2010 Energy Optimization for Many-Core Platforms: Communication and PVT Aware Voltage-Island Formation and Voltage Selection Algorithm
abstract
In this paper, we propose a novel approach to voltage-island formation, for the energy optimization of many-core architectures, which mitigates the impact of process, voltage, and temperature (PVT) variations. The islands are created by balancing their shape constraints imposed by intra and inter-island communication with the desire to limit the spatial extent of each island to minimize PVT impact. In addition, to reduce the number of voltage levels in the design, we propose an efficient voltage selection approach that provides near optimal results, for a set of 33 examined cases, with more than a ten times speedup compared to the best-known previous methods. This run-time improvement is important, especially for large many-core platforms. Finally, we present an evaluation platform considering pre-fabrication and post-fabrication PVT scenarios where multiple applications with hundreds to thousands of tasks are mapped onto many-core platforms with hundreds to thousands of cores to evaluate the proposed techniques. Results show that the average energy savings for 33 test cases using the proposed methods are 37% compared to 16% obtained using previous methods.
Sohaib Majzoub, Res Saleh, Steve Wilton, Rabab K. Ward
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2009 Wirelength modeling for homogeneous and heterogeneous FPGA architectural development
abstract
This paper describes an analytical model that relates the architectural parameters of an FPGA to the average prerouting wirelength of an FPGA implementation. Both homogeneous and heterogeneous FPGAs are considered. For homogeneous FPGAs, the model relates the lookup-table size, the cluster size, and the number of inputs per cluster to the expected wirelength. For heterogeneous FPGAs, the number and positioning of the embedded blocks, as well as the number of pins on each embedded block is considered. Two applications of the model to FPGA architectural design are also presented.
Alastair M. Smith, Steve Wilton, Joydip Das
FPGA2
2009 An analytical model relating FPGA architecture and place and route runtime
abstract
This paper presents an analytical model that relates the architectural parameters of an FPGA to the place-and-route runtimes of the FPGA CAD tools. We consider both a simulated annealing based placement algorithm employing a bounding box wirelength cost function, and a negotiation based A*router. We also show an example application of the model in early architecture evaluation.
Scott Y. L. Chin, Steve Wilton
FPL2
2009 Improving the memory footprint and runtime scalability of FPGA CAD algorithms
abstract
Advances in process technology have allowed for a dramatic increase in the capacity of FPGAs and this scaling is continuing at a steady pace. However, this scaling places increasing demands on the FPGA CAD tools. Already, for very large designs, compile-times of an entire work day are common, and memory requirements that exceed what would be found in a common desktop workstation are the norm. As FPGAs continue to grow, the problem will become worse. Unless the scalability of FPGA CAD tools is addressed, the long run times and large memory footprints will become a hindrance to future FPGA scaling, leading to increased costs and design times for companies who use these devices. The proposed research focuses on both the memory and runtime scalability of FPGA CAD tools. We have presented work on effective methods to improve the memory scalability. This work is summarized in Section 3. We are currently focusing on the runtime scalability portion of the project. Our current progress and proposed research on this part was summarized in Section 4.
Scott Y. L. Chin, Steve Wilton
FPL2
2009 Modeling post-techmapping and post-clustering FPGA circuit depth
abstract
This paper presents an analytical model that relates FPGA architectural parameters to the expected speed of FPGA implementation. More precisely, the model relates the lookup-table size, cluster size, and number of inputs per cluster to the depth of the circuit after technology mapping and after clustering. Comparison to experimental results with large MCNC circuits shows that our models are accurate. We show how the models can be used in FPGA architectural investigations to complement the more usual experimental approach.
Joydip Das, Steve Wilton, Philip H. W. Leong, Wayne Luk
FPL2
2009 A detailed delay path model for FPGAs
abstract
A complete circuit-level description of a representative FPGA is presented in this paper, from which a simple RC delay model as a function of architectural and technology parameters is derived. Using this model, the expression for the optimal delay of any path through the FPGA can be formulated. We distill our model into being purely architecture dependent, and use it to capture new insight into how FPGA parameters can directly affect its delay. Several applications of this model are: (1) to gain better intuition of how architecture and process parameters affect the delay path in an FPGA, (2) for initial studies into new circuit designs and integrated circuit technologies, (3) in CAD tools for optimisation and sensitivity analysis. The technique described can be applied to arbitrary circuits, and simulations show that our closed form equations give delay values that are accurate to approximately 10% when compared to HSPICE simulation.
Eddie Hung, Steve Wilton, Haile Yu, Thomas C. P. Chau, Philip H. W. Leong
FPT2
2009 Concurrently optimizing FPGA architecture parameters and transistor sizing: Implications for FPGA design
abstract
This paper presents a method that combines high-level and low-level architecture parameter exploration. The paper builds on an increasing body of work concerned with modeling reconfigurable architectures, and presents a full area and delay model of an FPGA. The optimization of this model is based on the use of geometric programming, and allows high-level architecture parameter selection and transistor sizing to be done concurrently. We use the framework to demonstrate that concurrent optimization of both high and low-level parameters can lead to significantly different architectural conclusions.
Alastair M. Smith, George A. Constantinides, Steve Wilton, Peter Y. K. Cheung
FPT3
2009 Static and Dynamic Memory Footprint Reduction for FPGA Routing Algorithms
abstract
This article presents techniques to reduce the static and dynamic memory requirements of routing algorithms that target field-programmable gate arrays. During routing, memory is required to store both architectural data and temporary routing data. The architectural data is static, and provides a representation of the physical routing resources and programmable connections on the device. We show that by taking advantage of the regularity in FPGAs, we can reduce the amount of information that must be explicitly represented, leading to significant memory savings. The temporary routing data is dynamic, and contains scoring parameters and traceback information for each routing resource in the FPGA. By studying the lifespan of the temporary routing data objects, we develop several memory management schemes to reduce this component. To make our proposals concrete, we applied them to the routing algorithm in VPR and empirically quantified the impact on runtime memory footprint, and place and route time.
Scott Y. L. Chin, Steve Wilton
ACM Trans. Reconfigurable Technol. Syst.2
2009 Floating-Point FPGA: Architecture and Modeling
abstract
This paper presents an architecture for a reconfigurable device that is specifically optimized for floating-point applications. Fine-grained units are used for implementing control logic and bit-oriented operations, while parameterized and reconfigurable word-based coarse-grained units incorporating word-oriented lookup tables and floating-point operations are used to implement datapaths. In order to facilitate comparison with existing FPGA devices, the virtual embedded block scheme is proposed to model embedded blocks using existing field-programmable gate array (FPGA) tools. This methodology involves adopting existing FPGA resources to model the size, position, and delay of the embedded elements. The standard design flow offered by FPGA and computer-aided design vendors is then applied and static timing analysis can be used to estimate the performance of the FPGA with the embedded blocks. On selected floating-point benchmark circuits, our results indicate that the proposed architecture can achieve four times improvement in speed and 25 times reduction in area compared with a traditional FPGA device.
Chun Hok Ho, Chi Wai Yu, Philip H. W. Leong, Wayne Luk, Steve Wilton
IEEE Trans. Very Large Scale Integr. Syst.5
2009 Programmable Logic Core Enhancements for High-Speed On-Chip Interfaces
abstract
Programmable logic cores (PLCs) offer a means of providing post-fabrication reconfigurability to a SoC design. This ability has the potential to significantly enhance the SoC design process by enabling post-silicon debugging, design error correction and post-fabrication feature enhancement. However, circuits implemented in general purpose programmable logic will inevitably have lower timing performance than fixed function circuits. This fundamental mismatch makes it difficult to use the PLC effectively. We address this problem by proposing changes to the structure of the PLC itself; these architectural enhancements enable circuit implementations with high performance interfaces. In previous work we addressed system bus interfaces, in this work we address direct synchronous interfaces. Our results show significant improvement in PLC interface timing, such that interaction with full-speed fixed-function SoC logic is possible. Our enhanced PLCs are able to implement direct synchronous interfaces running at, on average, 662 MHz (compared to 249 MHz in regular programmable logic). We are able to do this without compromising the basic structure or routiblity of the programmable fabric. At the same time, we show that the area overhead for these architectural changes was approximately 1%.
Bradley R. Quinton, Steve Wilton
IEEE Trans. Very Large Scale Integr. Syst.2
2008 BackSpace: Formal Analysis for Post-Silicon Debug
abstract
Post-silicon debug is the problem of determining what's wrong when the fabricated chip of a new design behaves incorrectly. This problem now consumes over half of the overall verification effort on large designs, and the problem is growing worse. We introduce a new paradigm for using formal analysis, augmented with some on-chip hardware support, to automatically compute error traces that lead to an observed buggy state, thereby greatly simplifying the post-silicon debug problem. Our preliminary simulation experiments demonstrate the potential of our approach: we can "backspace" hundreds of cycles from randomly selected states of some sample designs. Our preliminary architectural studies propose some possible implementations and show that the on-chip overhead can be reasonable. We conclude by surveying future research directions.
Flavio M. de Paula, Marcel Gort, Alan J. Hu, Steve Wilton, Jin Yang 0006
FMCAD4
2008 Rapid estimation of power consumption for hybrid FPGAs
abstract
A hybrid FPGA consists of island-style fine-grained units and domain-specific coarse-grained units. This paper describes an approach to estimate the power consumption of a set of hybrid FPGA architectures. The dynamic power consumption of the fine-grained units is obtained using standard FPGA tools, and the coarse-grained units using standard ASIC tools. Based on this approach, the dynamic power consumption of different hybrid FPGA architectures can be studied and we report on results over a set of floating point benchmark circuits.
Chun Hok Ho, Philip H. W. Leong, Wayne Luk, Steve Wilton
FPL4
2008 An analytical model describing the relationships between logic architecture and FPGA density
abstract
This paper describes an analytical model, based principally on Rentpsilas Rule, that relates logic architectural parameters to the area efficiency of an FPGA. In particular, the model relates the lookup-table size, the cluster size, and the number of inputs per cluster to the amount of logic that can be packed into each lookup-table and cluster, and the number of used inputs per cluster. Comparison to experimental results show that our models are accurate. This accuracy combined with the simple form of the equations make them a powerful tool for FPGA architects to better understand and guide the development of future FPGA architectures.
Andrew Lam, Steve Wilton, Philip H. W. Leong, Wayne Luk
FPL2
2008 A system-level stochastic circuit generator for FPGA architecture evaluation
abstract
We describe a stochastic circuit generator that can be used to automatically create benchmark circuits for use in FPGA architecture studies. The circuits consist of a hierarchy of interconnected modules, reflecting the structure of circuits designed using a system-on-chip design flow. Within each level of hierarchy, modules can be connected in a bus, star, or dataflow configuration. Our circuit generator is calibrated based on a careful study of existing SoC circuits. We compare our circuits to those generated by previous circuit generators, and characterize our circuits with respect to the type of network used to connect modules.
Cindy Mark, Ava Shui, Steve Wilton
FPT3
2008 Optimizing coarse-grained units in floating point hybrid FPGA
abstract
This paper introduces a novel methodology to optimize coarse-grained floating point units (FPUs) in a hybrid FPGA. We employ common subgraph extraction to determine the number of floating point adders/subtracters (FAs), multipliers (FMs) and wordblocks (WBs) in the FPUs. We flrst study the area, speed and utilization trade-off of the selected FPU subgraphs in a set of floating point benchmark circuits. We then explore the impact of density and flexibility of FPUs on the system in terms of area, speed and routing resources. We derive an optimized coarse-grained FPU by considering both architectural and system level issues. The results show that: (1) embedding more types of coarse-grained FPU in the system causes at most 21.3% increase in delay, (2) the area of the system can be reduced by 27.4% by embedding high density subgraphs, (3) the high density subgraphs requires 14.8% fewer routing resources.
Chi Wai Yu, Alastair M. Smith, Wayne Luk, Philip H. W. Leong, Steve Wilton
FPT5
2008 On the trade-off between power and flexibility of FPGA clock networks
abstract
FPGA clock networks consume a significant amount of power, since they toggle every clock cycle and must be flexible enough to implement the clocks for a wide range of different applications. The efficiency of FPGA clock networks can be improved by reducing this flexibility; however, reducing the flexibility introduces stricter constraints during the clustering and placement stages of the FPGA CAD flow. These constraints can reduce the overall efficiency of the final implementation. This article examines the trade-off between the power consumption and flexibility of FPGA clock networks. Specifically, this article makes three contributions. First, it presents a new parameterized clock-network framework for describing and comparing FPGA clock networks. Second, it describes new clock-aware placement techniques that are needed to find a legal placement satisfying the constraints imposed by the clock network. Finally, it performs an empirical study to examine the trade-off between the power consumption of the clock network and the impact of the CAD constraints for a number of different clock networks with varying amounts of flexibility. The results show that the techniques used to produce a legal placement can have a significant influence on power and the ability of the placer to find a legal solution. On average, circuits placed using the most effective techniques dissipate 5% less overall energy and are significantly more likely to be legal than circuits placed using other techniques. Moreover, the results show that the architecture of the clock network is also important. On average, FPGAs with an efficient clock network are up to 14.6% more energy efficient compared to other FPGAs.
Julien Lamoureux, Steve Wilton
ACM Trans. Reconfigurable Technol. Syst.2
2008 A Synthesizable Datapath-Oriented Embedded FPGA Fabric for Silicon Debug Applications
abstract
We present an architecture for a synthesizable datapath-oriented FPGA core that can be used to provide post-fabrication flexibility to an SoC. Our architecture is optimized for bus-based operations and employs a directional routing architecture, which allows it to be synthesized using standard ASIC design tools and flows. The primary motivation for this architecture is to provide an efficient mechanism to support on-chip debugging. The fabric can also be used to implement other datapath-oriented circuits such as those needed in signal processing and computation-intensive applications. We evaluate our architecture using a set of benchmark circuits and compare it to previous fabrics in terms of area, speed, and power.
Steve Wilton, Chun Hok Ho, Bradley R. Quinton, Philip H. W. Leong, Wayne Luk
ACM Trans. Reconfigurable Technol. Syst.1
2008 GlitchLess: Dynamic Power Minimization in FPGAs Through Edge Alignment and Glitch Filtering
abstract
This paper describes GlitchLess, a circuit-level technique for reducing power in field-programmable gate arrays (FPGAs) by eliminating unnecessary logic transitions called glitches. This is done by adding programmable delay elements to the logic blocks of the FPGA. After routing a circuit and performing static timing analysis, these delay elements are programmed to align the arrival times of the inputs of each lookup table (LUT), thereby preventing new glitches from being generated. Moreover, the delay elements also behave as filters that eliminate other glitches generated by upstream logic or off-chip circuitry. On average, the proposed implementation eliminates 87% of the glitching, which reduces overall FPGA power by 17%. The added circuitry increases the overall FPGA area by 6% and critical-path delay by less than 1%. Furthermore, since it is applied after routing, the proposed technique requires little or no modifications to the routing architecture or computer-aided design (CAD) flow.
Julien Lamoureux, Guy Lemieux, Steve Wilton
IEEE Trans. Very Large Scale Integr. Syst.3
2008 Practical Asynchronous Interconnect Network Design
abstract
The implementation of interconnect is becoming a significant challenge in modern integrated circuit (IC) design. Both synchronous and asynchronous strategies have been suggested to manage this problem. Creating a low skew clock tree for synchronous inter-block pipeline stages is a significant challenge. Asynchronous interconnect does not require a global clock, and therefore, it has a potential advantage in terms of design effort. This paper presents an asynchronous interconnect design that can be implemented using a standard application-specific IC flow. This design is considered across a range of IC interconnect scenarios. The results demonstrate that there is a region of the design space where the implementation provides an advantage over a synchronous interconnect by removing the need for clocked inter-block pipeline stages, while maintaining high throughput. Further results demonstrate a computer-aided design tool enhancement that would significantly increase this space. A detailed comparison of power, area, and latency of the two strategies is also provided for a range of IC scenarios.
Bradley R. Quinton, Mark R. Greenstreet, Steve Wilton
IEEE Trans. Very Large Scale Integr. Syst.3
2007 GlitchLess: an active glitch minimization technique for FPGAs
abstract
This paper describes a technique that reduces dynamic power in FPGAs by reducing the number of glitches in the global routing resources. The technique involves adding programmable delay elements within the logic blocks of an FPGA to programmably align the arrival times of early-arriving signals to the inputs of the lookup tables and to filter out glitches generated by earlier circuitry. On average, the proposed technique eliminates 91% of the glitching, which reduces overall FPGA power by 18%. The added circuitry increases overall area by 5% and critical-path delay by less than 1%. Furthermore, since it is applied after routing, the proposed technique requires no modifications to the existing FPGA routing architecture or CAD flow.
Julien Lamoureux, Guy Lemieux, Steve Wilton
FPGA3
2007 A synthesizable datapath-oriented embedded FPGA fabric
abstract
We present an architecture for a synthesizable datapath-oriented Field Programmable Gate Array (FPGA) core which can be used to provide post-fabrication flexibility to a System-on-Chip (SoC). Our architecture is optimized for bus-based operations that are common in signal processing and computation intensive applications. It employs a directional routing architecture, which allows it to be synthesized using standard ASIC design tools and flows. We also describe a proof-of-concept layout of our core. It is shown that the proposed architecture is significantly more area efficient than the best previously reported synthesizable programmable logic core.
Steve Wilton, Chun Hok Ho, Philip H. W. Leong, Wayne Luk, Bradley R. Quinton
FPGA1
2007 Domain-Specific Hybrid FPGA: Architecture and Floating Point Applications
abstract
This paper presents a novel architecture for domain-specific FPGA devices. This architecture can be optimised for both speed and density by exploiting domain-specific information to produce efficient reconfigurable logic with multiple granularity. In the reconfigurable logic, general-purpose finegrained units are used for implementing control logic and bit-oriented operations, while domain-specific coarse-grained units and heterogeneous blocks are used for implementing datapaths; the precise amount of each type of resources can be customised to suit specific application domains. Issues and challenges associated with the design flow and the architecture modelling are addressed. Examples of the proposed architecture for speeding up floating point applications are illustrated. Current results indicate that the proposed architecture can achieve 2.5 times improvement in speed and 18 times reduction in area on average, when compared with traditional FPGA devices on selected floating point benchmark circuits.
Chun Hok Ho, Chi Wai Yu, Philip H. W. Leong, Wayne Luk, Steve Wilton
FPL5
2007 Clock-Aware Placement for FPGAs
abstract
The programmable clock networks in FPGAs have a significant impact on overall power, area, and delay. Not only does the clock network itself dissipate a significant amount of power, since it connects to every latch on the FPGA and toggles every cycle, but the design of the clock network also affects how efficiently the rest of the application can be implemented since it imposes constraints on the CAD tools which map the application onto the FPGA. To examine this tradeoff, this paper describes and compares new clock-aware placement techniques and then examines how the clock network architecture affects overall power, area, and delay. Our results show that the placement techniques used to make placement clock-aware have a significant influence on power and delay. On average, circuits placed using the most effective techniques dissipate 9.9% less energy and were 2.4% faster than circuits placed using the least effective techniques. Moreover, the results show that the clock network architecture is also important. On average, FPGAs with an efficient clock network were up to 12.5% more energy efficient and 7.2% faster than other FPGAs.
Julien Lamoureux, Steve Wilton
FPL2
2007 Embedded Programmable Logic Core Enhancements for System Bus Interfaces
abstract
Programmable logic cores (PLCs) offer a means of providing post-fabrication re-configurability to a SoC design. Circuits implemented in a PLC will inevitably have lower timing performance and logic density than fixed function circuits. This fundamental mismatch makes the design of the interface between the PLC and the rest of the SoC a challenging problem. In this paper we focus on interfaces between circuits implemented in PLCs and SoC system busses. We demonstrate problems with existing implementation options and then propose modifications to parts of the PLC architecture to enable more efficient system bus interfaces. Our results show that, on average, this modified architecture improves interface timing by 36.4%, reduces CLB usage by 7.9% and improves routability by 28.8% for circuits that require system bus interfaces. We show that the area overhead is less than 0.5% for circuits that do not require bus interfaces.
Bradley R. Quinton, Steve Wilton
FPL2
2007 Memory Footprint Reduction for FPGA Routing Algorithms
abstract
In this paper, we present a technique to reduce the run-time memory footprint of FPGA routing algorithms. These algorithms require a representation of the physical routing resources and programmable connections on the device; this representation dominates the storage requirements of FPGA routers. We show that by taking advantage of the tile-based nature of FPGAs, we can reduce the amount of information that must be explicitly represented, leading to significant memory savings. To make our proposal concrete, we applied it to the routing algorithm in VPR and quantified the impact on run-time memory footprint, and place and route compile-time. We found that a memory reduction of 5X to 13X could be achieved at a routing runtime penalty of 2.26X and an overall place-and-route runtime penalty of 1.28X.
Scott Y. L. Chin, Steve Wilton
FPT2
2006 Virtual Embedded Blocks: A Methodology for Evaluating Embedded Elements in FPGAs
abstract
Embedded elements, such as block multipliers, are increasingly used in advanced field programmable gate array (FPGA) devices to improve efficiency in speed, area and power consumption. A methodology is described for assessing the impact of such embedded elements on efficiency. The methodology involves creating dummy elements, called virtual embedded blocks (VEBs), in the FPGA to model the size, position and delay of the embedded elements. The standard design flow offered by FPGA and CAD vendors can be used for mapping, placement, routing and retiming of designs with VEBs. The speed and resource utilisation of the resulting designs can then be inferred using the FPGA vendor's timing analysis tools. We illustrate the application of this methodology to the evaluation of various schemes of involving embedded elements that support floating-point computations
Chun Hok Ho, Philip H. W. Leong, Wayne Luk, Steve Wilton, Sergio López-Buedo
FCCM4
2006 FPGA clock network architecture: flexibility vs. area and power
abstract
This paper examines the tradeoffs between flexibility, area, and power dissipation of programmable clock networks for Field-Programmable Gate Arrays (FPGA's). The paper begins by describing a parameterized clock network model that describes a broad range of programmable clock network architectures. Specifically, the model supports architectures with multiple local and global clock domains and varying amounts of flexibility at various levels of the clock network. Using the model, the architectural parameters that control the flexibility of the clock network are varied to determine the cost of this flexibility in terms of area and power dissipation. From these experiments, the study finds that area and power costs are highest for networks with flexibility close to the logic blocks. Furthermore, it found that clock networks with local clock domains have little overhead and are significantly more efficient than clock networks without local clock domains for applications with multiple clocks.
Julien Lamoureux, Steve Wilton
FPGA2
2006 Power Implications of Implementing Logic Using FPGA Embedded Memory Arrays
abstract
This paper investigates the power and energy implications of using embedded FPGA memory arrays to implement logic. Previous studies have shown that this technique provides extremely dense implementations of some types of logic circuits, however, these previous studies did not evaluate the impact on power. The authors measure the effects on power and energy as a function of three architectural parameters: the number of available memory arrays, the size of the memory arrays, and the flexibility of the memory arrays. It was shown in this paper that although embedded memories provide area efficient implementations of many circuits, this technique results in additional power consumption. When power can be traded off for density, it was also shown that for most array sizes, the arrays should be as flexible as possible, and that smaller memory arrays are more power efficient than large arrays. When larger arrays are desired for more density improvement, non-square memories with more rows than columns are better. The results were obtained from fully place and routed circuits using modified versions of VPR and the Poon power model. Several results were also verified through measurements on a 0.13mum CMOS FPGA (Altera Stratix EP1S40)
Scott Y. L. Chin, Clarence S. P. Lee, Steve Wilton
FPL3
2006 Activity Estimation for Field-Programmable Gate Arrays
abstract
This paper examines various activity estimation techniques in order to determine which are most appropriate for use in the context of field-programmable gate arrays (FPGAs). Specifically, the paper compares how different activity estimation techniques affect the accuracy of FPGA power models and the ability of power-aware FPGA CAD tools to minimize power. After comparing various existing techniques, the most suitable existing techniques are combined with two novel enhancements to create a new activity estimation tool called ACE-2.0. Finally, the new publicly available tool is compared to existing tools to validate the improvements. Using activities estimated by ACE-2.0, the power estimates and power savings were both within 1% of the results obtained using simulated activities
Julien Lamoureux, Steve Wilton
FPL2
2006 Architecture and CAD for FPGA Clock Networks
abstract
Scaling of process technologies and innovations in FPGA architecture and CAD are allowing increasingly more sophisticated applications to be implemented on FPGAs. These applications include system-level designs consisting of many subcomponents that use separate clocks. To support a wide range of applications with different clocking schemes, commercial devices have incorporated high-speed, low-skew clock networks that supply multiple clock signals to logic, memory, and arithmetic elements in the FPGA. The design of these programmable clock networks and of the CAD tools that support them is of utmost importance. Not only does the clock network itself dissipate a significant amount of power, since it connects to every latch on the FPGA and toggles every cycle. But, the design of clock network also affects how efficiently the rest of the circuit can be implemented, since it imposed constraints on the CAD tools that map applications onto the FPGA.
Julien Lamoureux, Steve Wilton
FPL2
2006 Activity-based power estimation and characterization of DSP and multiplier blocks in FPGAs
abstract
This paper describes an activity-based strategy for estimating the average power dissipation of hard DSP and multiplier blocks embedded in FPGAs. We identified two technical challenges in creating a tool flow to do this: (1) estimating the activity of all nodes in designs containing DSP blocks, and (2) estimating the average power dissipated within the DSP block quickly and accurately. In this paper, we compare several methods to address each of these two challenges. We conclude with a description of our complete power estimation flow
Nathalie Chan King Choy, Steve Wilton
FPT2
2006 System-on-Chip: Reuse and Integration
abstract
Over the past ten years, as integrated circuits became increasingly more complex and expensive, the industry began to embrace new design and reuse methodologies that are collectively referred to as system-on-chip (SoC) design. In this paper, we focus on the reuse and integration issues encountered in this paradigm shift. The reusable components, called intellectual property (IP) blocks or cores, are typically synthesizable register-transfer level (RTL) designs (often called soft cores) or layout level designs (often called hard cores). The concept of reuse can be carried out at the block, platform, or chip levels, and involves making the IP sufficiently general, configurable, or programmable, for use in a wide range of applications. The IP integration issues include connecting the computational units to the communication medium, which is moving from ad hoc bus-based approaches toward structured network-on-chip (NoC) architectures. Design-for-test methodologies are also described, along with verification issues that must be addressed when integrating reusable components.
Res Saleh, Steve Wilton, Shahriar Mirabbasi, Alan J. Hu, Mark R. Greenstreet, Guy Lemieux, Partha Pratim Pande, Cristian Grecu, André Ivanov
Proc. IEEE2
2006 Product-Term-Based Synthesizable Embedded Programmable Logic Cores
abstract
As integrated circuits become increasingly complex, the ability to make post-fabrication changes will become more important and attractive. This capability can be realized by using programmable logic cores. Currently, such cores are available from vendors in the form of "hard" macro layouts. Previous work has suggested an alternative approach: vendors supply a synthesizable version of their programmable logic core and the integrated circuit designer synthesizes the programmable logic fabric using standard cells. Although this technique suffers increased delay, area, and power, the task of integrating such cores is far easier than the task of integrating "hard" cores into an ASIC or system-on-chip (SoC). When implementing a small amount of logic, this ease of use may be more important than the increased overhead. This paper presents a new family of architectures for these "synthesizable" cores; unlike previous architectures, which were based on lookup-tables (LUTs), the new family of architectures is based on a collection of product-term arrays. Compared to LUT-based architectures, the new architectures result in density improvements of 35% and speed improvements of 72% on standard benchmark circuits. The improvement is due to the inherent efficiency of product-term-based designs for small logic circuits. In addition, we describe novel ways of enhancing synthesizable architectures to support sequential logic. We show that directly embedding flip-flops as is done in stand-alone programmable cores will not suffice. Consequently, we present two novel architectures employing our solution and optimize and compare them. Finally, we describe a proof-of-concept layout employing one of our proposed architectures.
Andy Yan, Steve Wilton
IEEE Trans. Very Large Scale Integr. Syst.2
2005 Register File Architecture Optimization in a Coarse-Grained Reconfigurable Architecture
abstract
This paper investigates the impact of the local and global register file architecture on a reconfigurable system based on the ADRES architecture. The register files consume a significant amount of area on the reconfigurable device, and their architecture has a strong impact on the performance. We found that the global registers should be tightly connected to as many functional units as possible, while the connection of the local register files to their neighbours is less critical. We found that the global register file should contain between 12 and 16 registers, while each local register file should only contain one or two registers. We used these results to propose a new architecture that has between 60% and 95% higher performance per unit area compared to the original architecture over the set of benchmarks.
Zion S. Kwok, Steve Wilton
FCCM2
2005 Dynamic Voltage Scaling for Commercial FPGAs
Gary Chun Tak Chow, L. S. M. Tsui, Philip H. W. Leong, Wayne Luk, Steve Wilton
FPT5
2005 Post-Silicon Debug Using Programmable Logic Cores
Bradley R. Quinton, Steve Wilton
FPT2
2005 Asynchronous IC Interconnect Network Design and Implementation Using a Standard ASIC Flow
abstract
The implementation of interconnect is becoming a significant challenge in modern IC design. Both synchronous and asynchronous strategies have been suggested to manage this problem. Creating a low skew clock tree for synchronous inter-block pipeline stages is a significant challenge. Asynchronous interconnect does not require a global clock, and therefore, it has a potential advantage in terms of design effort. This paper presents an asynchronous interconnect design that can be implemented using a standard ASIC flow. This design is considered in the context of a simple interconnect network. The results demonstrate that there is a region of the design space where the implementation provides an advantage over a synchronous interconnect. A detailed comparison of power, area and latency of the two strategies is also provided for a range of IC scenarios.
Bradley R. Quinton, Mark R. Greenstreet, Steve Wilton
ICCD3
2005 Challenges and opportunities for low power FPGAs in nanometer technologies
abstract
In this session, we will first present an overview of new challenges in commercial FPGA architecture design with an emphasis on the circuit and architecture issues for power at current and upcoming process nodes. Today’s 90nm FPGAs utilize techniques such as programmable shut-down of unused resources at the architectural level and multiple threshold voltages and gate-oxides at the circuit level. At 65nm and 45nm new techniques will need to target not only power mitigation but process variation in power and timing and their impact on yield and manufacturability.
Lei He 0001, Mike Hutton, Tim Tuan, Steve Wilton
ISLPED4
2005 A detailed power model for field-programmable gate arrays
abstract
Power has become a critical issue for field-programmable gate array (FPGA) vendors. Understanding the power dissipation within FPGAs is the first step in developing power-efficient architectures and computer-aided design (CAD) tools for FPGAs. This article describes a detailed and flexible power model which has been integrated in the widely used Versatile Place and Route (VPR) CAD tool. This power model estimates the dynamic, short-circuit, and leakage power consumed by FPGAs. It is the first flexible power model developed to evaluate architectural tradeoffs and the efficiency of power-aware CAD tools for a variety of FPGA architectures, and is freely available for noncommercial use. The model is flexible, in that it can estimate the power for a wide variety of FPGA architectures, and it is fast, in that it does not require extensive simulation, meaning it can be used to explore a large architectural space. We show how the model can be used to investigate the impact of various architectural parameters on the energy consumed by the FPGA, focusing on the segment length, switch block topology, lookuptable size, and cluster size.
Kara K. W. Poon, Steve Wilton, Andy Yan
ACM Trans. Design Autom. Electr. Syst.2
2005 Routing architecture optimizations for high-density embedded programmable IP cores
abstract
Programmable logic cores differ from stand-alone field-programmable gate arrays in that they can take on a variety of shapes and sizes. With this in mind, we investigate the detailed routing architecture of rectangular programmable logic cores. We quantify the effects of having different X and Y channel capacities and show that the optimum ratio between the X and Y channel widths for a rectangular core is between 1.2 and 1.5. We also present a new switch block family optimized for rectangular cores. Further, we quantify the effects of logic block pin placement. Compared with a simple extension of an existing switch block, our new architecture leads to a density improvement of up to 11.9%. Finally, we show that, if the channel width, switch block, and pin placement are chosen carefully, then the penalty for using a rectangular core (compared to a square core with the same logic capacity) is small; for a core with an aspect ratio of 2:1, the area penalty is 1.6% and the speed penalty is 3.8%.
Peter Hallschmid, Steve Wilton
IEEE Trans. Very Large Scale Integr. Syst.2
2005 A novel FPGA architecture supporting wide, shallow memories
abstract
This paper investigates an architecture designed to implement wide, shallow memories on a field programmable gate array (FPGA). In the proposed architecture, existing configuration memory normally used to control the connectivity pattern of the FPGA is made user accessible. Typically, not all the switch blocks in an FPGA are used to transport signals. By adding only a modest amount of circuitry, the configuration memory in these unused switch blocks (or unused paths within used switch blocks) can be used to implement wide, shallow buffers and other similar memory structures. The size of FPGA required to implement a benchmark circuit that makes use of the wide, shallow memories, is 20% smaller than a standard memory architecture. In addition, the benchmark circuit is on average 40% faster using the proposed architecture.
Steven W. Oldridge, Steve Wilton
IEEE Trans. Very Large Scale Integr. Syst.2
2004 The Impact of Pipelining on Energy per Operation in Field-Programmable Gate Arrays
Steve Wilton, Su-Shin Ang, Wayne Luk
FPL1
2004 Interconnect architectures for modulo-scheduled coarse-grained reconfigurable arrays
abstract
The ability of a compiler to exploit loop-level parallelism in a reconfigurable array is significantly affected by the amount of flexibility in the interconnect architecture. A less flexible interconnect will make it more difficult for the compiler to find efficient loop-level pipelined schedules, leading to reduced instruction throughput, and larger configuration bit storage area. In this paper, we determine the optimum flexibility and topology for a point-to-point interconnect architecture in a reconfigurable system. We present four topologies, and show that their performance per unit area is significantly better than that that would be obtained if a fully-connected network had been used.
Steve Wilton, Noha Kafafi, Bingfeng Mei, Serge Vernalde
FPT1
2004 Placement and routing for non-rectangular embedded programmable logic cores in SoC design
abstract
As SoC design enters into mainstream usage, the ability to make post-fabrication changes will become more and more attractive. This ability can be realized using programmable logic cores. These cores are like any other IP in the SoC design methodology, except that their function can be changed after fabrication. In many cases, non-rectangular programmable logic cores are required, either to better mesh with the other IP cores, or because of I/O constraints. In order to use programmable logic cores, placement and routing algorithms are required to implement user circuits on the core. Existing placement and routing algorithms that target programmable logic were optimized for stand-alone FPGAs which are invariably square or rectangular. We show that these algorithms do not work well when targetting non-rectangular programmable logic cores, and we present enhancements to existing placement and routing algorithms that allow the algorithms to better target these cores. It is shown that the new algorithms lead to a 12% critical path improvement for "U"-shaped cores, and a 4% improvement for "O"-shaped cores. The density and speed penalty for using these non-rectangular cores is significant, compared to square cores, however, we show that the penalty would be significantly larger if the original algorithms were used.
Tony Wong, Steve Wilton
FPT2
2003 Architectures and algorithms for synthesizable embedded programmable logic cores
abstract
As integrated circuits become more and more complex, the ability to make post-fabrication changes will become more and more attractive. This ability can be realized using programmable logic cores. Currently, such cores are available from vendors in the form of a "hard" layout. In this paper, we focus on an alternative approach: vendors supply a synthesizable version of their programmable logic core (a "soft" core) and the integrated circuit designer synthesizes the programmable logic fabric using standard cells. Although this technique suffers increased speed, density, and power overhead, the task of integrating such cores is far easier than the task of integrating "hard" cores into an ASIC. For very small amounts of logic, this ease of use may be more important than the increased overhead. This paper presents two synthesizable programmable logic core architectures, describes the associated place and route CAD tools, and compares the two architectures to each other, and to a "hard" programmable logic core. It also shows how these cores can be made more efficient by creating a non-rectangular architecture, an option not available to "hard" core vendors.
Noha Kafafi, Kimberly A. Bozman, Steve Wilton
FPGA3
2003 Placement and routing for FPGA architectures supporting wide shallow memories
abstract
Today, FPGAs are being used to implement large, system-sized circuits. Systems often require significant memory resources, and vendors have responded to these needs by embedding block memories onto their FPGAs. In we presented an architecture designed to efficiently support the need for wide shallow memories on an FPGA by allowing switch block configuration memory to be read and written by the user circuit. This FPGA presents a unique placement and routing problem, since the embedded memories displace routing resources in switch blocks. In this paper, we present novel place and route algorithms for an FPGA containing these wide shallow memories. Using these tools, a comparison between FPGAs containing switch block memories and those containing standard memory architectures shows that switch block memory based solutions are 22% smaller and 40% faster, despite their overhead.
Steven W. Oldridge, Steve Wilton
FPT2
2003 Product-term based synthesizable embedded programmable logic cores
abstract
As integrated circuits become increasingly complex, the ability to make post-fabrication changes will become more important and attractive. This ability can be realized using programmable logic cores. Currently, such cores are available from vendors in the form of a "hard" layout. Previous work has suggested an alternative approach: vendors supply a synthesizable version of their programmable logic core and the integrated circuit designer synthesizes the programmable logic fabric using standard cells. This paper presents a new family of architectures for these synthesizable cores; unlike previous architectures which were based on lookup-tables, the new family of architectures is based on a collection of product-term arrays. Compared to lookup-table based architectures, the new architectures result in density improvements of 35% and speed improvements of 72% on standard benchmark circuits.
Andy Yan, Steve Wilton
FPT2
2003 On the Interaction Between Power-Aware FPGA CAD Algorithms
Julien Lamoureux, Steve Wilton
ICCAD2
2002 On the sensitivity of FPGA architectural conclusions to experimental assumptions, tools, and techniques
abstract
Recent years have seen a tremendous increase in the capacities and capabilities of Field-Programmable Gate Arrays (FPGA's). Much of this dramatic improvement has been the result of changes to the FPGAs' internal architectures. New architectural proposals are routinely generated in both academia and industry. For FPGA's to continue to grow, it is important that these new architectural ideas are fairly and accurately evaluated, so that those worthy ideas can be included in future chips. Typically, this evaluation is done using experimentation. However, the use of experimentation is dangerous, since it requires making assumptions regarding the tools and architecture of the device in question. If these assumptions are not accurate, the conclusions from the experiments may not be meaningful. In this paper, we investigate the sensitivity of FPGA architectural conclusions to experimental variations. To make our study concrete, we evaluate the sensitivity of four previously published and well-known FPGA architectural results: lookup-table size, switch block topology, cluster size, and memory size. It is shown that these experiments are significantly affected by the assumptions, tools, and techniques used in the experiments.
Andy Yan, Rebecca Cheng, Steve Wilton
FPGA3
2002 A Flexible Power Model for FPGAs
Kara K. W. Poon, Andy Yan, Steve Wilton
FPL3
2002 Sensitivity of FPGA power evaluation
abstract
Power dissipation is becoming a major concern among FPGA vendors. Recently, architectural studies have been published which attempt to quantify the effects of various architectural alternatives on the power dissipation of FPGAs. These studies are very sensitive to assumptions made during the experimentation. In this paper, we analyze the sensitivity of two of these assumptions: the primary input density and the routing algorithm. We show that both of these assumptions significantly impact the architectural results.
Kara K. W. Poon, Steve Wilton
FPT2
2002 Implementing logic in FPGA memory arrays: heterogeneous memory architectures
abstract
It has become clear that large embedded configurable memory arrays will be essential in future FPGAs. Embedded arrays provide high-density high-speed implementations of the storage parts of circuits. Unfortunately, they require the FPGA vendor to partition the device into memory and logic resources at manufacture-time. This leads to a waste of chip area for customers that do not use all of the storage provided This chip area need not be wasted, and can in fact be used very efficiently, if the arrays are configured as large multi-output ROMs, and used to implement logic. In this paper we investigate how the architecture of the FPGA embedded arrays affects their ability to implement logic. Specifically, we focus on architectures which contain more than one size of memory array. We show that these heterogeneous architectures result in significantly denser implementations of logic than architectures with only one size of memory array. We also show that the best heterogeneous architecture contains both 2048 bit arrays and 128 bit arrays.
Steve Wilton
FPT1
2001 Detailed routing architectures for embedded programmable logic IP cores
abstract
As the complexity of integrated circuits increases, the ability to make post-fabrication changes to fixed ASIC chips will become more and more attractive. This ability can be realized using programmable logic cores. These cores are blocks of programmable logic that can be embedded into a fixed-function ASIC or a custom chip. Such cores differ from stand-alone FPGAs in that they can take on a variety of shapes and sizes. With this in mind, we investigate the detailed routing characteristics of rectangular programmable logic cores. We quantify the effects of having different x and y channel capacities, and show that the optimum ratio between the x and y channel widths for a rectangular core is between 1.2 and 1.5. We also present a new switch block family optimized for rectangular cores. Compared to a simple extension of an existing switch block, our new architecture leads to an 8.7% improvement in density with little effect on speed. Finally, we show that if the channel widths and switch block are chosen carefully the penalty for using a rectangular core (compared to a square core with the same logic capacity) is small; for a core with an aspect ratio of 2:1, the area penalty is 1.6% and the speed penalty is 1.1%.
Peter Hallschmid, Steve Wilton
FPGA2
2001 A crosstalk-aware timing-driven router for FPGAs
abstract
As integrated circuits are migrated to more advanced technologies, it has become clear that crosstalk is an important physical phenomenon that must be taken into account. Crosstalk has primarily been a concern for ASICs, multi-chip modules, and custom chips, however, it will soon become a concern in FPGAs. In this paper, we describe the first published crosstalk-aware router that targets FPGAs. We show that, in a representative FPGA architecture implemented in a 0.18mm technology, the average routing delay in the presence of crosstalk can be reduced by 7.1% compared to a router with no knowledge of crosstalk. About half of this improvement is due to a tighter delay estimator, and half is due to an improved routing algorithm.
Steve Wilton
FPGA1
2001 Macrocell Architectures for Product Term Embedded Memory Arrays
Ernie Lin, Steve Wilton
FPL2
2001 Structural analysis and generation of synthetic digital circuits with memory
abstract
One of the most difficult aspects of experimental reconfigurable architecture or computer-aided design (CAD) tool research is obtaining sufficiently large benchmark circuits. One approach to obtaining such circuits is to generate them stochastically. Current circuit generators construct combinational and sequential logic circuits. Many of today's devices, however, are being used to implement entire systems, and often these systems contain on-chip storage. This paper describes a circuit generator that constructs circuits containing significant amounts of memory. To ensure the circuits are realistic, we have performed a detailed structural analysis of such circuits; this analysis is also described in this paper.
Steve Wilton, Jonathan Rose, Zvonko G. Vranesic
IEEE Trans. Very Large Scale Integr. Syst.1
2000 Heterogeneous technology mapping for FPGAs with dual-port embedded memory arrays
abstract
It has become clear that on-chip storage is an essential component of high-density FPGAs. These arrays were originally intended to implement storage, but recent work has shown that they can also be used to implement logic very efficiently. This previous work has only considered single-port arrays. Many current FPGAs, however, contain dual-port arrays. In this paper we present an algorithm that maps logic to these dual-port arrays. Our algorithm can either optimize area with no regard for circuit speed, or optimize area under the constraint that the combinational depth of the circuit does not increase. Experimental results show that, on average, our algorithm packs between 29% and 35% more logic than an algorithm that targets single-port arrays. We also show, however, that even with this algorithm, dual-port arrays are still not as area-efficient as single-port arrays when implementing logic.
Steve Wilton
FPGA1
2000 Heterogeneous technology mapping for area reduction in FPGAs withembedded memory arrays
abstract
It has become clear that large embedded configurable memory arrays will be essential in future field programmable gate arrays (FPGAs). Embedded arrays provide high-density high-speed implementations of the storage parts of circuits, Unfortunately, they require the FPGA vendor to partition the device into memory and logic resources at manufacture-time. This leads to a waste of chip area for customers that do not use all of the storage provided. This chip area need not be wasted, and can in fact be used very efficiently, if the arrays are configured as multioutput ROMs, and used to implement logic, In this paper, we describe two versions of a new technology mapping algorithm that identifies parts of circuits that can be efficiently mapped to an embedded array and performs this mapping, The first version of the algorithm places no constraints on the depth of the final circuit; on a set of 29 sequential and combinational benchmarks, the tool is able to map, on average, 59.7 4-LUTs into a single 2-Kbit memory array, while increasing the critical path by 7%, The second version of the algorithm places a constraint on the depth of the final circuit; it maps, on average, 56.7 4-LUTs into the same memory array, while increasing the critical path by only 2.3%. This paper also considers the effect of the memory array architecture on the ability of the algorithm to pack logic into memory, It is shown that the algorithm performs best when each array has between 512 and 2048 bits, and has a word width that can be configured as 1, 2, 4, or 8.
Steve Wilton
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1999 The memory/logic interface in FPGAs with large embedded memory arrays
abstract
As the capacities of field-programmable gate arrays (FPGAs) grow, they will be used to implement much larger circuits than ever before. These larger circuits often require significant amounts of storage. In order to address these storage requirements, FPGAs with large embedded memory arrays are now being developed by several vendors. One of the crucial components of an FPGA with on-chip memory is the routing structure between the memory arrays and logic resources. If this memory/logic interface is not flexible enough, many circuits will be unroutable, while if it is too flexible, it will be slower and consume more chip area than is necessary. In this paper, we show that an interconnect in which each memory pin can connect to between four and seven logic routing tracks is best in terms of both area and speed. We also show that by adding switches to support nets that connect multiple memory arrays, we can reduce the memory access time by up to 25% and improve the routability slightly.
Steve Wilton, Jonathan Rose, Zvonko G. Vranesic
IEEE Trans. Very Large Scale Integr. Syst.1
1998 SMAP: Heterogeneous Technology Mapping for Area Reduction in FPGAs with Embedded Memory Arrays
abstract
It has become clear that large embedded configurable memory arrays will be essential in future FPGAs. Embedded arrays provide high-density high-speed implementations of the storage parts of circuits. Unfortunately, they require the FPGA vendor to partition the device into memory and logic resources at manufacture-time. This leads to a waste of chip area for customers that do not use all of the storage provided. This chip area need not be wasted, and can in fact be used very efficiently, if the arrays are configured as large multi-output ROMs, and used to implement logic.
Steve Wilton
FPGA1
1997 Memory-to-Memory Connection Structures in FPGAs with Embedded Memory Arrays
abstract
This paper shows that the speed of FPGAs with large embedded memory arrays can be improved by adding direct programmable connections between the memories. Nets that connect to multiple memory arrays are often difficult to route, and are often part of the critical path of circuit implementations. The memory-to-memory connection structure proposed in this paper allows for the efficient implementation of these nets, resulting in a reduction in memory access time of up to 25% and a slight improvement in routability. 1 Introduction As FPGAs become larger, they will be used to implement entire systems, rather than small logic subcircuits. One of the key differences between these large systems and the smaller logic subcircuits is that the systems often contain memory. Architectural support for the efficient implementation of memory in nextgeneration FPGAs, therefore, is crucial. Several vendors offer FPGAs with architectural support for memory [1, 2, 3, 4, 5, 6, 7]. The memory resources in ...
Steve Wilton, Jonathan Rose, Zvonko G. Vranesic
FPGA1
1995 Architecture of Centralized Field-Configurable Memory
abstract
As the capacities of FPGAs grow, it becomes feasible to implement the memory portions of systems directly on an FPGA together with logic. We believe that such an FPGA must contain specialized architectural support in order to implement memories efficiently. The key feature of such architectural support is that it must be flexible enough to accommodate many different memory shapes (widths and depths) as well as allowing different numbers of independently-addressed memory blocks. This paper describes a family of centralized Field-Configurable Memory architectures which consist of a number of memory arrays and dedicated mapping blocks to combine these arrays. We also present a method for comparing these architectures, and use this method to examine the tradeoffs involved in choosing the array size and mapping block capabilities.
Steve Wilton, Jonathan Rose, Zvonko G. Vranesic
FPGA1
1994 Tradeoffs in Two-Level On-Chip Caching
abstract
The performance of two-level on-chip caching is investigated for a range of technology and architecture assumptions. The area and access time of each level of cache is modeled in detail. The results indicate that for most workloads, two-level cache configurations (with a set-associative second level) perform marginally better than single-level cache configurations that require the same chip area once the first-level cache sizes are 64 KB or larger. Two-level configurations become even more important in systems with no off-chip cache and in systems in which the memory cells in the first-level caches are multiported and hence larger than those in the second-level cache. Finally, a new replacement policy called two-level exclusive caching is introduced. Two-level exclusive caching improves the performance of two-level caching organizations by increasing the effective associativity and capacity.>
Norman P. Jouppi, Steve Wilton
ISCA2
1993 A VLSI Implementation of a Cascade Viterbi Decoder with Traceback
Gennady Feygin, Paul Chow, P. Glenn Gulak, John Chappel, Grant Goodes, Oswin Hall, Ahmad Sayes, Satwant Singh, Michael B. Smith, Steve Wilton
ISCAS10