VLDB 2026 Research / reviewers in the wild / expert
Jason Helge Anderson
dblp:46/4753 · also Jason Anderson 0001
· DBLP profile ↗
122ranked-venue papers
21as first author
23since 2021 · last 2026
0000-0001-9083-6853ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 120 · 21 first-author · 23 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ARCS: Architecture-Responsive CGRA SchedulingabstractScheduling is a key aspect in mapping applications to coarse-grained reconfigurable architectures (CGRA). During scheduling, the number of pipeline registers on each path is determined to ensure that the data for an operation arrives at the correct cycle. Traditional scheduling methods, such as as soon as possible (ASAP) and as late as possible (ALAP), determine the pipeline registers without considering the placement of operations. This can restrict the mapping algorithm, forcing placement to accommodate scheduling and limiting routing options to satisfy both scheduling and placement. To overcome the limitations of traditional schedulers, we propose a method that adaptively adjusts the schedule during the initial routing phase. In this approach, the mapping algorithm begins with an ASAP schedule to establish initial schedule constraints. We then utilize simulated annealing for placement, and we employ a architecture-responsive scheduling algorithm post-placement to update the schedule of each edge based on the placement and generate an initial routing solution with overlaps. Afterwards, the PathFinder algorithm is applied to the initial routing solution, along with the generated schedule, to find a valid routing with no overlaps. Our results demonstrate that the architecture-responsive scheduling approach maintains a quality comparable to that of conventional ASAP scheduling. Furthermore, architecture-responsive scheduling enables generic mapping of applications onto restricted architectures that do not allow routes to bypass pipeline registers, a challenge that traditional schedulers do not address. Omar Ragheb, Jason Helge Anderson |
ASP-DAC | 2 |
| 2026 | CUT-MC: Optimizing the Relationship Between Context Count, Unrolling and Throughput in Multi-Context CGRAsabstractWe propose techniques to raise throughput for compute-intensive loops when implemented in multi-context CGRAs via judicious choices for context count and loop-unroll factor. We show that using a non-minimal initiation interval (II) can yield higher throughput in many applications (vs. using the minimum II), provided there is sufficient loop unrolling. Average throughput improvements of 36% are achieved across a range of applications, CGRA sizes and context counts. Stephen Wicklund, Jason Helge Anderson |
DATE | 2 |
| 2025 | FRIDA: Reconfigurable Arrays for Dynamically Scheduled High-Level SynthesisabstractReconfigurable computing fabrics include FPGAs and CGRAs. FPGAs offer flexible bit-level reconfigurability and can map almost any program via high-level synthesis (HLS) compilers, but they incur high area and speed overheads compared to ASICs. CGRAs, in contrast, provide ASIC-like performance but limited flexibility, typically supporting only feedforward programs with unambiguous memory accesses, far from the capabilities of HLS compilers. This work introduces a new class of reconfigurable arrays inspired by modern dynamically scheduled HLS (DHLS) tools. Unlike traditional HLS, DHLS compilers no longer produce explicit state machines, eliminating the need for look-up tables. Instead, they delegate scheduling decisions to a set of coarse-grained primitives. Our arrays leverage these primitives as processing elements and combine FPGA-style interconnect topology for high routing flexibility with CGRA-like bus-based interconnect. We present a framework to explore these arrays and evaluate a preliminary architecture using DHLS benchmarks. The results show an average of ~2× speed improvement, but unfortunately only a ~20% area reduction compared to an FPGA implemented on the same technology node. Louis Coulon, Lucas Ramirez, Jason Helge Anderson, Mirjana Stojilovic, Paolo Ienne |
FPGA | 3 |
| 2025 | FRESCO: Efficient Subgraph Enumeration for Scalable Clustering in Heterogeneous CGRAsabstractIn recent years, there has been a trend towards reconfigurable fabrics at the intersection between field-programmable gate arrays (FPGAs) and coarse-grained reconfigurable arrays (CGRAs): using FPGA-like interconnect but word-based and built around coarse-grained primitives. These architectures often employ complex clusters with far more heterogeneous resources than FPGAs or typical CGRAs— sometimes over a hundred primitives, most of which are bypassable. As a result, clustering, the problem of covering the application netlist with architecture clusters, is a key challenge for design tools targeting these fabrics. Clustering is analogous to the instruction selection problem in CISC architectures, albeit with orders of magnitude more complex "instructions". In this work, we propose a two-phase, architecture-agnostic clustering algorithm that scales to highly complex architecture clusters. The first phase enumerates potential cluster matches in the application netlist using a strategy based on an abstract decision tree. The second phase selects a cover from the enumerated matches. We show that our algorithm effectively prunes the search space for complex clusters, scales well to circuits composed of many clusters, and achieves better clustering quality than CLUMAP, a state-of-the-art CGRA clustering algorithm, for the simple cases that CLUMAP can handle. Louis Coulon, Adham Ragab, Jason Helge Anderson, Mirjana Stojilovic, Paolo Ienne |
ICCAD | 3 |
| 2024 | Raising Compute Density of Molecular Dynamics Simulation Through Approximate MemoizationabstractMolecular dynamics (MD) simulation involves simulating the interactions of particles. MD has many applications in basic biological sciences, drug discovery, materials science, and other fields. Simulating$1\mu \mathrm{s}$of a 100K-atom system can take hours or days11https://www.bdr.riken.jp/en/research/labs/taiji-mlmdgrape4.html, where the compute-heavy aspect of MD is calculating the long-range forces between pairs of particles. In this paper, we explore the application of approximate computing in MD as a means to improve compute density. Specifically, we employ approximate memoization, where previously computed forces (and more) are stored in a table, and are retrieved in subsequent force calculations, provided the inputs to the force calculation are the same or similar. If the prior-computed table values can be used, significant computational work is avoided. In an experimental study, we apply software simulation to understand the degree to which approximation is feasible. We then propose a hardware implementation of memoization to be used within an ASIC MD simulator, MDGRAPE-4A [18]. We show that compute density, measured as$\text{pair-interactions}/(s\cdot\mu m^{2})$is improved substantially, between 40 % and 70 % for the studied cases. This is contingent on the particular system being simulated, the table size, and permitted level of approximation. Salim Khemira, Xinyuan Wang 0003, Yutaka Tamiya, Makoto Taiji, Takahide Yoshikawa, Jason Helge Anderson |
ASAP | 7 |
| 2024 | MLIR-to-CGRA: A Versatile MLIR-Based Compiler Framework for CGRAsabstractCoarse-grained reconfigurable architectures (CGRAs) are programmable hardware platforms with coarse-grained programmable logic blocks and word-wide configurable interconnect. In this paper, we describe a high-level compilation framework tailored for CGRAs. The input to the framework is a C-language program and description of the available CG RA architectural features. This work primarily focuses on automatically handling sequential rectangular loop access patterns. It achieves this by running automated design space exploration (DSE) written within the MLIR compiler representation [7] to leverage both spatial and temporal parallelism inherent in CG RAs. Furthermore, the compiler is engineered to be architecture-agnostic. It employs standard loop (and other) compiler optimization passes alongside CG RA-specific passes to generate a data flow graph (DFG) tailored for a CG RA implementation. In an experimental study, we optimize kernels to increase CG RA mappability and show the impact of automated DSE on the kernel performance. Moreover, the framework eases the task of compilation from software to RISC-V+CGRA hybrid systems. Omar Ragheb, Stephen Wicklund, Jason Helge Anderson |
ASAP | 4 |
| 2024 | CLUMAP: Clustered Mapper for CGRAs with PredicationabstractCoarse-grained reconfigurable architectures (CGRAs) have gained popularity as accelerators for compute-intensive kernels. Complex CGRA architectures that support key features such as multi-context and predication are being developed to support a wider range of kernels. However, mapping applications on these complex architectures poses significant challenges. In this paper, we provide an architecture-agnostic clustered mapping technique and a new cost function tailored for simulated-annealing placement. The mapper simplifies placement and routing phases, demonstrating significant speedup for popular CGRA architectures: HyCUBE and ADRES. Additionally, our method demonstrates an increase in mapping success for the ADRES architecture. Omar Ragheb, Jason Helge Anderson |
DAC | 2 |
| 2024 | Mapping Enumeration for Multi-Context CGRAs Using Zero-Suppressed Binary Decision DiagramsabstractA primary aim of Coarse-Grained Reconfigurable Arrays (CGRAs), compared to FPGAs, is to maximize the portion of the die used for computational resources, while minimizing the complexity of control and steering logic, leading to inherently constrained routing architectures. This challenge has compelled CAD developers to utilize exact solutions, such as integer linear programming (ILP), in formulating and solving the mapping problem. Those solutions have been shown not to scale, especially for larger devices with intricate architectural features, such as multiple contexts and optional pipeline registers. Even if an exact or a greedy approach yields a feasible solution, it often fails to optimize multifaceted objective criteria. In this work, we have devised a framework for systematically enumerating mapping solutions of a subject kernel on a target CGRA using Zero- Suppressed Binary Decision Diagrams (ZDDs). To effectively manage runtime, we developed a linear algorithm that retains the best$k$solutions at each stage of the mapping flow, where both the objective function and$k$are user defined. Experimental results on a diverse range of application kernels targeting two CGRA architectures show how we can enumerate hundreds of thousands of solutions within seconds. When compared against prior methodologies, and while generating dozens of solutions, our mapper exhibits a remarkable speed advantage, ranging from one to three orders of magnitude faster than exact and heuristic approaches. Notably, when allocated the same runtime as the fastest heuristic, our framework demonstrates its efficacy by generating an impressive 105 solutions. Rami Beidas, Jason Helge Anderson |
FCCM | 2 |
| 2024 | Exploration of Trade-offs Between General-Purpose and Specialized Processing Elements in HPC-Oriented CGRAabstractCoarse-Grained Reconfigurable Arrays (CGRAs) are a class of reconfigurable accelerators traditionally used in embedded computing. Recently, CGRA-like devices have gained traction for HPC and AI acceleration; however, typical HPC and AI workloads often require operations that current CGRAs cannot implement, such as complex mathematical calculations. In this work, we present a broad architectural study exploring potential heterogeneous computational resources in CGRA architectures for HPC, which are not commonly considered in typical CGRA architecture research. We first improved the general-purpose Processing Element (PE) of a baseline CGRA to optimize computational resources and then developed a new specialized PE for mathematical functions commonly found in HPC applications. Finally, we evaluated multiple CGRA configurations concerning floorplan, size, general-purpose/specialized PE ratio, and Power, Performance, and Area (PPA) results from hardware synthesis. Emanuele Del Sozzo, Xinyuan Wang 0003, Boma Anantasatya Adhi, Carlos Cortes, Jason Helge Anderson, Kentaro Sano |
IPDPS | 5 |
| 2023 | Area-Driven FPGA Logic Synthesis Using Reinforcement LearningabstractLogic synthesis involves a rich set of optimization algorithms applied in a specific sequence to a circuit netlist prior to technology mapping. A conventional approach is to apply a fixed "recipe" of such algorithms deemed to work well for a wide range of different circuits. We apply reinforcement learning (RL) to determine a unique recipe of algorithms for each circuit. Feature-importance analysis is conducted using a random-forest classifier to prune the set of features visible to the RL agent. We demonstrate conclusive learning by the RL agent and show significant FPGA area reductions vs. the conventional approach (resyn2). In addition to circuit-by-circuit training and inference, we also train an RL agent on multiple circuits, and then apply the agent to optimize: 1) the same set of circuits on which it was trained, and 2) an alternative set of "unseen" circuits. In both scenarios, we observe that the RL agent produces higher-quality implementations than the conventional approach. This shows that the RL agent is able to generalize, and perform beneficial logic synthesis optimizations across a variety of circuits. Guanglei Zhou, Jason Helge Anderson |
ASP-DAC | 2 |
| 2023 | GRAMM: Fast CGRA Application Mapping Based on A Heuristic for Finding Graph MinorsabstractA graph$H$is a minor of a second graph$G$if$G$can be transformed into$H$by two operations: 1) deleting nodes and/or edges, or 2) contracting edges. Coarse-grained reconfigurable array (CGRA) application mapping is closely related to the graph minor problem, where$H$is the application's dataflow graph and$G$is the CGRA's device-model graph. A heuristic algorithm to find graph minors has proven to be practical for sparse graphs with hundreds of vertices in a quantum computing application. In this work, we adapt the heuristic to CGRA application mapping, where the graphs have directed edges, and the vertices have unique types (e.g., representing ALUs or interconnect). Additionally, we alter the original cost function, taking inspiration from PathFinder, an iterative negotiated-congestion routing algorithm. In an experimental study comparing with a CGRA mapper based on integer linear programming, we demonstrate a higher rate of successful mappings and from 80× to up to orders of magnitude lower runtime. Guanglei Zhou, Mirjana Stojilovic, Jason Helge Anderson |
FPL | 3 |
| 2022 | CGRA Mapping Using Zero-Suppressed Binary Decision DiagramsabstractThe restricted routing networks of coarse-grained reconfigurable arrays (CGRAs) have motivated CAD developers to utilize exact solutions, such as integer linear programming (ILP), in formu-lating and solving the mapping problem. Such so-lutions that rely on general purpose optimizers have not been shown to scale. In this work, we formu-late CGRA mapping as a solution enumeration and selection problem, relying on the efficiency of zero-suppressed binary decision diagrams (ZDDs) [22] to capture the solution space. For small-to-moderate size problems, it is possible to capture every possible map-ping in a few megabytes. For larger problems, thou-sands if not millions of solutions can be enumerated. The final mapping is a simple linear-time DAG traver-sal of the enumeration ZDD. The proposed solution was implemented in the CGRA-ME [6] framework. A speedup of two orders of magnitude was obtained when compared with past solutions targeting smaller CGRA devices. Larger devices beyond the capacity of those solutions are now accessible. Rami Beidas, Jason Helge Anderson |
ASP-DAC | 2 |
| 2022 | Streaming Accuracy: Characterizing Early Termination in Stochastic ComputingabstractStochastic computing has garnered interest in the research community due to its ability to implement complicated compute with very small area footprints, at the cost of some accuracy and higher latency. With its unique tradeoffs between area, accuracy and latency, one commonly used technique to minimize area and latency is to early-terminate computation. Given this, it is useful to be able to measure and characterize how amenable a bitstream is to early termination. We present Streaming Accuracy, a metric that measures how far a bitstream is from its most early-terminable form. We show that it overcomes limitations of prior studies, and we characterize the design space for building stochastic circuits with early termination. We then propose a new hardware bitstream generator that produces bitstreams with optimal streaming accuracy. Hsuan Hsiao, Joshua San Miguel, Jason Helge Anderson |
ASP-DAC | 3 |
| 2022 | Modeling and Exploration of Elastic CGRAsabstractElastic design concepts have the potential to bring multiple benefits to coarse-grained reconfigurable arrays (CGRAs) architecture, including the ability to interface with memories, having unknown latencies, incorporate run-time variable-latency processing elements, and ease the CGRA mapping challenges of scheduling, placement and routing. However, there are overheads in terms of power, performance and area (PPA) associated with the design and implementation of elastic circuits. In this paper, we quantify these overheads in the CGRA context by first extending an open-source CGRA modelling and exploration framework (CGRA-ME) [4] to allow elastic circuit primitives (e.g. fork, join, merge, diverge, etc.) to be used when composing/modelling a CGRA architecture. We then use this new capability to “elasticize” two widely studied CGRA architectures, ADRES [11] and HyCUBE [8]. The PPA of the elastic versions of the CGRAs are compared with their traditional statically scheduled counterparts. We also evaluate the PPA “cost” of several elastic-circuit design points, such as elastic buffer length and inclusion of merge and diverge components. Omar Ragheb, David Ma, Jason Helge Anderson |
FPL | 4 |
| 2022 | Efficient Memory Arbitration in High-Level Synthesis From Multi-Threaded CodeabstractHigh-level synthesis (HLS) is an increasingly popular method for generating hardware from a description written in a software language like C/C++. Traditionally, HLS tools have operated on sequential code, however in recent years there has been a drive to synthesise multi-threaded code. In this context, a major challenge facing HLS tools is how to automatically partition memory among parallel threads to fully exploit the bandwidth available on an FPGA device and minimise memory contention. Existing partitioning approaches require inefficient arbitration circuitry to serialise accesses to each bank because they make conservative assumptions about which threads might access which memory banks. In this article, we design a static analysis that can prove certain memory banks are only accessed by certain threads, and use this analysis to simplify or even remove the arbiters while preserving correctness. We show how this analysis can be implemented using the Microsoft Boogie verifier on top of satisfiability modulo theories (SMT) solver, and propose a tool named EASY using automatic formal verification. Our work supports arbitrary input code with any irregular memory access patterns and indirect array addressing forms. We implement our approach in LLVM and integrate it into the LegUp HLS tool. For a set of typical application benchmarks our results have shown that EASY can achieve 0.13× (avg. 0.43×) of area and 1.64× (avg. 1.28×) of performance compared to the baseline, with little additional compilation time relative to the long time in hardware synthesis. Jianyi Cheng, Shane T. Fleming, Yu Ting Chen, Jason Helge Anderson, John Wickerson, George A. Constantinides |
IEEE Trans. Computers | 4 |
| 2021 | CGRA-ME: An Open-Source Framework for CGRA Architecture and CAD Research : (Invited Paper)abstractCoarse-grained reconfigurable arrays (CGRAs) are programmable hardware platforms that can be used to realize application-specific accelerators for higher performance and energy efficiency. A CGRA is a 2D array of configurable logic blocks & interconnect, where the logic blocks are typically large & ALU-like, and the interconnect is word-wide. CGRA-ME is a software framework that enables the modelling and exploration of CGRA architectures, as well as research on CGRA CAD algorithms. With CGRA-ME, an architect can specify a CGRA architecture at a high level of abstraction. A set of applications can be mapped onto the architecture to assess the mappability, power, performance and cost. CGRA-ME also allows one to generate synthesizable Verilog RTL for the modelled CGRA, permitting its implementation as an ASIC or FPGA overlay. In this paper, we describe the CGRA-ME framework [5] and overview its capabilities and current limitations. We discuss ongoing and prior research conducted with the framework, as well as outline future plans. We believe CGRA-ME will be a valuable contribution to the community, enabling new research on CGRA CAD & architectures. Jason Helge Anderson, Rami Beidas, Vimal Chacko, Hsuan Hsiao, Xiaoyi Ling, Omar Ragheb, Xinyuan Wang 0003 |
ASAP | 1 |
| 2021 | Power, Performance and Area Consequences of Multi-Context Support in CGRAsabstractA feature associated with coarse-grained reconfigurable arrays (CGRAs) is dynamic reconfigurability, wherein the CGRA supports multiple contexts. The multiple contexts form a set of configuration bitstreams that are loaded into the CGRA simultaneously and cycled through according to a schedule. Multi-context allows the CGRA hardware to be time-multiplexed: the logic blocks and interconnect can perform different functions according to the context selected in a given clock cycle. We consider how multi-context may be implemented at the circuit level, and evaluate three circuit implementations from the power, performance and area (PPA) perspectives. Results show that the choice of multi-context circuit implementation has an appreciable impact on the overall CGRA PPA. We also quantify the PPA overhead of the multi-context feature in CGRAs vs. a single-context device. Vimal Chacko, Jason Helge Anderson |
ASAP | 2 |
| 2021 | Double-Pumping the Interconnect for Area Reduction in Coarse-Grained Reconfigurable ArraysabstractWe consider double-pumped interconnect as a means of area reduction in coarse-grained reconfigurable arrays (CGRAs). Interconnect multiplexers comprise a considerable portion of CGRA area. We apply double-pumping to halve the word-width of the interconnect multiplexers, saving area. The interconnect is operated at twice the system clock frequency, where the top and bottom half-words of a value are communicated in the first and second half of a clock cycle. Several circuit-level approaches for double-pumping are considered, and evaluated in different CGRA architectures with varied interconnect richness. Area and performance consequences are assessed through a 45nm standard-cell ASIC implementation. Overall CGRA area improvements of up to 16% are observed, depending on the CGRA architecture and double-pumping implementation. Xinyuan Wang 0003, Hsuan Hsiao, Jason Helge Anderson |
ASAP | 4 |
| 2021 | Zero Correlation Error: A Metric for Finite-Length Bitstream Independence in Stochastic ComputingabstractStochastic computing (SC), with its probabilistic data representation format, has sparked renewed interest due to its ability to use very simple circuits to implement complex operations. Though unlike traditional binary computing, SC needs to carefully handle correlations that exist across data values to avoid the risk of unacceptably inaccurate results. With many SC circuits designed to operate under the assumption that input values are independent, it is important to provide the ability to accurately measure and characterize independence of SC bitstreams. We propose zero correlation error (ZCE), a metric that quantifies how independent two finite-length bitstreams are, and show that it addresses fundamental limitations in metrics currently used by the SC community. Through evaluation at both the functional unit level and application level, we demonstrate how ZCE can be an effective tool for analyzing SC bitstreams, simulating circuits and design space exploration. Hsuan Hsiao, Joshua San Miguel, Yuko Hara-Azumi, Jason Helge Anderson |
ASP-DAC | 4 |
| 2021 | High-Level Synthesis of Transactional Memory
Omar Ragheb, Jason Helge Anderson |
ASP-DAC | 2 |
| 2021 | An Open-Source Framework for the Generation of RISC-V Processor + CGRA Accelerator SystemsabstractWe describe a framework for automated generation of hybrid processor/accelerator systems comprising a RISC-V processor, and a coarse-grained reconfigurable array (CGRA) for realizing compute-kernel acceleration. CGRAs are programmable hardware platforms having an array of coarse ALU-like processing elements, and word-wide programmable interconnect. The proposed framework integrates CGRAs generated by the open-source CGRA-ME tool [1], with the RISC-V processor from the PULP project [2]. In an experimental study, we use the framework to generate RISC-V+CGRA systems that provide an order-of-magnitude speedup vs. software and considerable speedup vs. a vector processor on several applications by leveraging the CGRA spatial and pipeline parallelism. As CGRA-ME permits a variety of different CGRAs to be modelled and mapped to, we believe the proposed framework represents a powerful open-source platform, enabling a variety of new research on processor/CGRA system architectures and programming models. Xiaoyi Ling, Takahiro Notsu, Jason Helge Anderson |
DSD | 3 |
| 2021 | Post-LUT-Mapping Implementation of General Logic on Carry Chains Via a MIG-Based Circuit RepresentationabstractCarry chains on FPGAs have traditionally been only used for fast binary arithmetic operations. In this paper, we propose using the carry chain to implement general logic as a means of reducing the critical path delay and raising performance. To achieve this, we use a Majority-Inverter Graph (MIG) to represent the application during technology mapping, since carry functionality directly maps to the majority logic function. This aligns the subject graph of technology mapping with the capabilities of the carry chain. We first map an application to LUTs, then determine a chain of critical LUTs containing paths of majority “gates” that we deem beneficial for mapping onto the carry chain. We place such paths onto the carry chains, with the remaining logic in LUTs. In an experimental study using a suite of benchmarks, we observe that the proposed approach yields a post-place-and-route critical path delay that is superior to using delay-optimized mapping, yet without the significant area penalty. With carry-chain optimizations, area-delay product is improved by 9% vs. baseline LUT mappings. Jin Hee Kim, Jason Helge Anderson |
FPL | 2 |
| 2021 | Profiling-Based Control-Flow Reduction in High-Level SynthesisabstractControl flow in a program can be represented in a directed graph, called the control flow graph (CFG). Nodes in the graph represent straight-line segments of code, basic blocks, and directed edges between nodes correspond to transfers of control. We present a methodology to selectively reduce control flow by collapsing basic blocks into their parent blocks, revealing increased instruction-level parallelism to a high-level synthesis (HLS) scheduler, thereby raising circuit performance.We evaluate our approach within an HLS tool that allows a C-language software program to be automatically synthesized into a hardware circuit, using the CHStone benchmark suite [1], targeting an Intel Cyclone V FPGA. For individual benchmark circuits we observe cycle count reductions up to 20.7% and wall-clock time reductions up to 22.6%, and 6% on average. Austin Liolli, Omar Ragheb, Jason Helge Anderson |
FPT | 3 |
| 2020 | Optimizing FPGA Logic Block Architectures for ArithmeticabstractHardened adder and carry logic is widely used in commercial field-programmable gate arrays (FPGAs) to improve the efficiency of arithmetic functions. There are many design choices and complexities associated with such hardening, including circuit design, FPGA architectural choices, and the computer-aided design (CAD) flow. However, these choices have not been studied much and hence we explore a number of possibilities. We also highlight front-end elaboration optimization that helps ameliorate the restrictions placed on logic synthesis by hardened arithmetic. We show that hard adders and carry chains increase the performance of simple adders by a factor of 4 or more, but on larger benchmark designs that contain arithmetic improve the overall performance by 15%. Our results also show that for complete application circuits simple hardened ripple-carry adders perform as well as more complex carry-lookahead adders. Our best non-fracturable lookup table (non-fLUT) architecture with hardened arithmetic yields 12% better area-delay product than architectures without hardened arithmetic. We also investigate the impact of fLUTs and their interaction with hardened arithmetic. We find that fLUTs offer significant (12%-15%) area reduction, which is complementary to the delay reduction of hardened arithmetic. Therefore, our best fLUT architectures which use two bits of hardened arithmetic achieve 25% better area-delay product than non-fLUT architectures without hardened arithmetic. Kevin E. Murray, Jason Luu, Matthew J. P. Walker, Conor McCullough, Safeen Huda, Charles Chiasson, Kenneth B. Kent, Jason Helge Anderson, Jonathan Rose, Vaughn Betz |
IEEE Trans. Very Large Scale Integr. Syst. | 10 |
| 2019 | XOMA: exclusive on-chip memory architecture for energy-efficient deep learning accelerationabstractState-of-the-art deep neural networks (DNNs) require hundreds of millions of multiply-accumulate (MAC) computations to perform inference, e.g. in image-recognition tasks. To improve the performance and energy efficiency, deep learning accelerators have been proposed, realized both on FPGAs and as custom ASICs. Generally, such accelerators comprise many parallel processing elements, capable of executing large numbers of concurrent MAC operations. From the energy perspective, however, most consumption arises due to memory accesses, both to off-chip external memory, and on-chip buffers. In this paper, we propose an on-chip DNN co-processor architecture where minimizing memory accesses is the primary design objective. To the maximum possible extent, off-chip memory accesses are eliminated, providing lowest-possible energy consumption for inference. Compared to a state-of-the-art ASIC, our architecture requires 36% fewer external memory accesses and 53% less energy consumption for low-latency image classification. Hyeon Uk Sim, Jason Helge Anderson, Jongeun Lee |
ASP-DAC | 2 |
| 2019 | Thread Weaving: Static Resource Scheduling for Multithreaded High-Level SynthesisabstractIn high-level synthesis (HLS), software multithreading constructs can be used to explicitly specify coarse-grained parallelism for multiple accelerators. While software threads typically operate independently and in isolation of each other on CPUs, HLS threads/accelerators are sub-components of one circuit. Since these components generally reside in the same clock domain, we can schedule their execution statically to avoid shared-resource contention among threads. We propose thread weaving, a technique that statically interleaves requests from different threads through scheduling constraints. With the guarantee of a contention-free schedule, we eliminate replication/arbitration of shared resources, reducing the area footprint of the circuit and improving its maximum operating frequency (Fmax). Hsuan Hsiao, Jason Helge Anderson |
DAC | 2 |
| 2019 | Impact of FPGA Architecture on Area and Performance of CGRA OverlaysabstractCoarse-grained reconfigurable arrays (CGRAs) are programmable logic devices with ALU-style processing elements and datapath interconnect. CGRAs can be realized as custom ASICs or implemented on FPGAs as overlays . A key element of CGRAs is that they are typically software programmable with rapid compile times - an advantage arising from their coarse-grained characteristics, simplifying CAD mapping tasks. We implement two previously published CGRAs as overlays on two commercial FPGAs (Intel and Xilinx), and consider the impact of the underlying FPGA architecture on the CGRA area and performance. We present optimizations for the overlays to take advantage of the FPGA architectural features and show a peak performance improvement of 1.93x, as well as maximum area savings of 31.1% and 48.5% for Intel and Xilinx, respectively, relative to a naive first-cut implementation. We also present a novel technique for a configurable multiplexer implementation, which embeds the select signals into SRAM configuration, saving 35.7% in area. The research is conducted using the open-source CGRA-ME (modeling and exploration) framework [1]. Ian Taras, Jason Helge Anderson |
FCCM | 2 |
| 2019 | Generic Connectivity-Based CGRA Mapping via Integer Linear ProgrammingabstractCoarse-grained reconfigurable architectures (CGRAs) are programmable logic devices with large coarsegrained ALU-like logic blocks, and multi-bit datapath-style routing. CGRAs often have relatively restricted data routing networks, so they attract CAD mapping tools that use exact methods, such as Integer Linear Programming (ILP). However, tools that target general architectures must use large constraint systems to fully describe an architecture's flexibility, resulting in lengthy run-times. In this paper, we propose to derive connectivity information from an otherwise generic device model, and use this to create simpler ILPs, which we combine in an iterative schedule and retain most of the exactness of a fully-generic ILP approach. This new approach has a speed-up geometric mean of 5.88x when considering benchmarks that do not hita time-limit of 7.5 hours on the fully-generic ILP, and 37.6x otherwise. This was measured using the set of benchmarks used to originally evaluate the fully-generic approach and several more benchmarks representing computation tasks, over three different CGRA architectures. All run-times of the new approach are less than 20 minutes, with 90th percentile time of 410 seconds. The proposed mapping techniques are integrated into, and evaluated using the open-source CGRA-ME architecture modelling and exploration framework. Matthew J. P. Walker, Jason Helge Anderson |
FCCM | 2 |
| 2019 | EASY: Efficient Arbiter SYnthesis from Multi-threaded CodeabstractHigh-Level Synthesis (HLS) tools automatically transform a high-level specification of a circuit into a low-level RTL description. Traditionally, HLS tools have operated on sequential code, however in recent years there has been a drive to synthesize multi-threaded code. A major challenge facing HLS tools in this context is how to automatically partition memory amongst parallel threads to fully exploit the bandwidth available on an FPGA device and avoid memory contention. Current automatic memory partitioning techniques have inefficient arbitration due to conservative assumptions regarding which threads may access a given memory bank. In this paper, we address this problem through formal verification techniques, permitting a less conservative, yet provably correct circuit to be generated. We perform a static analysis on the code to determine which memory banks are shared by which threads. This analysis enables us to optimize the arbitration efficiency of the generated circuit. We apply our approach to the LegUp HLS tool and show that for a set of typical application benchmarks we can achieve up to 87% area savings, and 39% execution time improvement, with little additional compilation time. Jianyi Cheng, Shane T. Fleming, Yu Ting Chen, Jason Helge Anderson, George A. Constantinides |
FPGA | 4 |
| 2019 | A Dynamic Memory Allocation Library for High-Level SynthesisabstractOne impediment to the uptake of high-level synthesis (HLS) design methodologies is their lack of support for constructs frequently employed by software engineers - a primary example being dynamic memory allocation routines. No commercial HLS tool supports these constructs, forcing designers to rewrite programs to remove any dynamic memory allocation function calls (e.g.malloc(), free()), replacing them with statically allocated data. This shortcoming limits the portability of C/C++ descriptions, may introduce software bugs, and forces users to overestimate memory requirements, consuming precious on-chip BRAM resources. We address these problems by extending the capabilities of modern HLS tools through introduction of a tool-independent, HLS-friendly C library of five dynamic memory allocation schemes. Additionally, we developed a benchmark suite to evaluate and compare all five allocation schemes for their performance, area and memory trade-offs. We use the high-level synthesis tool, LegUp, to conduct our experiments. Our results indicate that each allocator in our library is best-suited for certain applications, in terms of performance, area and memory usage. We provide usage guidelines to assist HLS developers in selecting an appropriate allocation scheme. Nicholas V. Giamblanco, Jason Helge Anderson |
FPL | 2 |
| 2018 | An architecture-agnostic integer linear programming approach to CGRA mappingabstractCoarse-grained reconfigurable architectures (CGRAs) have gained traction as a potential solution to implement accelerators for compute-intensive kernels, particularly in domains requiring hardware programmability. Architecture and CAD for CGRAs are tightly intertwined, with many prior works having combined architectures and tools. In this work, we present an architecture-agnostic integer linear programming (ILP) approach for CGRA mapping, integrated within an open-source CGRA architecture evaluation framework. The mapper accepts an application and an architecture description as input and can generate an optimal mapping, if indeed mapping is feasible. An experimental study demonstrates its effectiveness over a range of CGRA architectures. S. Alexander Chin, Jason Helge Anderson |
DAC | 2 |
| 2018 | High-level synthesis of software-customizable floating-point coresabstractParameterized cores with fixed capabilities are typically used for floating-point (FP) operations on FPGAs. However, such standard cores can be over provisioned or lack specific specializations as required by applications. We consider FP cores described in the C language, synthesized to hardware using the LegUp high-level synthesis (HLS) tool [1]. Their software specification permits straightforward customization to non-compliant variants having superior area and performance characteristics, such as reduced-precision floating point, or cores without full IEEE 754 exceptions support. We create and evaluate the IEEE 754 FP standard cores for the key operations of addition, subtraction, division and multiplication, targeted to an FPGA and compare with widely used optimized RTL FP cores from Altera [7] and FloPoCo [3]. The software-specified HLS-generated cores are surprisingly close to the optimized RTL cores in terms of area/performance, and superior in certain cases, such as FP division. Samridhi Bansal, Hsuan Hsiao, Tomasz S. Czajkowski, Jason Helge Anderson |
DATE | 4 |
| 2018 | Sensei: An area-reduction advisor for FPGA high-level synthesisabstractHigh-level synthesis (HLS) provides an easy-to-use abstraction for designing hardware circuits. However, standard datatypes in high-level languages are over provisioned for typical applications, incurring extra area since the underlying FPGA hardware can support arbitrary bitwidths. This area inefficiency can be overcome by enabling the use of arbitrary-width datatypes at the source code level. However, this requires that HLS users spend time and effort on examining all program variables and quantifying their area impact, which can be intractable especially with large, complex programs and time-consuming synthesis. We propose Sensei, an advisor that predicts the post-synthesis area savings brought about by reducing bitwidth and presents users with a ranking of program variables and their area impact. Equipped with a convolutional neural network (CNN)-based predictor, Sensei achieves high area-prediction accuracy and enables rapid exploration of area-saving opportunities. Hsuan Hsiao, Jason Helge Anderson |
DATE | 2 |
| 2018 | High-Level Synthesis of FPGA Circuits with Multiple Clock DomainsabstractWe consider the high-level synthesis of circuits with multiple clock domains in a bid to raise circuit performance. A profiling-based approach is used to select time-intensive sub-circuits within a larger circuit to operate on separate clock domains. This isolates the critical paths of the sub-circuits from the larger circuit, allowing the sub-circuits to be clocked at the highest-possible speed. The open-source LegUp high-level synthesis tool (HLS) is modified to automatically insert clock-domain-crossing circuitry for signals crossing between two domains. The scheduling and binding phases of HLS were changed to reflect the impact of multiple clock domains on memory. Namely, the block RAMs in FPGAs are dual-port, where each port can operate on a different domain, implying that sub-circuits on different domains can access shared memory provided the domains of the memory ports are consistent with the sub-circuit domains. In an experimental study, we apply multi-clock domain HLS to the CHStone benchmark suite and demonstrate average wall-clock time improvements of 33%. Omar Ragheb, Jason Helge Anderson |
FCCM | 2 |
| 2018 | Software-Specified FPGA Accelerators for Elementary FunctionsabstractWe use a high-level synthesis (HLS) methodology for the design of hardware accelerators for two elementary functions: reciprocal and square root. The functions are described in C-language software and synthesized into Verilog RTL using the LegUp HLS tool from the University of Toronto [1]. The accelerators are designed to deliver high accuracy, and provide less than 1 ULP error in comparison with GNU software (math.h). Through changes to the HLS constraints, hardware implementations with different speed/area trade-offs can be generated rapidly. In an experimental study, our HLS-generated accelerators are targeted to the Altera/Intel Cyclone V FPGA and compared with hand-designed cores from the FPGA vendor. Results show that our cores offer considerably better resource usage (area) (i.e. ALMs, DSPs, memory bits), while commercial cores operate at a modestly higher FMax. Jason Helge Anderson |
FPT | 3 |
| 2018 | Synthesizable Heterogeneous FPGA FabricsabstractWe present an automated framework for the generation of synthesizable FPGAs with heterogeneous functional blocks and carry chains, as modelled with the open-source Verilog-to-Routing (VTR) FPGA architecture evaluation framework. VTR's modelling of hardened blocks, such as DSPs and BRAMs, is leveraged to generate synthesizeable FPGAs mappable via VTR's Verilog frontend. The generated Verilog source for the FPGA can be synthesized to target any conventional semiconductor process via an industry-standard ASIC toolflows with minimal implementation effort. We model a Stratix IV-style FPGA architecture, complete with carry chains, DSPs and BRAMs, and compare area/performance with the commercial Stratix IV FPGA. The area and performance gap between the fully synthesizable and commercial fabrics for a set of benchmarks using the heterogeneous blocks is 3.2× and 2.3×, respectively. Optimizations to reduce the gap are discussed. Brett Grady, Jason Helge Anderson |
FPT | 2 |
| 2018 | FPGA Architecture Enhancements for Efficient BNN ImplementationabstractBinarized neural networks (BNNs) are ultra-reduced precision neural networks, having weights and activations restricted to single-bit values. BNN computations operate on bitwise data, making them particularly amenable to hardware implementation. In this paper, we first analyze BNN implementations on contemporary commercial 20nm FPGAs. We then propose two lightweight architectural changes that significantly improve the logic density of FPGA BNN implementations. The changes involve incorporating additional carry-chain circuitry into logic elements, where the additional circuitry is connected in a specific way to benefit BNN computations. The architectural changes are evaluated in the context of state-of-the-art Intel and Xilinx FPGAs and shown to provide over 2x area reduction in the key BNN computational task (the XNOR-popcount sub-circuit), at a modest performance cost of less than 2%. Jin Hee Kim, Jongeun Lee, Jason Helge Anderson |
FPT | 3 |
| 2018 | Compact Area and Performance Modelling for CGRA Architecture EvaluationabstractWe present area and performance models for use in coarse-grained reconfigurable array (CGRAs) architectural exploration. The area and performance models can be computed rapidly and are incorporated into the open-source CGRA-ME architecture evaluation framework. Area is modelled by synthesizing (into standard cells) commonly occurring CGRA primitives in isolation, and then aggregating the component-wise areas. For performance, we incorporate a fully fledged static-timing analysis (STA) framework into CGRA-ME. The delays in the STA timing graph are annotated based on: 1) a library component-wise delays for logic/memory, and 2) a fanout-based delay estimation model for interconnect. Performance and area are modelled for both performance-optimized and area-optimized standard-cell CGRA implementations. Accuracy of the area and performance models is within 7% and 10%, respectively, of a fully laid-out standard-cell CGRA implementation. Kuang Ping Niu, Jason Helge Anderson |
FPT | 2 |
| 2018 | Architecture Exploration of Standard-Cell and FPGA-Overlay CGRAs Using the Open-Source CGRA-ME FrameworkabstractWe describe an open-source software framework,CGRA-ME, for the modeling and exploration of coarse-grained reconfigurable architectures (CGRAs). CGRAs are programmable hardware devices having large ALU-like logic blocks, and datapath bus-style inter-connect. CGRAs are positioned between fine-grained FPGAs and standard-cell ASICs on the spectrum of programmability - they are less flexible than FPGAs, yet are more flexible than ASICs. With CGRA-ME, an architect can describe a CGRA architecture in an XML-based language. The framework also allows the architect to map benchmarks onto the architecture and provides automatic generation of Verilog RTL for the modeled architecture. This allows the architect to simulate for verification purposes, and perform synthesis to either an ASIC or FPGA-overlay implementation of the CGRA, assessing performance, area, and power consumption. In an experimental study, we use CGRA-ME to model, map benchmarks onto, and evaluate several variants of a widely known CGRA, considering both standard-cell and FPGA-overlay physical realizations of the CGRA. S. Alexander Chin, Kuang Ping Niu, Matthew J. P. Walker, Shizhang Yin, Alexander Mertens, Jongeun Lee, Jason Helge Anderson |
ISPD | 7 |
| 2017 | CGRA-ME: A unified framework for CGRA modelling and explorationabstractCoarse-grained reconfigurable arrays (CGRAs) are a style of programmable logic device situated between FPGAs and custom ASICs on the spectrum of programmability, performance, power and cost. CGRAs have been proposed by both academia and industry; however, prior works have been mainly self-contained without broad architectural exploration and comparisons with competing CGRAs. We present CGRA-ME - a unified CGRA framework that encompasses generic architecture description, architecture modelling, application mapping, and physical implementation. Within this framework, we discuss our architecture description language CGRA-ADL, a generic LLVM-based simulated annealing mapper, and a standard cell flow for physical implementation. An architecture exploration case study is presented, highlighting the capabilities of CGRA-ME by exploring a variety of architectures with varying functionality, interconnect, array size, and execution contexts through the mapping of application benchmarks and the production of standard cell designs. S. Alexander Chin, Noriaki Sakamoto, Allan Rui, Jim Zhao, Jin Hee Kim, Yuko Hara-Azumi, Jason Helge Anderson |
ASAP | 7 |
| 2017 | Automated generation of banked memory architectures in the high-level synthesis of multi-threaded softwareabstractSome modern high-level synthesis (HLS) tools [1] permit the synthesis of multi-threaded software into parallel hardware, where concurrent software threads are realized as concurrently operating hardware units. A common performance bottleneck in any parallel implementation (whether it be hardware or software) is memory bandwidth — parallel threads demand concurrent access to memory resulting in contention which hurts performance. FPGAs contain an abundance of independently accessible memories offering high internal memory bandwidth. We describe an approach for leveraging such bandwidth in the context of synthesizing parallel software into hardware. Our approach applies trace-based profiling to determine how a program's arrays should be automatically partitioned into sub-arrays, which are then implemented in separate on-chip RAM blocks within the target FPGA. The partitioning is accomplished in a way that requires a single HLS execution and logic simulation for trace extraction. The end result is that each thread, when implemented in hardware, has exclusive access to its own memories to the extent possible, significantly reducing contention and arbitration and thus raising performance. Yu Ting Chen, Jason Helge Anderson |
FPL | 2 |
| 2017 | Synthesizable Standard Cell FPGA Fabrics Targetable by the Verilog-to-Routing CAD FlowabstractIn this article, we consider implementing field-programmable gate arrays (FPGAs) using a standard cell design methodology and present a framework for the automated generation of synthesizable FPGA fabrics. The open-source Verilog-to-Routing (VTR) FPGA architecture evaluation framework [Rose et al. 2012] is extended to generate synthesizable Verilog for its in-memory FPGA architectural device model. The Verilog can subsequently be synthesized into standard cells, placed and routed using an ASIC design flow. A second extension to VTR generates a configuration bitstream for the FPGA, where the bitstream configures the FPGA to realize a user-provided placed and routed design. The proposed framework and methodology makes possible the silicon implementation of a wide range of VTR-modeled FPGA fabrics. In an experimental study, area and timing-optimized FPGA implementations in 65nm TSMC standard cells are compared to a 65nm Altera commercial FPGA. In addition, we consider augmenting the generic standard-cell library from TSMC with a manually designed and laid-out FPGA-specific cell. We demonstrate the utility of the custom cell in reducing the area of the synthesized FPGA fabric. Jin Hee Kim, Jason Helge Anderson |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2017 | Microarchitectural Comparison of the MXP and Octavo Soft-Processor FPGA OverlaysabstractField-Programmable Gate Arrays (FPGAs) can yield higher performance and lower power than software solutions on CPUs or GPUs. However, designing with FPGAs requires specialized hardware design skills and hours-long CAD processing times. To reduce and accelerate the design effort, we can implement an overlay architecture on the FPGA, on which we then more easily construct the desired system but at a large cost in performance and area relative to a direct FPGA implementation. In this work, we compare the micro-architecture, performance, and area of two soft-processor overlays: the Octavo multi-threaded soft-processor and the MXP soft vector processor. To measure the area and performance penalties of these overlays relative to the underlying FPGA hardware, we compare direct FPGA implementations of the micro-benchmarks written in C synthesized with the LegUp HLS tool and also written in the Verilog HDL. Overall, Octavo’s higher operating frequency and MXP’s more efficient code execution results in similar performance from both, within an order of magnitude of direct FPGA implementations, but with a penalty of an order of magnitude greater area. Charles Eric LaForest, Jason Helge Anderson |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2017 | The First 25 Years of the FPL Conference: Significant PapersabstractA summary of contributions made by significant papers from the first 25 years of the Field-Programmable Logic and Applications conference (FPL) is presented. The 27 papers chosen represent those which have most strongly influenced theory and practice in the field. Philip H. W. Leong, Hideharu Amano, Jason Helge Anderson, Koen Bertels, João M. P. Cardoso, Oliver Diessel, Guy Gogniat, Mike Hutton, Wayne Luk, Patrick Lysaght, Marco Platzner, Viktor Prasanna 0001, Tero Rissa, Cristina Silvano, Hayden Kwok-Hay So, Yu Wang 0002 |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2017 | From Pthreads to Multicore Hardware Systems in LegUp High-Level Synthesis for FPGAsabstractIn the last decade, processor speeds have remained fairly stagnant, and to improve performance further, the industry started to increase the number of processor cores. The use of specialized hardware, such as field-programmable gate arrays (FPGAs), has also been on the rise. The traditional design methodology for FPGAs, however, requires hardware knowledge, which makes the platform inaccessible to software engineers. High-level synthesis (HLS) tools aim to resolve this issue by allowing software design methodologies to be used for FPGAs. However, HLS remains difficult to use for many software engineers, as there are tasks, such as system integration, which is still mostly a manual process. Consequently, creating a multicore hardware system on an FPGA is not feasible for most software engineers. To this end, we provide an HLS framework, which can automatically generate a multicore hardware system from software. We provide support for POSIX threads, which can be compiled to concurrently executing hardware cores that can be used in a processor-accelerator hybrid system, or in a hardware-only system without a processor. With this, we show that we can create multicore FPGA systems that can provide significant benefits in performance and energy-efficiency compared with hardware executing sequentially, and software executing on MIPS/ARM/x86 processors. Jongsok Choi, Stephen Brown 0003, Jason Helge Anderson |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | Leveraging Unused Resources for Energy Optimization of FPGA InterconnectabstractConventional field-programmable gate arrays are typically overprovisioned with routing resources to ensure that they meet routeability targets, which results in increased routing static and dynamic power. In this paper, we leverage the excess routing conductors to reduce dynamic and static power. To reduce dynamic power, we propose to ensure that used routing conductors are adjacent to unused routing conductors, which are left floating to reduce the effective capacitance seen by active nets. To reduce static power, we observe that leakage in routing multiplexers is dominated by specific paths; if the routing conductors, which connect to the input pins on these paths, are unused and left floating, the leakage of the multiplexer may be significantly reduced. To ensure that unused conductors are allowed to float requires the use of tristate routing buffers, and thus we propose two low-cost tristate buffer topologies with different power and area-overhead tradeoffs. We also introduce CAD techniques to optimize the overall energy dissipation in the routing network using the proposed techniques. Results show that interconnect dynamic power reductions of up to 25%, interconnect static power reductions of up to 81%, and overall interconnect energy reductions ranging between 14.9%-42.7% are expected, with a critical path degradation of <;1.8% and area-overhead of 2.6%-4.8%. Safeen Huda, Jason Helge Anderson |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | A unified software approach to specify pipeline and spatial parallelism in FPGA hardwareabstractHigh-level synthesis (HLS) is increasingly becoming a mainstream design methodology for FPGAs. Whereas its previous applications were mostly limited to research and simple designs, it is now being used to tape-out real-world chips in production [1]. Advances in compiler and HLS research continue to improve the quality of HLS-generated hardware. Despite this, the ease-of-use of HLS tools remains a hurdle to its broad uptake, particularly by engineers without hardware skills. To this end, we propose using a well-known software technique to infer streaming parallel hardware in HLS. Specifically, we use the producer-consumer pattern, commonly used in multi-threaded programming, to infer the generation of hardware that can exploit both pipeline and spatial parallelism on FPGAs. Our proposed methodology allows one to create a design in software, using only standard software methodologies, that cannot only synthesize to streaming hardware, but also model the generated hardware more accurately than existing solutions from other state-of-the-art C-based HLS tools. We use four different real-world benchmarks to illustrate the use of our methodology, and how it can create circuits that are either pipelined, or pipelined and replicated, all from software. For comparison, we also use a commercial HLS tool to synthesize one of the benchmarks, and show that our methodology can produce competitive results to that of the commercial tool. Jongsok Choi, Ruolong Lian, Stephen Brown 0003, Jason Helge Anderson |
ASAP | 4 |
| 2016 | Effect of LFSR seeding, scrambling and feedback polynomial on stochastic computing accuracy
Jason Helge Anderson, Yuko Hara-Azumi, Shigeru Yamashita |
DATE | 1 |
| 2016 | Towards PVT-Tolerant Glitch-Free Operation in FPGAsabstractGlitches are unnecessary transitions on logic signals that needlessly consume dynamic power. Glitches arise from imbalances in the combinational path delays to a signal, which may cause the signal to toggle multiple times in a given clock cycle before settling to its final value. In this paper, we propose a low-cost circuit structure that is able to eliminate a majority of glitches. The structure, which is incorporated into the output buffers of FPGA logic elements, suppresses pulses on buffer outputs whose duration is shorter than a configurable time window (set at the time of FPGA configuration). Glitches are thereby eliminated "at the source" ensuring they do not propagate into the high-capacitance FPGA interconnect, saving power. An experimental study, using Altera commercial tools for power analysis, demonstrates that the proposed technique reduces 70% of glitches, at a cost of 1% reduction in speed performance. Safeen Huda, Jason Helge Anderson |
FPGA | 2 |
| 2016 | PrefaceabstractAnother year and another step forward in reconfigurable computing as it joins the computing technology mainstream. For the past 26 years FPL has reflected this progress in its technical program, its keynote addresses, and the programs of its workshops and tutorials. This year FPL is hosted by the École polytechnique fédérale de Lausanne (EPFL) on the shores of the largest Alpine lake, Lake Geneva. Paolo Ienne, Walid A. Najjar, Jason Helge Anderson, Philip Brisk, Walter Stechele |
FPL | 3 |
| 2016 | High-level synthesis - the right side of historyabstractHigh-level synthesis (HLS) was first proposed in the 1980s. After spending decades on the sidelines of mainstream RTL digital design, there has been tremendous buzz around HLS technology in recent years. Indeed, HLS is on the upswing as a design methodology for field-programmable gate arrays (FPGAs) to improve designer productivity and ultimately, to make FPGA technology accessible to software engineers having limited hardware expertise. The hope is that down the road, software developers could use HLS to realize FPGA-based accelerators customized to applications that work in tandem with standard processors to raise computational throughput and energy efficiency. And, the further hope is that such HLS-generated accelerators operate close to the speed and energy efficiency of human-expert-designed accelerators. In this talk, I will overview the trends behind the recent drive towards FPGA HLS and why the need for, and use of, HLS will only become more pronounced in the coming years. I will argue that HLS, as opposed to traditional RTL design, is on the “right side of history”. The talk will highlight current HLS research directions and expose some of the challenges for HLS that may hinder its update in the digital design community. I will also describe work underway in the LegUp HLS project at the University of Toronto - a publicly available HLS tool that has been downloaded by over 4000 groups from around the world. Jason Helge Anderson |
FPT | 1 |
| 2016 | Power Optimization of FPGA Interconnect Via Circuit and CAD TechniquesabstractWe target power dissipation in field-programmable gate array (FPGA) interconnect and present three approaches that leverage a unique property of FPGAs, namely, the presence of unused routing conductors. A first technique attacks dynamic power by placing unused conductors, adjacent to used conductors, into a high-impedance state, reducing the effective capacitance seen by used conductors. A second technique, charge recycling, re-purposes unused conductors as charge reservoirs to reduce the supply current drawn for a positive transition on a used conductor. A third approach reduces leakage current in interconnect buffers by pulse-based signalling, allowing a driving buffer to be placed into a high impedance stage after a logic transition. All three techniques require CAD support in the routing stage to encourage specific positionings of unused conductors relative to used conductors. Safeen Huda, Jason Helge Anderson |
ISPD | 2 |
| 2016 | A Survey and Evaluation of FPGA High-Level Synthesis ToolsabstractHigh-level synthesis (HLS) is increasingly popular for the design of high-performance and energy-efficient heterogeneous systems, shortening time-to-market and addressing today’s system complexity. HLS allows designers to work at a higher-level of abstraction by using a software program to specify the hardware functionality. Additionally, HLS is particularly interesting for designing field-programmable gate array circuits, where hardware implementations can be easily refined and replaced in the target device. Recent years have seen much activity in the HLS research community, with a plethora of HLS tool offerings, from both industry and academia. All these tools may have different input languages, perform different internal optimizations, and produce results of different quality, even for the very same input description. Hence, it is challenging to compare their performance and understand which is the best for the hardware to be implemented. We present a comprehensive analysis of recent HLS tools, as well as overview the areas of active interest in the HLS research community. We also present a first-published methodology to evaluate different HLS tools. We use our methodology to compare one commercial and three academic tools on a common set ofCbenchmarks, aiming at performing an in-depth evaluation in terms of performance and the use of resources. Razvan Nane, Vlad Mihai Sima, Christian Pilato, Jongsok Choi, Blair Fort, Andrew Canis, Yu Ting Chen, Hsuan Hsiao, Stephen Brown 0003, Fabrizio Ferrandi, Jason Helge Anderson, Koen Bertels |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 11 |
| 2016 | Hybrid LUT/Multiplexer FPGA Logic ArchitecturesabstractHybrid configurable logic block architectures for field-programmable gate arrays that contain a mixture of lookup tables and hardened multiplexers are evaluated toward the goal of higher logic density and area reduction. Multiple hybrid configurable logic block architectures, both nonfracturable and fracturable with varying MUX:LUT logic element ratios are evaluated across two benchmark suites (VTR and CHStone) using a custom tool flow consisting of LegUp-HLS, Odin-II front-end synthesis, ABC logic synthesis and technology mapping, and VPR for packing, placement, routing, and architecture exploration. Technology mapping optimizations that target the proposed architectures are also implemented within ABC. Experimentally, we show that for nonfracturable architectures, without any mapper optimizations, we naturally save up to ~8% area postplace and route; both accounting for complex logic block and routing area while maintaining mapping depth. With architecture-aware technology mapper optimizations in ABC, additional area is saved, post-place-and-route. For fracturable architectures, experiments show that only marginal gains are seen after place-and-route up to ~2%. For both nonfracturable and fracturable architectures, we see minimal impact on timing performance for the architectures with best area-efficiency. S. Alexander Chin, Jason Luu, Safeen Huda, Jason Helge Anderson |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2015 | Profiling-driven multi-cycling in FPGA high-level synthesis
Stefan Hadjis, Andrew Canis, Ryoya Sobue, Yuko Hara-Azumi, Hiroyuki Tomiyama, Jason Helge Anderson |
DATE | 6 |
| 2015 | Synthesizable-from-C Embedded Processor Based on MIPS-ISA and OISCabstractWe describe a lightweight open-source MIPS-ISA processor, wherein performance and area can be flexibly traded-off with one another. The processor contains an ultra-low-cost co-processor capable of executing programs comprised of SUBLEQ instructions (subtract and branch if the difference is ≤ 0), which recent work has shown to be sufficient for any computation. Area/performance trade-offs are realized by implementing a user-selectable subset of MIPS instructions with functionally equivalent SUBLEQ sub-routines that run on the coprocessor. Silicon area is reduced as more MIPS instructions are implemented with the co-processor, rather than "natively" using functional units within the host MIPS. The processor is described in the C language and synthesized to an FPGA hardware implementation with high-level synthesis (HLS). Since it is specified at a high level of abstraction, it is straightforward to tailor to any application. As such, the processor can be viewed as a family of processors with different area/performance/power characteristics. In an experimental study, we compare a variety of processor variants, wherein different subsets of MIPS instructions are handled by the co-processor. We also compare the proposed synthesizable processor with a hand-designed 5-pipeline-stage MIPS implementation, and achieve area reductions ranging from 2.5 - 4×. Tanvir Ahmed 0004, Noriaki Sakamoto, Jason Helge Anderson, Yuko Hara-Azumi |
EUC | 3 |
| 2015 | Synthesizable FPGA fabrics targetable by the Verilog-to-Routing (VTR) CAD flowabstractWe consider implementing FPGAs using a standard cell design methodology, and present a framework for the automated generation of synthesizable FPGA fabrics. The open-source Verilog-to-Routing (VTR) FPGA architecture evaluation framework [1] is extended to generate synthesizable Verilog for its in-memory FPGA architectural device model. The Verilog can be synthesized into standard cells, placed and routed using an ASIC design flow. A second extension to VTR generates a configuration bitstream for the FPGA; that is, the bitstream configures the FPGA to realize a user-provided placed and routed design. The proposed framework and methodology opens the door to silicon implementation of a wide range of VTR-modelled FPGA fabrics. In an experimental study, area and timing-optimized FPGA implementations in 65nm TSMC standard cells are compared with a 65nm Altera commercial FPGA. Jin Hee Kim, Jason Helge Anderson |
FPL | 2 |
| 2015 | Significant papers from the first 25 years of the FPL conferenceabstractThe list of significant papers from the first 25 years of the Field-Programmable Logic and Applications conference (FPL) is presented in this paper. These 27 papers represent those which have most strongly influenced theory and practice in the field. Philip H. W. Leong, Hideharu Amano, Jason Helge Anderson, Koen Bertels, João M. P. Cardoso, Oliver Diessel, Guy Gogniat, Mike Hutton, Wayne Luk, Patrick Lysaght, Marco Platzner, Viktor Prasanna 0001, Tero Rissa, Cristina Silvano, Hayden Kwok-Hay So, Yu Wang 0002 |
FPL | 3 |
| 2015 | Resource and memory management techniques for the high-level synthesis of software threads into parallel FPGA hardwareabstractRecent work has proposed the high-level synthesis of parallel software programs (specified using Pthreads or OpenMP) into concurrently operating parallel hardware modules [6]. In this paper, we describe resource and memory management techniques for improving performance and area of hardware generated by such software thread synthesis. One direction investigated pertains to how modules in the HLS-generated parallel hardware should connect to one another: 1) with a nested topology, or 2) with a flat topology. In the nested topology, hardware modules are created in a hierarchical manner: modules are instantiated inside within modules that use them. Conversely, the flat topology instantiates all hardware modules at the same level of hierarchy. For the flat topology, we describe a system generator that automatically generates the required interconnect between all hardware modules, as well as flexibly shares or replicates functions, functional units, and memories. We also explore methods to reduce memory contention among hardware units that operate in parallel, by investigating three different memory architectures which use: 1) a global memory controller, 2) local memories, and 3) shared-local memories. Local and shared-local memories are dedicated RAM blocks for a single or a set of hardware modules, and help to increase memory bandwidth by allowing concurrent memory accesses. We also consider memory replication to localize memories in hardware modules, and convert small memories to registers to further improve performance and memory usage. Finally, we describe implementing locks and barriers in HLS hardware: synchronization constructs used in parallel programming. We show that with our resource and memory management techniques, we can improve the geomean performance, area, and area-delay product of parallel HLS-generated hardware up to 41.6%, 38.3%, and 63.3%, respectively, for a set of 15 benchmarks. Jongsok Choi, Stephen Brown 0003, Jason Helge Anderson |
FPT | 3 |
| 2015 | The Effect of Compiler Optimizations on High-Level Synthesis-Generated HardwareabstractWe consider the impact of compiler optimizations on the quality of high-level synthesis (HLS)-generated field-programmable gate array (FPGA) hardware. Using an HLS tool implemented within the state-of-the-art LLVM compiler, we study the effect of compiler optimizations on the hardware metrics of circuit area, execution cycles, FMax , and wall-clock time. We evaluate 56 different compiler optimizations implemented within LLVM and show that some optimizations significantly affect hardware quality. Moreover, we show that hardware quality is also affected by some optimization parameter values, as well as the order in which optimizations are applied. We then present a new HLS-directed approach to compiler optimizations, wherein we execute partial HLS and profiling at intermittent points in the optimization process and use the results to judiciously undo the impact of optimization passes predicted to be damaging to the generated hardware quality. Results show that our approach produces circuits with 16% better speed performance, on average, versus using the standard -O3 optimization level. Qijing Huang 0001, Ruolong Lian, Andrew Canis, Jongsok Choi, Ryan Xi, Nazanin Calagar, Stephen Brown 0003, Jason Helge Anderson |
ACM Trans. Reconfigurable Technol. Syst. | 8 |
| 2014 | Automating the Design of Processor/Accelerator Embedded Systems with LegUp High-Level SynthesisabstractLegUp [1] is an open-source high-level synthesis (HLS) tool that accepts a C program as input and automatically synthesizes it into a hybrid system. The hybrid system comprises an embedded processor and custom accelerators that realize user-designated compute-intensive parts of the program with improved throughput and energy efficiency. In this paper, we overview the LegUp framework and describe several recent developments: 1) support for an embedded ARM processor, as is available on Altera's recently released SoC FPGA, 2) HLS support for software parallelization schemes -- pthreads and OpenMP, 3) enhancements to LegUp's core HLS algorithms that raise the quality of the auto-generated hardware, and, 4) a preliminary debugging and verification framework providing C source-level debugging of HLS hardware. Since its first release in 2011, LegUp has been downloaded over 1000 times by groups around the world, providing a powerful platform for new research in high-level synthesis algorithms and embedded systems design. Blair Fort, Andrew Canis, Jongsok Choi, Nazanin Calagar, Ruolong Lian, Stefan Hadjis, Yu Ting Chen, Mathew Hall, Bain Syrowik, Tomasz S. Czajkowski, Stephen Brown 0003, Jason Helge Anderson |
EUC | 12 |
| 2014 | On Hard Adders and Carry Chains in FPGAsabstractHardened adder and carry logic is widely used in commercial FPGAs to improve the efficiency of arithmetic functions. There are many design choices and complexities associated with such hardening, including circuit design, FPGA architectural choices, and the CAD flow. There has been very little study, however, on these choices and hence we explore a number of possibilities for hard adder design. We also highlight optimizations during front-end elaboration that help ameliorate the restrictions placed on logic synthesis by hardened arithmetic. We show that hard adders and carry chains, when used for simple adders, increase performance by a factor of four or more, but on larger benchmark designs that contain arithmetic, improve overall performance by roughly 15%. We measure an average area increase of 5% for architectures with carry chains but believe that better logic synthesis should reduce this penalty. Interestingly, we show that adding dedicated inter-logic-block carry links or fast carry look-ahead hardened adders result in only minor delay improvements for complete designs. Jason Luu, Conor McCullough, Safeen Huda, Charles Chiasson, Kenneth B. Kent, Jason Helge Anderson, Jonathan Rose, Vaughn Betz |
FCCM | 8 |
| 2014 | Optimizing effective interconnect capacitance for FPGA power reductionabstractWe propose a technique to reduce the effective parasitic capacitance of interconnect routing conductors in a bid to simultaneously reduce power consumption and improve delay. The parasitic capacitance reduction is achieved by ensuring routing conductors adjacent to those used by timing critical or high activity nets are left floating - disconnected from either VDD or GND. In doing so, the effective coupling capacitance between the conductors is reduced, because the original coupling capacitance between the conductors is placed in series with other capacitances in the circuit (series combinations of capacitors correspond to lower effective capacitance). To ensure unused conductors can be allowed to float requires the use of tri-state routing buffers, and to that end, we also propose low-cost tri-state buffer circuitry. We also introduce CAD techniques to maximize the likelihood that unused routing conductors are made to be adjacent to those used by nets with high activity or low slack, improving both power and speed. Results show that interconnect dynamic power reductions of up to ~15.5% are expected to be achieved with a critical path degradation of ~1%, and a total area overhead of ~2.1%. Safeen Huda, Jason Helge Anderson, Hirotaka Tamura |
FPGA | 2 |
| 2014 | Towards interconnect-adaptive packing for FPGAsabstractIn order to investigate new FPGA logic blocks, FPGA architects have traditionally needed to customize CAD tools to make use of the new features and characteristics of those blocks. The software development effort necessary to create such CAD tools can be a time-consuming process that can significantly limit the number and variety of architectures explored. Thus, architects want flexible CAD tools that can, with few or no software modifications, explore a diverse space. Existing flexible CAD tools suffer from impractically long runtimes and/or fail to efficiently make use of the important new features of the logic blocks being investigated. This work is a step towards addressing these concerns by enhancing the packing stage of the open-source VTR CAD flow [17] to efficiently deal with common interconnect structures that are used to create many kinds of useful novel blocks. These structures include crossbars, carry chains, dedicated signals, and others. To accomplish this, we employ three techniques in this work: speculative packing, pre-packing, and interconnect-aware pin counting. We show that these techniques, along with three minor modifications, result in improvements to runtime and quality of results across a spectrum of architectures, while simultaneously expanding the scope of architectures that can be explored. Compared with VTR 1.0 [17], we show an average 12-fold speedup in packing for fracturable LUT architectures with 20% lower minimum channel width and 6% lower critical path delay. We obtain a 6 to 7-fold speedup for architectures with non-fracturable LUTs and architectures with depopulated crossbars. In addition, we demonstrate packing support for logic blocks with carry chains. Jason Luu, Jonathan Rose, Jason Helge Anderson |
FPGA | 3 |
| 2014 | Source-level debugging for FPGA high-level synthesisabstractWe describe a source-level debugging framework for FPGA high-level synthesis (HLS) that offers gdb-like step, break, and data inspection functionality for an HLS-generated hardware circuit. With the proposed framework, the user can inspect the values of logic signals in the hardware from the C source code perspective. The logic signal values come from one of two sources: 1) a logic simulation of the RTL, or 2) an actual execution of the hardware on an FPGA. In addition to the software-like ecosystem for FPGA HLS debugging, the framework provides the user with insight on the RTL produced by the HLS tool for each C statement, and permits concurrent hardware and software debugging to discover the first point at which any logic signal in the hardware mismatches with its corresponding variable in software. Nazanin Calagar, Stephen Brown 0003, Jason Helge Anderson |
FPL | 3 |
| 2014 | Modulo SDC scheduling with recurrence minimization in high-level synthesisabstractLoop pipelining is a high-level synthesis scheduling technique that overlaps loop iterations to achieve higher performance. However, industrial designs often have resource constraints and other constraints imposed by cross-iteration dependencies. The interaction between multiple constraints can pose a challenge for HLS modulo scheduling algorithms, which, if not handled properly can lead to a loop pipeline schedule that fails to achieve the minimum possible initiation interval. We propose a novel modulo scheduler based on an SDC formulation that includes a backtracking mechanism to properly handle multiple scheduling constraints and still achieve the minimum possible initiation interval. The SDC formulation has the advantage of being a mathematical framework that supports flexible constraints that are useful for more complex loop pipelines. Furthermore, we describe how to specifically apply associative expression transformations during modulo scheduling to restructure recurrences in complex loops to enable better scheduling. We compared our techniques to existing prior work in modulo scheduling in HLS and also compared against a state-of-art commercial tool. Over a suite of benchmarks, we show that our scheduler and proposed optimizations can result in a geomean wall-clock time reduction of 32% versus prior work and 29% versus a commercial tool. Andrew Canis, Stephen Brown 0003, Jason Helge Anderson |
FPL | 3 |
| 2014 | Design re-use for compile time reduction in FPGA high-level synthesis flowsabstractHigh-level synthesis (HLS) raises the level of abstraction for hardware design through the use of software methodologies. An impediment to productivity in HLS flows, however, is the run-time of the back-end toolflow - synthesis, packing, placement and routing - which can take hours or days for the largest designs. We propose a new back-end flow for HLS that makes use of pre-synthesized and placed "macros" for portions of the design, thereby reducing the amount of work to be done by the back-end tools, lowering run-time. A key aspect of our work is an analytical placement algorithm capable of handling large macros whose internal blocks have fixed relative placements, in conjunction with placing the surrounding individual logic blocks. In an experimental study, we consider the impact on run-time and quality-of-results of using macros: 1) in synthesis alone, and 2) in synthesis, packing and placement. Results show that the proposed approach reduces run-time by ~3x, on average, with a negative performance impact of ~5%. Marcel Gort, Jason Helge Anderson |
FPT | 2 |
| 2014 | Approaching overhead-free execution on FPGA soft-processorsabstractImplementing systems on FPGA soft-processors, rather than as custom hardware, eases and accelerates the development process, but at the cost of a great reduction in performance. Orthogonal to limitations in parallelism or clock frequency, this reduction in performance primarily originates in the intrinsic addressing and flow-control overheads of scalar microprocessors, which expend a considerable number of cycles interleaving address calculations and branch decisions within the actual useful work. We present an improved FPGA soft-processor architecture which statically overlaps "overhead" computations and executes them in parallel with the "useful" computations, significantly reducing the number of processor cycles needed to execute sequential programs, while reducing maximum clock frequency to 0.939x of its original value. In addition to eliminating almost all overhead computations, the proposed soft-processor can operate at 500 MHz on the Altera Stratix IV FPGA - 0.909x of the absolute maximum rating. Combined, the high speed and execution efficiency increase the range of FPGA designs amenable to soft-processors rather than custom hardware. We evaluate our cycle count improvements with multiple benchmarks, achieving speedups ranging from 1.07x for control-heavy code, to 1.92x for looping code, never performing worse than the original sequential code, and always performing better than a totally unrolled loop. Charles Eric LaForest, Jason Helge Anderson, J. Gregory Steffan |
FPT | 2 |
| 2014 | Introduction to the Special Issue on the 11th International Conference on Field-Programmable Technology (FPT'12)abstractNo abstract available. Jason Helge Anderson, Kiyoung Choi |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2014 | VTR 7.0: Next Generation Architecture and CAD System for FPGAsabstractExploring architectures for large, modern FPGAs requires sophisticated software that can model and target hypothetical devices. Furthermore, research into new CAD algorithms often requires a complete and open source baseline CAD flow. This article describes recent advances in the open source Verilog-to-Routing (VTR) CAD flow that enable further research in these areas. VTR now supports designs with multiple clocks in both timing analysis and optimization. Hard adder/carry logic can be included in an architecture in various ways and significantly improves the performance of arithmetic circuits. The flow now models energy consumption, an increasingly important concern. The speed and quality of the packing algorithms have been significantly improved. VTR can now generate a netlist of the final post-routed circuit which enables detailed simulation of a design for a variety of purposes. We also release new FPGA architecture files and models that are much closer to modern commercial architectures, enabling more realistic experiments. Finally, we show that while this version of VTR supports new and complex features, it has a 1.5× compile time speed-up for simple architectures and a 6× speed-up for complex architectures compared to the previous release, with no degradation to timing or wire-length quality. Jason Luu, Jeffrey B. Goeders, Michael Wainberg, Andrew Somerville, Thien Yu, Konstantin Nasartschuk, Miad Nasr, Tim Liu, Nooruddin Ahmed, Kenneth B. Kent, Jason Helge Anderson, Jonathan Rose, Vaughn Betz |
ACM Trans. Reconfigurable Technol. Syst. | 12 |
| 2013 | Range and bitmask analysis for hardware optimization in high-level synthesisabstractWe consider the extent to which the bit-level representation of variables can be used to optimize hardware generated by high-level synthesis (HLS). Two approaches to bit-level optimization are considered (individually and together): 1) range analysis, and 2) bitmask analysis. Range analysis aims to predetermine min/max ranges for variables to reduce the bitwidth required to represent variables in hardware. Bitmask analysis characterizes individual bits within a word as either constants (1 or 0), sign bits, or unknowns, where constants/don't-cares permit hardware to be eliminated under certain conditions. Static compiler-based analysis is contrasted with dynamic profiling-based analysis in terms of their potential to impact area and speed of HLS-generated hardware. For a set of benchmarks implemented in the Altera Cyclone II FPGA, results show bit-level optimizations in HLS based on static analysis reduce circuit area by 9%, on average, while additional optimizations based on dynamic analysis provide 34% area reduction. Marcel Gort, Jason Helge Anderson |
ASP-DAC | 2 |
| 2013 | From software to accelerators with LegUp high-level synthesisabstractEmbedded system designers can achieve energy and performance benefits by using dedicated hardware accelerators. However, implementing custom hardware accelerators for an application can be difficult and time intensive. LegUp is an open-source high-level synthesis framework that simplifies the hardware accelerator design process [8]. With LegUp, a designer can start from an embedded application running on a processor and incrementally migrate portions of the program to hardware accelerators implemented on an FPGA. The final application then executes on an automatically-generated software/hardware coprocessor system. This paper presents on overview of the LegUp design methodology and system architecture, and discusses ongoing work on profiling, hardware/software partitioning, hardware accelerator quality improvements, Pthreads/OpenMP support, visualization tools, and debugging support. Andrew Canis, Jongsok Choi, Blair Fort, Ruolong Lian, Qijing Huang 0001, Nazanin Calagar, Marcel Gort, Jia Jun Qin, Mark Aldham, Tomasz S. Czajkowski, Stephen Brown 0003, Jason Helge Anderson |
CASES | 12 |
| 2013 | Multi-pumping for resource reduction in FPGA high-level synthesisabstractResource sharing is a classic high-level synthesis (HLS) optimization that saves area by mapping multiple operations to a single functional unit. With resource sharing, only operations scheduled in separate cycles can be assigned to shared hardware, which can result in longer schedules. In this paper, we propose a new approach to resource sharing that allows multiple operations to be performed by a single functional unit in one clock cycle. Our approach is based on multi-pumping, which operates functional units at a higher frequency than the surrounding system logic, typically 2×, allowing multiple computations to complete in a single system cycle. Our approach is particularly effective for DSP blocks on an FPGA, which are used to perform multiply and/or accumulate operations. Our results show that resource sharing using multi-pumping is comparable to traditional resource sharing in terms of area saved, but provides significant performance advantages. Specifically, when targeting a 50% reduction in DSP blocks, traditional resource sharing decreases circuit speed performance by 80%, on average, whereas multi-pumping decreases circuit speed by just 5%. Multi-pumping is a viable approach to achieve the area reductions of resource sharing, with considerably less negative impact to circuit performance. Andrew Canis, Jason Helge Anderson, Stephen Brown 0003 |
DATE | 2 |
| 2013 | The Effect of Compiler Optimizations on High-Level Synthesis for FPGAsabstractWe consider the impact of compiler optimizations on the quality of high-level synthesis (HLS)-generated FPGA hardware. Using a HLS tool implemented within the state-of-the-art LLVM [1] compiler, we study the effect of compiler optimizations on the hardware metrics of circuit area, execution cycles, Fmax, and wall-clock time. We evaluate 56 different compiler optimizations implemented within LLVM and show that some optimizations significantly affect hardware quality. Moreover, we show that hardware quality is also affected by the order in which optimizations are applied. We then present a new HLS-directed approach to compiler optimizations, wherein we execute partial HLS and profiling at intermittent points in the optimization process and use the results to judiciously undo the impact of optimization passes predicted to be damaging to the generated hardware quality. Results show that our approach produces circuits with 16% better speed performance, on average, versus using the standard -O3 optimization level. Qijing Huang 0001, Ruolong Lian, Andrew Canis, Jongsok Choi, Ryan Xi, Stephen Brown 0003, Jason Helge Anderson |
FCCM | 7 |
| 2013 | High-level synthesis with LegUp: a crash course for users and researchersabstractHigh-level synthesis (HLS) has been gaining traction recently as a design methodology for FPGAs, with the promise of raising the productivity of FPGA hardware designers, and ultimately, opening the door to the use of FPGAs as computing devices targetable by software engineers. In this tutorial, we introduce LegUp, an open-source HLS tool for FPGAs developed at the University of Toronto. With LegUp, a user can compile a C program completely to hardware, or alternately, he/she can choose to compile the program to a hybrid hardware/software system comprising a processor along with one or more accelerators. LegUp supports the synthesis of most of the C language to hardware, including loops, structs, multi-dimensional arrays, pointer arithmetic, and floating point operations. The LegUp distribution includes the CHStone HLS benchmark suite, as well as a test suite and associated infrastructure for measuring quality of results, and for verifying the functionality of LegUp-generated circuits. LegUp is freely downloadable at www.legup.org, providing a powerful platform that can be leveraged for new high-level synthesis research. Jason Helge Anderson, Stephen Brown 0003, Andrew Canis, Jongsok Choi |
FPGA | 1 |
| 2013 | Charge recycling for power reduction in FPGA interconnectabstractWe propose charge recycling (CR) to reduce power consumption in FPGAs. We take advantage of the property that many routing conductors are left unused in any FPGA implementation of an application. Charge recycling via the unused conductors reduces the amount of charge drawn from the supply, lowering energy consumption. We present a routing switch that operates in two modes: normal and CR, and describe the CAD tool changes needed to support CR at the routing and post-routing stages of the flow. Results show that dynamic power in the FPGA interconnect can be reduced by up to ~15-18.4% by the proposed techniques, depending on the performance constraints. Safeen Huda, Jason Helge Anderson, Hirotaka Tamura |
FPL | 2 |
| 2013 | From C to Blokus Duo with LegUp high-level synthesisabstractWe apply high-level synthesis (HLS) to generate Blokus Duo game-playing hardware for the FPT 2013 Design Competition [3]. Our design, written in C, is synthesized using the LegUp open-source HLS tool to Verilog, then subsequently mapped using vendor tools to an Altera Cyclone IV FPGA on DE2 board. Our software implementation is designed to be amenable to high-level synthesis, and includes a custom stack implementation, uses only integer arithmetic, and employs the use of bitwise logical operations to improve overall computational performance. The underlying AI decision making is based on alpha-beta pruning [2]. The performance of our synthesizable solution is gauged by playing against the Pentobi [8] - a “known good” C++ software implementation. Jiu Cheng Cai, Ruolong Lian, Andrew Canis, Jongsok Choi, Blair Fort, Eric Hart, Emily Miao, Nazanin Calagar, Stephen Brown 0003, Jason Helge Anderson |
FPT | 12 |
| 2013 | A case for hardened multiplexers in FPGAsabstractThis paper presents a case for a hybrid configurable logic block that contains a mixture of LUTs and hardened multiplexers towards the goal of higher logic density and area reduction. Technology mapping optimizations, called MuxMap, that target the proposed architecture are implemented using a modified version of the mapper in the ABC logic synthesis tool. VPR is used to model the new hybrid configurable logic block and verify post place and route implementation. Multiple hybrid configurable logic block architectures with varying MUX:LUT ratios are evaluated across three benchmark suites with both Quartus II and Odin-II front-end RTL synthesis tools. Experimentally, we show that without any mapper optimizations we naturally save ~4% area post place and route and with MuxMap optimizations in ABC yielding ~6% area reduction post place and route while maintaining mapping depth, overall configurable logic block count, and routing demand. S. Alexander Chin, Jason Helge Anderson |
FPT | 2 |
| 2013 | From software threads to parallel hardware in high-level synthesis for FPGAsabstractWe describe the support within high-level hardware synthesis (HLS) for two standard software parallelization paradigms: Pthreads and OpenMP. Parallel code segments, as specified in the software, are automatically synthesized by our HLS tool into parallel-operating hardware sub-circuits. Both data parallelism and task-level parallelism are supported, as is the combined use of both Pthreads and OpenMP. Moreover, our work also provides automated synthesis for commonly occurring synchronization constructs within the Pthreads/OpenMP library: mutual exclusion (mutex) and barriers. Essentially, our framework allows a software engineer to specify parallelism to an HLS tool using methodologies they are likely to be familiar with. An experimental study considers a variety of parallelization scenarios, including demonstrated speedups of up to 12.9× in circuit wall-clock time for the 16-thread case and area-delay product as low as 12% (~8× improvement) when using 4 pipelined hardware threads. Jongsok Choi, Stephen Brown 0003, Jason Helge Anderson |
FPT | 3 |
| 2013 | Bitwidth-optimized hardware accelerators with software fallbackabstractWe propose the high-level synthesis of an FPGA-based hybrid computing system, where the implementations of compute-intensive functions are available in both software, and as hardware accelerators. The accelerators are optimized to handle common-case inputs, as opposed to worst-case inputs, allowing accelerator area to be reduced by 28%, on average, while retaining the majority of performance advantages associated with a hardware versus software implementation. When inputs exceed the range that the hardware accelerators can handle, a software fallback is automatically triggered. Optimization of the accelerator area is achieved by reducing datapath widths based on application profiling of variable ranges in software (under typical datasets). The selected widths are passed to a high-level synthesis tool which generates the accelerator for a given function. The optimized accelerators with software fallback capability are generated automatically by our framework, with minimal user intervention. Our study explores the trade-offs of delay and area for benchmarks implemented on an Altera Cyclone II FPGA. Ana Klimovic, Jason Helge Anderson |
FPT | 2 |
| 2013 | DistCL: A Framework for the Distributed Execution of OpenCL KernelsabstractGPUs are used to speed up many scientific computations, however, to use several networked GPUs concurrently, the programmer must explicitly partition work and transmit data between devices. We propose DistCL, a novel framework that distributes the execution of penCL kernels across a GPU cluster. DistCL makes multiple distributed compute devices appear to be a single compute device. DistCL abstracts and manages many of the challenges associated with distributing a kernel across multiple devices including: (1) partitioning work into smaller parts, (2) scheduling these parts across the network, (3) partitioning memory so that each part of memory is written to by at most one device, and (4) tracking and transferring these parts of memory. Converting an OpenCL application to DistCL is straightforward and requires little programmer effort. This makes it a powerful and valuable tool for exploring the distributed execution of OpenCL kernels. We compare DistCL to SnuCL, which also facilitates the distribution of OpenCL kernels. We also give some insights: distributed tasks favor more compute bound problems and favour large contiguous memory accesses. DistCL achieves a maximum speedup of 29.1 and average speedups of 7.3 when distributing kernels among 32 peers over an Infiniband cluster. Tahir Diop, Steven Gurfinkel, Jason Helge Anderson, Natalie D. Enright Jerger |
MASCOTS | 3 |
| 2013 | Latch-Based Performance Optimization for Field-Programmable Gate ArraysabstractWe explore using pulsed latches for timing optimization in field-programmable gate arrays (FPGAs). Pulsed latches are transparent latches driven by a clock with a nonstandard (i.e., not 50%) duty cycle. As latches are already present on commercial FPGAs, their use for timing optimization can avoid the power or area drawbacks associated with other techniques such as clock skew and retiming. We propose algorithms that automatically replace certain flip-flops with latches for performance gains. Under conservative short path or minimum delay assumptions, our latch-based optimization, operating on already routed designs, provides all the benefit of clock skew in most cases and increases performance by 9%, on average, without area penalties or significant netlist changes. We show that short paths greatly hinder the ability of using pulsed latches, and that further improvements in performance are possible by increasing the delay of certain short paths. Bill Teng, Jason Helge Anderson |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2013 | LegUp: An open-source high-level synthesis tool for FPGA-based processor/accelerator systemsabstractIt is generally accepted that a custom hardware implementation of a set of computations will provide superior speed and energy efficiency relative to a software implementation. However, the cost and difficulty of hardware design is often prohibitive, and consequently, a software approach is used for most applications. In this article, we introduce a new high-level synthesis tool called LegUp that allows software techniques to be used for hardware design. LegUp accepts a standard C program as input and automatically compiles the program to a hybrid architecture containing an FPGA-based MIPS soft processor and custom hardware accelerators that communicate through a standard bus interface. In the hybrid processor/accelerator architecture, program segments that are unsuitable for hardware implementation can execute in software on the processor. LegUp can synthesize most of the C language to hardware, including fixed-sized multidimensional arrays, structs, global variables, and pointer arithmetic. Results show that the tool produces hardware solutions of comparable quality to a commercial high-level synthesis tool. We also give results demonstrating the ability of the tool to explore the hardware/software codesign space by varying the amount of a program that runs in software versus hardware. LegUp, along with a set of benchmark C programs, is open source and freely downloadable, providing a powerful platform that can be leveraged for new research on a wide range of high-level synthesis topics. Andrew Canis, Jongsok Choi, Mark Aldham, Victor Zhang, Ahmed Kammoona, Tomasz S. Czajkowski, Stephen Brown 0003, Jason Helge Anderson |
ACM Trans. Embed. Comput. Syst. | 8 |
| 2013 | Combined Architecture/Algorithm Approach to Fast FPGA RoutingabstractWe propose a new field-programmable gate array (FPGA) routing approach, which, when combined with a low-cost architecture change, results in a 40% reduction in router runtime, at the cost of a 6% area overhead and with no increase in critical path delay. Our approach begins with PathFinder-style routing, which we run on a coarsened representation of the routing architecture. This leads to fast generation of a partial routing solution where the signals are assigned to groups of wire segments rather than individual wire segments. A Boolean satisfiability (SAT)-based stage follows, generating a legal routing solution from the partial solution. We explore approximately 165 000 FPGA switch block architectures, showing that the choice of the architecture has a significant impact on the complexity of the SAT formulation, and by extension, on routing runtime. Our approach points to a new research direction, namely, reducing FPGA computer-aided design runtime by exploring FPGA architectures and algorithms together. Marcel Gort, Jason Helge Anderson |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | Leveraging reconfigurability to raise productivity in FPGA functional debugabstractWe propose new hardware and software techniques for FPGA functional debug that leverage the inherent reconfigurability of the FPGA fabric to reduce functional debugging time. The functionality of an FPGA circuit is represented by a programming bitstream that specifies the configuration of the FPGA's internal logic and routing. The proposed methodology allows different sets of design internal signals to be traced solely by changes to the programming bitstream followed by device reconfiguration and hardware execution. Evidently, the advantage of this new methodology vs. existing debug techniques is that it operates without the need of iterative executions of the computationally-intensive design re-synthesis, placement and routing tools. In essence, with a single execution of the synthesis flow, the new approach permits a large number of internal signals to be traced for an arbitrary number of clock cycles using a limited number of external pins. Experimental results using commercial FPGA vendor tools demonstrate productivity (i.e. run-time) improvements of up to 30× vs. a conventional approach to FPGA functional debugging. These results demonstrate the practicality and effectiveness of the proposed approach. Zissis Poulos, Yu-Shen Yang, Jason Helge Anderson, Andreas G. Veneris, Bao Le |
DATE | 3 |
| 2012 | Impact of Cache Architecture and Interface on Performance and Area of FPGA-Based Processor/Parallel-Accelerator SystemsabstractWe describe new multi-ported cache designs suitable for use in FPGA-based processor/parallel-accelerator systems, and evaluate their impact on application performance and area. The baseline system comprises a MIPS soft processor and custom hardware accelerators with a shared memory architecture: on-FPGA L1 cache backed by off-chip DDR2 SDRAM. Within this general system model, we evaluate traditional cache design parameters (cache size, line size, associativity). In the parallel accelerator context, we examine the impact of the cache design and its interface. Specifically, we look at how the number of cache ports affects performance when multiple hardware accelerators operate (and access memory) in parallel, and evaluate two different hardware implementations of multi-ported caches using: 1) multi-pumping, and 2) a recently-published approach based on the concept of a live-value table. Results show that application performance depends strongly on the cache interface and architecture: for a system with 6 accelerators, depending on the cache design, speed up swings from 0.73× to 6.14×, on average, relative to a baseline sequential system (with a single accelerator and a direct-mapped, 2KB cache with 32B lines). Considering both performance and area, the best architecture is found to be a 4-port multi-pump direct-mapped cache with a 16KB cache size and a 128B line size. Jongsok Choi, Kevin Nam, Andrew Canis, Jason Helge Anderson, Stephen Brown 0003, Tomasz S. Czajkowski |
FCCM | 4 |
| 2012 | Impact of FPGA architecture on resource sharing in high-level synthesisabstractResource sharing is a key area-reduction approach in high-level synthesis (HLS) in which a single hardware functional unit is used to implement multiple operations in the high-level circuit specification. We show that the utility of sharing depends on the underlying FPGA logic element architecture and that different sharing trade-offs exist when 4-LUTs vs. 6-LUTs are used. We further show that certain multi-operator patterns occur multiple times in programs, creating additional opportunities for sharing larger composite functional units comprised of patterns of interconnected operators. A sharing cost/benefit analysis is used to inform decisions made in the binding phase of an HLS tool, whose RTL output is targeted to Altera commercial FPGA families: Stratix IV (dual-output 6-LUTs) and Cyclone II (4-LUTs). Stefan Hadjis, Andrew Canis, Jason Helge Anderson, Jongsok Choi, Kevin Nam, Stephen Brown 0003, Tomasz S. Czajkowski |
FPGA | 3 |
| 2012 | The VTR project: architecture and CAD for FPGAs from verilog to routingabstractTo facilitate the development of future FPGA architectures and CAD tools -- both embedded programmable fabrics and pure-play FPGAs -- there is a need for a large scale, publicly available software suite that can synthesize circuits into easily-described hypothetical FPGA architectures. These circuits should be captured at the HDL level, or higher, and pass through logical and physical synthesis. Such a tool must provide detailed modelling of area, performance and energy to enable architecture exploration. As software flows themselves evolve to permit design capture at ever higher levels of abstraction, this downstream full-implementation flow will always be required. This paper describes the current status and new release of an ongoing effort to create such a flow - the 'Verilog to Routing' (VTR) project, which is a broad collaboration of researchers. There are three core tools: ODIN II for Verilog Elaboration and front-end hard-block synthesis, ABC for logic synthesis, and VPR for physical synthesis and analysis. ODIN II now has a simulation capability to help verify that its output is correct, as well as specialized synthesis at the elaboration step for multipliers and memories. ABC is used to optimize the 'soft' logic of the FPGA. The VPR-based packing, placement and routing is now fully timing-driven (the previous release was not) and includes new capability to target complex logic blocks. In addition we have added a set of four large benchmark circuits to a suite of previously-released Verilog HDL circuits. Finally, we illustrate the use of the new flow by using it to help architect a floating-point unit in an FPGA, and contrast it with a prior, much longer effort that was required to do the same thing. Jonathan Rose, Jason Luu, Chi Wai Yu, Opal Densmore, Jeffrey B. Goeders, Andrew Somerville, Kenneth B. Kent, Peter Jamieson, Jason Helge Anderson |
FPGA | 9 |
| 2012 | Analyzing and predicting the impact of CAD algorithm noise on FPGA speed performance and powerabstractFPGA CAD algorithms are heuristic, and generally make use of cost functions to gauge the value of one potential circuit implementation over another. At times, such algorithms must decide between two or more implementation options of apparently equal cost. This work explores the variations in circuit quality, i.e. noise, that arise when CAD algorithms are altered to choose randomly when faced with such equal-cost alternatives. Noise sources are identified in logic synthesis and technology mapping algorithms, and experimental results are presented which show standard deviations of 3.3% and 3.7% from the mean in post-routed delay and power. As a means of dealing with this variation, early timing and power prediction metrics can be applied after technology mapping to find the best circuits in the presence of noise. When applied to designs with over 1.5% variation in delay and power, the best prediction models have a 40% probability of capturing the best circuit when predicting the top 10% of circuits in a group of noise-injected circuits. Warren Wai-Kit Shum, Jason Helge Anderson |
FPGA | 2 |
| 2012 | Analytical placement for heterogeneous FPGAsabstractWe present HeAP, an analytical placement algorithm for heterogeneous FPGAs comprised of LUT-based logic blocks, multiplier/DSP blocks and block RAMs. Specifically, we adapt a state-of-the-art ASIC-based analytical placer to target FPGAs with heterogeneous blocks located at discrete locations throughout the fabric. Our placer also handles macros of LUT-based blocks with specific layout requirements, such as carry chains. Results show that our placer delivers a 4× speedup, on average, compared to Altera's non-timing driven flow, at the cost of a 5% increase in postrouted wirelength, and an 11× speedup compared to Altera's timing-driven flow, at the cost of a 4% increase in post-routed wirelength and a 9% reduction in maximum operating frequency. We also compare with an academic simulated annealing-based placer and demonstrate a 7.4× runtime advantage with 6% better placement quality. Marcel Gort, Jason Helge Anderson |
FPL | 2 |
| 2012 | FPGA power reduction by guarded evaluation considering physical informationabstractWe reconsider guarded evaluation as a means to reduce FPGA dynamic power consumption. We augment and evaluate guarded evaluation as proposed in [1] after different stages of the FPGA CAD flow. Guarding later in the flow provides more feedback to the algorithm and yields a more effective cost-benefit analysis of newly added signals. Numerical results show that guarding later in the flow yields slightly less power savings versus guarding after technology mapping. However, fewer guards are inserted which results in fewer netlist changes and less impact on routing resource usage. Chirag Ravishankar, Andrew A. Kennings, Jason Helge Anderson |
VLSI-SoC | 3 |
| 2012 | Accelerating FPGA Routing Through Parallelization and Engineering Enhancements Special Section on PAR-CAD 2010abstractWe present parallelization and heuristic techniques to reduce the run-time of field-programmable gate array (FPGA) negotiated congestion routing. Two heuristic optimizations provide over 3× speedup versus a sequential baseline. In our parallel approach, sets of design signals are assigned to different processor cores and routed concurrently. Communication between cores is through the message passing interface communications protocol. We propose a geographic partitioning of signals into independent sets to help minimize the communication overhead. Our parallel implementation provides approximately 2.3× speedup using four cores and produces deterministic/repeatable results. When combined, the parallel and heuristic techniques provide over 7× speedup with four cores versus the router in the widely used Versatile Place and Route (VPR) FPGA placement/routing framework, with no significant impact on circuit speed or wirelength. Marcel Gort, Jason Helge Anderson |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | FPGA Power Reduction by Guarded Evaluation Considering Logic ArchitectureabstractGuarded evaluation is a power reduction technique that involves identifying subcircuits (within a larger circuit) whose inputs can be held constant (guarded) at specific times during circuit operation, thereby reducing switching activity and lowering dynamic power. The concept is rooted in the property that under certain conditions, some signals within digital designs are not “observable” at design outputs, making the circuitry that generates such signals a candidate for guarding. Guarded evaluation has been demonstrated successfully for application-specific integrated circuits (ASICs); in this paper, we apply the technique to field-programmable gate arrays (FPGAs). In ASICs, guarded evaluation entails adding additional hardware to the design, increasing silicon area and cost. Here, we apply the technique in a way that imposes minimal area overhead by leveraging existing unused circuitry within the FPGA. The primary challenge in guarded evaluation is in determining the specific conditions under which a subcircuit's inputs can be held constant without impacting the larger circuit's functional correctness. We propose a simple solution to this problem based on discovering gating inputs using “noninverting” and “partial noninverting” paths in a circuit's AND-inverter graph representation. Experimental results show that guarded evaluation can reduce switching activity on average by as much as 32% and 25% for 6-input look-up table (6-LUT) and 4-LUT architectures, respectively. Dynamic power consumption in the FPGA interconnect is reduced on average by as much as 24% and 22% for 6-LUT and 4-LUT architectures, respectively. The impact to critical path delay ranges from 1% to 43%, depending on the guarding scenario and the desired power/delay tradeoff. Chirag Ravishankar, Jason Helge Anderson, Andrew A. Kennings |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | Raising FPGA Logic Density Through Synthesis-Inspired ArchitectureabstractWe leverage properties of the logic synthesis netlist to define both a new field-programmable gate-array (FPGA) logic element (function generator) architecture and an associated technology mapping algorithm that together provide improved logic density. We demonstrate that an “extended” logic element with slightly modifiedK-input lookup tables (LUTs) achieves much of the benefit of an architecture withK+1-input LUTs, while consuming silicon area close to aK-LUT (aK-LUT requires half the area of aK+1-LUT). We introduce the notion of “non-inverting paths” in a circuit's and-inverter graph (AIG) and show their utility in mapping into the proposed logic element architectures. We propose a general family of logic element architectures, and present results showing that they offer a variety of area/performance tradeoffs. One of our key results demonstrates that while circuits mapped to a traditional 5-LUT architecture need 15% more LUTs and have 14% more depth than a 6-LUT architecture, our extended 5-LUT architecture requires only 7% more LUTs and 5% more depth than 6-LUTs, on average. Nearly all of the depth reduction associated with moving fromK-input toK+1 -input LUTs can be achieved with considerably less area using extendedK-LUTs. We further show that 6-LUT optimal mapping depths can be achieved with a small fraction of the LUTs in hardware being 6-LUTs and the remainder being extended 5-LUTs, suggesting that a heterogeneous logic block architecture may prove to be advantageous. Jason Helge Anderson, Chirag Ravishankar |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2011 | Low-cost hardware profiling of run-time and energy in FPGA embedded processorsabstractThis paper introduces a low-overhead hardware profiling architecture, called LEAP, that attains real-time cycle and energy profiles of an FPGA-based soft processor. A novel technique is used to associate profiling data with specific functions in a way that is area- and power-efficient. Results show that relative to a previously-published hardware profiler, our design uses up to 18× less area and 8.6× less energy. LEAP is designed to be extensible for a variety of profiling tasks, three of which are investigated in this paper. We also demonstrate the utility of LEAP in the context of hardware/software co-design of processor/accelerator FPGA-based systems. Mark Aldham, Jason Helge Anderson, Stephen Brown 0003, Andrew Canis |
ASAP | 2 |
| 2011 | Area-efficient FPGA logic elements: Architecture and synthesisabstractWe consider architecture and synthesis techniques for FPGA logic elements (function generators) and show that the LUT-based logic elements in modern commercial FPGAs are over-engineered. Circuits mapped into traditional LUT-based logic elements have speeds that can be achieved by alternative logic elements that consume considerably less silicon area. We introduce the concept of a trimming input to a logic function, which is an input to a K-variable function about which Shannon decomposition produces a cofactor having fewer than K -1 variables. We show that trimming inputs occur frequently in circuits and we propose low-cost asymmetric FPGA logic element architectures that leverage the trimming input concept, as well as some other properties of a circuit's AND-inverter graph (AIG) functional representation. We describe synthesis techniques for the proposed architectures that combine a standard cut-based FPGA technology mapping algorithm with two straightforward procedures: 1) Shannon decomposition, and 2) finding non-inverting paths in the circuit's AIG. The proposed architectures exhibit improved logic density versus traditional LUT-based architectures with minimal impact on circuit speed. Jason Helge Anderson |
ASP-DAC | 1 |
| 2011 | An integer programming placement approach to FPGA clock power reductionabstractClock signals are responsible for a significant portion of dynamic power in FPGAs owing to their high toggle frequency and capacitance. Clock signals are distributed to loads through a programmable routing tree network, designed to provide low delay and low skew. The placement step of the FPGA CAD flow plays a key role in influencing clock power, as clock tree branches are connected based solely on the placement of the clock loads. In this paper, we present a placement-based approach to clock power reduction based on an integer linear programming (ILP) formulation. Our technique is intended to be used as an optimization post-pass executed after traditional placement, and it offers fine-grained control of the amount by which clock power is optimized versus other placement criteria. Results show that the proposed technique reduces clock network capacitance by over 50% with minimal deleterious impact on post-routed wirelength and circuit speed. Alireza Rakhshanfar, Jason Helge Anderson |
ASP-DAC | 2 |
| 2011 | LegUp: high-level synthesis for FPGA-based processor/accelerator systemsabstractIn this paper, we introduce a new open source high-level synthesis tool called LegUp that allows software techniques to be used for hardware design. LegUp accepts a standard C program as input and automatically compiles the program to a hybrid architecture containing an FPGA-based MIPS soft processor and custom hardware accelerators that communicate through a standard bus interface. Results show that the tool produces hardware solutions of comparable quality to a commercial high-level synthesis tool. Andrew Canis, Jongsok Choi, Mark Aldham, Victor Zhang, Ahmed Kammoona, Jason Helge Anderson, Stephen Brown 0003, Tomasz S. Czajkowski |
FPGA | 6 |
| 2011 | Architecture description and packing for logic blocks with hierarchy, modes and complex interconnectabstractThe development of future FPGA fabrics with more sophisticated and complex logic blocks requires a new CAD flow that permits the expression of that complexity and the ability to synthesize to it. In this paper, we present a new logic block description language that can depict complex intra-block interconnect, hierarchy and modes of operation. These features are necessary to support modern and future FPGA complex soft logic blocks, memory and hard blocks. The key part of the CAD flow associated with this complexity is the packer, which takes the logical atomic pieces of the complex blocks and groups them into whole physical entities. We present an area-driven generic packing tool that can pack the logical atoms into any heterogeneous FPGA described in the new language, including many different kinds of soft and hard logic blocks. We gauge its area quality by comparing the results achieved with a lower bound on the number of blocks required, and then illustrate its explorative capability in two ways: on fracturable LUT soft logic architectures, and on hard block memory architectures. The new infrastructure attaches to a flow that begins with a Verilog front-end, permitting the use of benchmarks that are significantly larger than the usual ones, and can target heterogenous FPGAs. Jason Luu, Jason Helge Anderson, Jonathan Rose |
FPGA | 2 |
| 2011 | Reducing FPGA Router Run-Time through Algorithm and ArchitectureabstractWe propose a new FPGA routing approach that, when combined with a low-cost architecture change, results in a 34% reduction in router run-time, at the cost of a 3% area overhead, with no increase in critical path delay. Our approach begins with traditional PathFinder-style routing, which we run on a coarsened representation of the routing architecture. This leads to fast generation of a partial routing solution where signals are assigned to groups of wire segments rather than individual wire segments. A boolean satisfiability (SAT)-based stage follows, generating a legal routing solution from the partial solution. Our approach points to a new research direction: reducing FPGA CAD run-time by exploring FPGA architectures and algorithms together. Marcel Gort, Jason Helge Anderson |
FPL | 2 |
| 2011 | Latch-Based Performance Optimization for FPGAsabstractWe explore using pulsed latches for timing optimization -- a first in the FPGA community. Pulsed latches are transparent latches driven by a clock with a non-standard (non-50%) duty cycle. We exploit existing functionality within commercial FPGA chips to implement latch-based optimizations that do not have the power or area drawbacks associated with other timing optimization approaches, such as clock skew and retiming. We propose an algorithm that iteratively replaces certain flip-flops in a logic design with latches for an improvement in circuit speed. Results show that much of the performance improvement achieved by using multiple skewed clocks can also be achieved using a single clock and latches. We also consider the impact of short delay paths (i.e. minimum delays), which can cause hold-time violations. Under conservative minimum delay assumptions, our latch-based optimization, operating on the routed design, provides a 5% performance improvement, on average, essentially for "free" (i.e. without any re-routing/delay padding). We show that short paths greatly hinder the ability of using latches for speed improvement, motivating further work to reduce their effects. Bill Teng, Jason Helge Anderson |
FPL | 2 |
| 2011 | FPGA glitch power analysis and reduction
Warren Wai-Kit Shum, Jason Helge Anderson |
ISLPED | 2 |
| 2010 | A PUF design for secure FPGA-based embedded systemsabstractThe concept of having an integrated circuit (IC) generate its own unique digital signature has broad application in areas such as embedded systems security, and IP/IC counter-piracy. Physically unclonable functions (PUFs) are circuits that compute a unique signature for a given IC based on the process variations inherent in the IC manufacturing process. This paper presents the first PUF design specifically targeted for field-programmable gate arrays (FPGAs). Our novel design makes use of the underlying FPGA architecture, and unlike prior published PUFs, the proposed PUF can be naturally embedded into a design's HDL, consuming very little area, and does not require the use of "hard macros" with fixed routing. Measured results on the Xilinx Virtex-5 65 nm FPGA demonstrate PUF signatures to be both unique and reliable under temperature variation. Jason Helge Anderson |
ASP-DAC | 1 |
| 2010 | FPGA power reduction by guarded evaluationabstractGuarded evaluation is a power reduction technique that in-volves identifying sub-circuits (within a larger circuit) whose inputs can be held constant (guarded) at specific times dur-ing circuit operation, thereby reducing switching activity and lowering dynamic power. The concept is rooted in the property that under certain conditions, some signals within digital designs are not “observable ” at design outputs, mak-ing the circuitry that generates such signals a candidate for guarding. Guarded evaluation has been demonstrated successfully for custom ASICs; in this paper, we apply the technique to FPGAs. In ASICs, guarded evaluation entails adding additional hardware to the design, increasing sili-con area and cost. Here, we apply the technique in a way that imposes minimal area overhead by leveraging existing unused circuitry within the FPGA. The primary challenge in guarded evaluation is in determining the specific condi-tions under which a sub-circuit’s inputs can be held con-stant without impacting the larger circuit’s functional cor-rectness. We propose a simple solution to this problem based on discovering “non-inverting paths ” in the circuit’s AND-inverter graph representation. Experimental results show that guarded evaluation can reduce switching activity by 22%, on average, and can reduce power consumption in the FPGA interconnect by 14%. Jason Helge Anderson, Chirag Ravishankar |
FPGA | 1 |
| 2010 | Parallelizing FPGA placement using Transactional MemoryabstractTo capitalize on the growing abundance of multicore hardware, FPGA vendors have begun to parallelize the most compute intensive algorithms in their CAD software. However, parallelization is a painstaking and hence expensive process that limits the number of algorithms that can be cost-effectively parallelized. Transactional Memory (TM) promises an easier-to-use alternative to locks for critical sections in threaded code-allowing programmers to avoid deadlocks and data races, and also allowing critical sections to execute in parallel as long as they dynamically access independent data. In this paper, we present our work on using TM to parallelize simulated annealing-based placement for FPGAs. In particular, we use a software TM (TinySTM) to parallelize the placement phase of Versatile Place and Route (VPR) 5.0.2. With TM we very quickly produced a parallel and correct version of the software, allowing us to focus on incrementally tuning performance. We describe our experiences in tuning the TM system and CAD software, and the interesting algorithmic trade-offs that exist. In the end, we found that optimized transactional placement has the potential for scalable performance: our non-deterministic implementation achieves self-relative speedups over a single thread of 1.82x, 3.62x and 7.27x at 2, 4, and 8 threads respectively with little quality degradation. However, hardware support for TM is likely required to overcome the overheads of STM, as our implementation's single thread performance is 8x slower than sequential VPR. Steven Birk, J. Gregory Steffan, Jason Helge Anderson |
FPT | 3 |
| 2010 | Deterministic multi-core parallel routing for FPGAsabstractWe consider coarse and fine-grained techniques for parallel FPGA routing on modern multi-core processors. In the coarse-grained approach, sets of design signals are assigned to different processor cores and routed concurrently. Communication between cores is through the MPI (message passing interface) communications protocol. In the fine-grained approach, the task of routing an individual load pin on a signal is parallelized using threads. Specifically, as FPGA routing resources are traversed during maze expansion, delay calculation, costing and priority queue insertion for these resources execute concurrently. The proposed techniques provide deterministic/repeatable results. Moreover, the coarse and fine-grained approaches are not mutually exclusive and can be used in tandem. Results show that on a 4-core processor, the techniques improve router run-time by ~2.1×, on average, with no significant impact on circuit speed performance or interconnect resource usage. Marcel Gort, Jason Helge Anderson |
FPT | 2 |
| 2009 | Emerging application domains: research challenges and opportunities for FPGAsabstractCommunications infrastructure, data processing and industrial electronics are the cornerstone application areas for programmable logic today. But what are the application domains of tomorrow? What nascent application areas could explode the growth of programmable logic usage and expand the programmable market? In this workshop, we will hear speakers from industry and academia talk about the emerging application areas for FPGAs and the challenges and opportunities in these areas. We will consider how programmable hardware and the associated tools should be enhanced to become better-suited to tomorrow's applications. The overarching aim of the workshop is to seed ideas in the research community by giving an applications perspective of the fertile topics for future research on FPGA architecture, CAD and applications. Jason Helge Anderson |
FPGA | 1 |
| 2009 | Clock power reduction for virtex-5 FPGAsabstractClock network power in field-programmable gate arrays (FPGAs) is considered and two complementary approaches for clock power reduction in the Xilinx Virtex-5 FPGA are presented. The approaches are unique in that they leverage specific architectural aspects of Virtex-5 to achieve reductions in dynamic power consumed by the clock network. The first approach comprises a placement-based technique to reduce interconnect resource usage on the clock network, thereby reducing capacitance and power (up to 12%). The second approach borrows the "clock gating" notion from the ASIC domain and applies it to FPGAs. Clock enable signals on flip-flops are selectively migrated to use the dedicated clock enable available on the FPGA's built-in clock network, leading to reduced toggling on the clock interconnect and lower power (up to 28%). Power reductions are achieved without any performance penalty, on average. Subodh Gupta, Jason Helge Anderson |
FPGA | 3 |
| 2009 | Improving logic density through synthesis-inspired architectureabstractWe leverage properties of the logic synthesis netlist to define both a logic element architecture and an associated technology mapping algorithm that together provide improved logic density. We demonstrate that an ldquoextendedrdquo logic element with slightly modified K-input LUTs achieves much of the benefit of an architecture with K+1-input LUTs, while consuming silicon area close to a K-LUT (a K-LUT requires half the area of a K+1-LUT).We introduce the notion of ldquonon-inverting pathsrdquo in a circuit's AND-inverter graph (AIG) and show their utility in mapping into the proposed logic element. Results show that while circuits mapped to a traditional 5-LUT architecture need 14% more LUTs and have 12% more depth than a 6-LUT architecture, our extended 5-LUT architecture requires only 7%more LUTs and 2.5% more depth than 6-LUTs, on average. Nearly all of the depth reduction associated with moving from K-input to K+1-input LUTs can be achieved with considerably less area using extended K-LUTs. We further show that 6-LUT optimal mapping depths can be achieved with a small fraction of the LUTs in hardware being 6-LUTs and the remainder being extended 5-LUTs, suggesting that a heterogeneous logic block architecture may prove to be advantageous. Jason Helge Anderson |
FPL | 1 |
| 2009 | Clock gating architectures for FPGA power reductionabstractClock gating is a power reduction technique that has been used successfully in the custom ASIC domain. Clock and logic signal power are saved by temporarily disabling the clock signal on registers whose outputs do not affect circuit outputs. We consider and evaluate FPGA clock network architectures with built-in clock gating capability and describe a flexible placement algorithm that can operate with various gating granularities (various sizes of device regions containing clock loads that can be gated together). Results show that depending on the clock gating architecture and the fraction of time clock signals are enabled, clock power can be reduced by over 50%, and results suggest that a fine granularity gating architecture yields significant power benefits. Safeen Huda, Muntasir Mallick, Jason Helge Anderson |
FPL | 3 |
| 2009 | Packing Techniques for Virtex-5 FPGAsabstractPacking is a key step in the FPGA tool flow that straddles the boundaries between synthesis, technology mapping and placement. Packing strongly influences circuit speed, density, and power, and in this article, we consider packing in the commercial FPGA context and examine the area and performance trade-offs associated with packing in a state-of-the-art FPGA---the Xilinx ® Virtex TM -5 FPGA. In addition to look-up-table (LUT)-based logic blocks, modern FPGAs also contain large IP blocks. We discuss packing techniques for both types of blocks. Virtex-5 logic blocks contain dual-output 6-input LUTs. Such LUTs can implement any single logic function of up to 6 inputs, or any two logic functions requiring no more than 5 distinct inputs. The second LUT output has reduced speed, and therefore, must be used judiciously. We present techniques for dual-output LUT packing that lead to improved area-efficiency, with minimal performance degradation. We then describe packing techniques for large IP blocks, namely, block RAMs and DSPs. We pack circuits into the large blocks in a way that leverages the unique block RAM and DSP layout/architecture in Virtex-5, achieving significantly improved design performance. Taneem Ahmed, Paul D. Kundarewich, Jason Helge Anderson |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2009 | Low-Power Programmable FPGA Routing CircuitryabstractWe consider circuit techniques for reducing field-programmable gate-array (FPGA) power consumption and propose a family of new FPGA routing switch designs that are programmable to operate in three different modes: high-speed, low-power, or sleep. High-speed mode provides similar power and performance to traditional FPGA routing switches. In low-power mode, speed is curtailed in order to reduce power consumption. Leakage is reduced by 28%-52% in low-power versus high-speed mode, depending on the particular switch design selected. Dynamic power is reduced by 28%-31% in low-power mode. Leakage power in sleep mode, which is suitable for unused routing switches, is 61%-79% lower than in high-speed mode. Each of the proposed switch designs has a different power/area/speed tradeoff. All of the designs require only minor changes to a traditional routing switch and involve relatively small area overhead, making them easy to incorporate into current commercial FPGAs. The applicability of the new switches is motivated through an analysis of timing slack in industrial FPGA designs. It is observed that a considerable fraction of routing switches may be slowed down (operate in low-power mode), without impacting overall design performance. Jason Helge Anderson, Farid N. Najm |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2008 | Architecture-specific packing for virtex-5 FPGAsabstractWe consider packing in the commercial FPGA context and examine the speed, performance and power trade-offs associated with packing in a state-of-the art FPGA -- the Xilinx Virtex-5 FPGA. Two aspects of packing are discussed: 1)packing for general logic blocks, and 2 packing for large IP blocks. Virtex-5 logic blocks contain dual-output 6-input look-up-tables (LUTs). Such LUTs can implement any single logic function requiring no more than 6 inputs, or any two logic functions requiring no more than 5 distinct inputs. The second LUT output is associated with slower speed, and therefore, must be used judiciously. We present placement-based techniques for dual-output LUT packing that lead to improved area-efficiency and power, with minimal performance degradation. We then move on to address packing for large IP blocks, specifically, block RAMs and DSPs. We present a packing optimization that is widely applicable in DSP designs that leads to significantly improved design performance Taneem Ahmed, Paul D. Kundarewich, Jason Helge Anderson, Brad L. Taylor, Rajat Aggarwal |
FPGA | 3 |
| 2006 | Active leakage power optimization for FPGAsabstractActive leakage power dissipation is considered in field-programmable gate arrays (FPGAs) and two "no cost" approaches for active leakage reduction are presented. It is well known that the leakage power consumed by a digital CMOS circuit depends strongly on the state of its inputs. The authors' first leakage reduction technique leverages a fundamental property of basic FPGA logic elements [look-up tables (LUTs)] that allows a logic signal in an FPGA design to be interchanged with its complemented form without any area or delay penalty. This property is applied to select polarities for logic signals so that FPGA hardware structures spend the majority of time in low-leakage states. In an experimental study, active leakage power is optimized in circuits mapped into a state-of-the-art 90-nm commercial FPGA. Results show that the proposed approach reduces active leakage by 25%, on average. The authors' second approach to leakage optimization consists of altering the routing step of the FPGA computer-aided design (CAD) flow to encourage more frequent use of routing resources that have low leakage power consumptions. Such "leakage-aware routing" allows active leakage to be further reduced, without compromising design performance. Combined, the two approaches offer a total active leakage power reduction of 30%, on average. Jason Helge Anderson, Farid N. Najm |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2004 | Interconnect capacitance estimation for FPGAs
Jason Helge Anderson, Farid N. Najm |
ASP-DAC | 1 |
| 2004 | Active leakage power optimization for FPGAsabstractWe consider active leakage power dissipation in FPGAs and present a "no cost" approach for active leakage reduction. It is well-known that the leakage power consumed by a digital CMOS circuit depends strongly on the state of its inputs. Our leakage reduction technique leverages a fundamental property of basic FPGA logic elements (look-up-tables) that allows a logic signal in an FPGA design to be interchanged with its complemented form without any area or delay penalty. We apply this property to select polarities for logic signals so that FPGA hardware structures spend the majority of time in low leakage states. In an experimental study, we optimize active leakage power in circuits mapped into a state-of-the-art 90nm commercial FPGA. Results show that the proposed approach reduces active leakage by 25%, on average. Jason Helge Anderson, Farid N. Najm, Tim Tuan |
FPGA | 1 |
| 2004 | Run-Time-Conscious Automatic Timing-Driven FPGA Layout Synthesis
Jason Helge Anderson, Sudip Nag, Kamal Chaudhary, Sandor Kalman, Chari Madabhushi, Paul Cheng |
FPL | 1 |
| 2004 | Low-power programmable routing circuitry for FPGAsabstractWe propose two new FPGA routing switch designs that are programmable to operate in three different modes: high-speed, low-power or sleep. High-speed mode provides similar power and performance to a traditional routing switch. In low-power mode, speed is curtailed in order to reduce power consumption. Our first switch design reduces leakage power consumption by 36-40% in low-power vs. high-speed mode (on average); dynamic power is reduced by up to 28%. Leakage power in sleep mode is 61% lower than in high-speed mode. A second switch design offers a 36% smaller area overhead and reduces leakage by 28-30% in low-power vs. high-speed mode. The proposed switch designs require only minor changes to a traditional routing switch, making them easy to incorporate into current FPGA interconnect. The applicability of the new switches is motivated through an analysis of timing slack in industrial FPGA designs. Specifically, we show that a considerable fraction of routing switches may be slowed down (operate in low-power mode), without impacting overall design performance. Jason Helge Anderson, Farid N. Najm |
ICCAD | 1 |
| 2004 | Power estimation techniques for FPGAsabstractThe dynamic power consumed by a digital CMOS circuit is directly proportional to both switching activity and interconnect capacitance. In this paper, we consider early prediction of net activity and interconnect capacitance in field-programmable gate array (FPGA) designs. We develop empirical prediction models for these parameters, suitable for use in power-aware layout synthesis, early power estimation/planning, and other applications. We examine how switching activity on a net changes when delays are zero (zero delay activity) versus when logic delays are considered (logic delay activity) versus when both logic and routing delays are considered (routed delay activity). We then describe a novel approach for prelayout activity prediction that estimates a net's routed delay activity using only zero or logic delay activity values, along with structural and functional circuit properties. For capacitance prediction, we show that prediction accuracy is improved by considering aspects of the FPGA interconnect architecture in addition to generic parameters, such as net fanout and bounding box perimeter length. We also demonstrate that there is an inherent variability (noise) in the switching activity and capacitance of nets that limits the accuracy attainable in prediction. Experimental results show the proposed prediction models work well given the noise limitations. Jason Helge Anderson, Farid N. Najm |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2002 | Power-aware technology mapping for LUT-based FPGAsabstractWe present a new power-aware technology mapping technique for LUT-based FPGAs which aims to keep nets with high switching activity out of the FPGA routing network and takes an activity-conscious approach to logic replication. Logic replication is known to be crucial for optimizing depth in technology mapping; an important contribution of our work is to recognize the effect of logic replication on circuit structure and to show its consequences on power. In an experimental study, we examine the power characteristics of mapping solutions generated by several publicly available technology mappers. Results show that for a specific depth of mapping solution, the power consumption can vary considerably, depending on the technology mapping approach used. Furthermore, results show that our proposed mapping algorithm leads to circuits with substantially less power dissipation than previous approaches. Jason Helge Anderson, Farid N. Najm |
FPT | 1 |
| 1998 | Technology Mapping for Large Complex PLDsabstractIn this paper we present a new technology mapping algorithm for use with complex PLDs (CPLDs), which consists of a large number of PLA-style logic blocks. Although the traditional synthesis approach for such devices uses two-level minimization, the complexity of recently-produced CPLDs has resulted in a trend toward multi-level synthesis. We describe an approach that allows existing multi-level synthesis techniques [13] to be adapted to produce circuits that are well-suited for implementation in CPLDs. Our algorithm produces circuits that require up to 90% fewer logic blocks than the circuits produced by a recently-published algorithm. Jason Helge Anderson, Stephen Brown 0003 |
DAC | 1 |
| 1998 | An LPGA with Foldable PLA-style Logic BlocksabstractLaser-programmed gate arrays (LPGAs) represent a new approach to application specific integrated circuit prototyping and implementation. This paper proposes a new LPGA logic block architecture called a foldable PLA-style logic block. The proposed logic block architecture is similar to that found in commercially available CPLDs. The term foldable means that the granularity of the logic block can be varied. This is achieved using the LPGA laser disconnect methodology. A custom CAD tool has been developed to map circuits into the new logic block architecture. An experimental study shows that LPGAs with foldable logic blocks are more area-efficient than those based on normal unfoldable logic blocks. Jason Helge Anderson, Stephen Brown 0003 |
FPGA | 1 |