EDBT 2026 Demo / reviewers in the wild / expert
Brad L. Hutchings
dblp:h/BradLHutchings
· DBLP profile ↗
59ranked-venue papers
11as first author
4since 2021 · last 2023
0000-0002-2991-0230ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 56 · 10 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Improving the Reliability of FPGA CRO PUFsabstractThis paper presents a novel technique that greatly improves the reliability of FPGA-based CRO PUFs. We improve upon existing CRO implementations and increase the number of configurations per CLB tile from 16 384 to 1.1 × 1012• To maximize reliability, each CRO pair must be configured to maximize its frequency difference. This requires using a novel technique that reduces the configuration search space from 1.1 × 1012to 256. Our CRO PUF achieves 100% reliability within the FPGA's maximum rated voltages. We believe that this is the first FPGA PUF that can achieve this level of reliability without the use of post-processing. We also show that in some cases, our CRO may be reliable enough to omit the ECC that is usually required in PUF-based key generation circuits. This allows our CRO PUF to provide the reliability required for key generation while reducing the latency, complexity, and area overhead of ECC algorithms. Hayden Cook, Zephram Tripp, Brad L. Hutchings, Jeffrey B. Goeders |
FPL | 3 |
| 2022 | Cloning the Unclonable: Physically Cloning an FPGA Ring-Oscillator PUFabstractThis work presents a novel technique to physically clone a ring oscillator physically unclonable function (RO PDF) onto another distinct FPG A die, using precise, targeted aging. The resulting cloned RO PDF provides a response that is identical to its copied FPGA counterpart, i.e., the FPGA and its clone are indistinguishable from each other. Targeted aging is achieved by: 1) heating the FPGA using bitstream-Iocated short circuits, and 2) enabling/disabling ROs in the same FPGA bitstream. During self heating caused by short-circuits contained in the FPGA bitstream, circuit areas containing oscillating ROs (enabled) degrade more slowly than circuit areas containing non-oscillating ROs (disabled), due to bias temperature instability effects. This targeted aging technique is used to swap the relative frequencies of two ROs that will, in turn, flip the corresponding bit in the PUF response. Two experiments are described. The first experiment uses targeted aging to create an FPGA that exhibits the same PUF response as another FPGA, i.e., a clone of an FPGA PUF onto another FPGA device. The second experiment demonstrates that this aging technique can create an RO PUF with any desired response. Hayden Cook, Jonathan Thompson, Zephram Tripp, Brad L. Hutchings, Jeffrey B. Goeders |
FPT | 4 |
| 2022 | Approaches for FPGA Design AssuranceabstractField-Programmable Gate Arrays (FPGAs) are widely used for custom hardware implementations, including in many security-sensitive industries, such as defense, communications, transportation, medical, and more. Compiling source hardware descriptions to FPGA bitstreams requires the use of complex computer-aided design (CAD) tools. These tools are typically proprietary and closed-source, and it is not possible to easily determine that the produced bitstream is equivalent to the source design. In this work, we present various FPGA design flows that leverage pre-synthesizing or pre-implementing parts of the design, combined with open-source synthesis tools, bitstream-to-netlist tools, and commercial equivalence-checking tools, to verify that a produced hardware design is equivalent to the designer’s source design. We evaluate these different design flows on several benchmark circuits and demonstrate that they are effective at detecting malicious modifications made to the design during compilation. We compare our proposed design flows with baseline commercial design flows and measure the overheads to area and runtime. Eli Cahill, Brad L. Hutchings, Jeffrey B. Goeders |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2022 | Inducing Non-uniform FPGA Aging Using Configuration-based Short CircuitsabstractThis work demonstrates a novel method of accelerating FPGA aging by configuring FPGAs to implement thousands of short circuits, resulting in high on-chip currents and temperatures. Patterns of ring oscillators are placed across the chip and are used to characterize the operating frequency of the FPGA fabric. Over the course of several months of running the short circuits on two-thirds of the reconfigurable fabric, with daily characterization of the FPGA 6 performance, we demonstrate a decrease in FPGA frequency of 8.5%. We demonstrate that this aging is induced in a non-uniform manner. The maximum slowdown outside of the shorted regions is 2.1%, or about a fourth of the maximum slowdown that is experienced inside the shorted region. In addition, we demonstrate that the slowdown is linear after the first two weeks of the experiment and is unaffected by a recovery period. Additional experiments involving short circuits are also performed to demonstrate the results of our initial experiments are repeatable. These experiments also use a more fine-grained characterization method that provides further insight into the non-uniformed nature of the aging caused by short circuits. Hayden Cook, Jacob Arscott, Brent George, Tanner Gaskin, Jeffrey B. Goeders, Brad L. Hutchings |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2020 | Using Novel Configuration Techniques for Accelerated FPGA AgingabstractIn this work we demonstrate a novel method of accelerating FPGA aging by configuring the FPGA to implement thousands of short circuits, resulting in high on-chip currents and temperatures. Three ring oscillators are placed across the chip and are used to characterize the operating frequency of the FPGA fabric. Over the course of several weeks of running the short circuits, with daily characterization of the FPGA performance, we measured a decrease in FPGA frequency greater than 5%. After aging, the FPGA part was repeatedly characterized during a two week idle period. Results indicated that the slowdown did not change, and the aging appeared to be permanent. In addition, we demonstrated that this aging could be induced in a non-uniform manner. In our experiments, the short circuits were all placed in the lower two-thirds of the chip, and one of the characterization ring oscillators was placed at the top of the chip, outside of the region with the short circuits. The fabric at this location exhibited a 1.36% slowdown, only one-quarter the slowdown measured in the targeted region. Tanner Gaskin, Hayden Cook, Wesley Stirk, Robert Lucas, Jeffrey B. Goeders, Brad L. Hutchings |
FPL | 6 |
| 2019 | Preallocating Resources for Distributed Memory Based FPGA DebugabstractMost internal FPGA debug methods require the use of Block-RAM (BRAM) memory for trace buffers. Recent work has shown the viability of replacing BRAMs with distributed, LUT based memory. Distributed memory (DIME) trace buffers are lean and can be utilized in large designs where other debug methods are unlikely to fit. Since LUTs are abundant on FPGA devices, there are nearly always some left unused after the user's design is placed, even for designs that utilize more than 90% of the FPGA's resources. DIME trace buffers are inserted into highly utilized designs within minutes using RapidWright. In this paper we contrast the previously used method of scavenging leftover LUT resources with a preallocation scheme that ensures a certain amount of memory LUTs are left available for distributed memory trace buffers. While causing virtually no penalty to the user design, preallocating memory LUT resources allows the very largest designs to utilize higher numbers of distributed memory trace buffers at lower timing penalties. We also show that depth of DIME trace buffers can be extended from 16 to 256 bits. Robert Hale, Brad L. Hutchings |
FPL | 2 |
| 2018 | Demand Driven Assembly of FPGA Configurations Using Partial Reconfiguration, Ubuntu Linux, and PYNQabstractThe PYNQ system (Python Productivity for Zynq) is notable for combining a monolithic preconfigured bitstream, Ubuntu Linux, Python, and Jupyter notebooks to form an FPGA-based system that is far more accessible to non-FPGA experts than previous systems. In this work, the monolithic pre-configured PYNQ bitstream is replaced with a combination of a simple base bitstream containing several partial reconfiguration regions and a library of partial bitstreams that implement a variety of hardware interfaces such as: GPIO, UART, Timer, IIC, SPI, Real-Time Clock, etc., that interface to various Pmod-based peripherals. When peripherals are plugged into a Pmod socket at run-time, corresponding partial reconfigurations and standard device drivers can be automatically loaded into the Ubuntu kernel using device-tree overlays. This demand-driven, partially-reconfigured approach is found to be advantageous to the monolithic bitstream because: 1) it provides similar functionality to the monolithic bitstream while consuming less area, 2) it provides a way for users to modify or augment hardware functionality without requiring the user to develop a new monolithic bitstream, 3) run-time demand loading of partial bitstreams makes the system more responsive to changing conditions, and 4) implementation issues such as timing-closure, etc., are simplified because the base bitstream circuitry is smaller and less complex. Jeffrey B. Goeders, Tanner Gaskin, Brad L. Hutchings |
FCCM | 3 |
| 2018 | Enabling Low Impact, Rapid Debug for Highly Utilized FPGA DesignsabstractInserting soft logic analyzers into FPGA circuits is a common way to provide signal visibility at run-time, helping users locate bugs in their designs. However, this can become infeasible for highly (70-90+%) utilized designs, which leave few logic resources or block RAMs available for internal logic analyzers. This paper presents a fast, low-impact method of enabling signal visibility in these situations using LUT-based distributed memory. Trace-buffers are inserted post-PAR allowing users to quickly change the set of observed nets. Results from routing-based experiments are presented which demonstrate that, even in highly utilized designs, many design signals can be observed with this technique. Robert Hale, Brad L. Hutchings |
FPL | 2 |
| 2018 | Distributed-Memory Based FPGA Debug: Design Timing ImpactabstractIn FPGAs, debug observability is often achievedby attaching memory-based recording circuitry to user signals. Block-RAM (BRAM)-based embedded logic analyzers are ofteninserted into user circuits to observe circuit behavior. Incontrast with BRAM-based approaches, distributed memory:1) is almost always available (user circuits may consume allBRAMs but even highly utilized circuits contain unused LUTs), and 2) can usually be physically located very near to user signals(LUTs are spread across the entire device while BRAMs arelocated only in specific columns). Previous work has shownbasic feasibility and demonstrated that distributed memoriescan provide debug observability for highly utilized circuits. Thispaper focuses on timing impacts and describes the quantitativetradeoff between FPGA device utilization, debug probe count, and clock frequency. For example, a design with 70% of LUTsutilized, with no debug logic, can operate at a minimum clockperiod of 5ns. Instrumenting 300 debug probes increases thisperiod to 7ns, and 1500 probes to 8ns. Placing trace bufferswith a simulated annealing algorithm improved success ratesfrom 20% to 50% depending on the design and probe count. Robert Hale, Brad L. Hutchings |
FPT | 2 |
| 2018 | Enhancing debug observability for HLS-based FPGA circuits through source-to-source compilation
Joshua S. Monson, Brad L. Hutchings |
J. Parallel Distributed Comput. | 2 |
| 2017 | Rapid implementation of a partially reconfigurable video system with PYNQabstractUndergraduate students rapidly implement a partially-reconfigured, real-time video processor on the Xilinx PYNQ board. The video processor performs various real-time operations including Sobel edge detection, embossing, averaging, an interactive Pong game, etc., using a separate partially-reconfigurable bit-stream for each distinct function. Selection of image-processing functions is accomplished by a Python-based graphical user interface that is accessed via a Jupyter notebook. As users select image-processing functions the appropriate partial bit-stream is automatically downloaded to the FPGA. The resulting system is easily and quickly developed by several undergraduate students over a period of about 10 weeks with very little supervision. All files related to the project are available for download on GitHub. The productivity benefits provided by PYNQ, including Jupyter-based documentation, tutorials, and executable Python code greatly ease development effort making PYNQ an excellent FPGA platform for education. Brad L. Hutchings, Michael J. Wirthlin |
FPL | 1 |
| 2015 | RapidSmith 2: A Framework for BEL-level CAD Exploration on Xilinx FPGAsabstractRapidSmith is an open-source framework that allows for the exploration of novel approaches to the FPGA CAD flow for Xilinx devices. However, RapidSmith has poor support for manipulating designs below the slice level. In this paper, we highlight many of the projects RapidSmith enables and present extensions incorporated into "RapidSmith 2" that expose LUTs and flip-flops for direct manipulation in custom-built CAD tools. To demonstrate the utility of RapidSmith 2 we present the results of work to identify BELs in a design which must be clustered together and a tool that does pre-packing clustering accordingly. Travis Haroldsen, Brent E. Nelson, Brad L. Hutchings |
FPGA | 3 |
| 2015 | Using Source-Level Transformations to Improve High-Level Synthesis Debug and Validation on FPGAsabstractThis paper proposes a method for extending source-level visibility into the RTL of an HLS-generated design using automated source-level transformations. Using our method, source-level visibility can be extended into co-simulation, in-system simulation, and hardware execution of any HLS tool that provides the ability to infer top-level ports. Experimental results show the feasibility of our method in situations where visibility needs to be added without modifying the timing, latency, or throughput of the design. Joshua S. Monson, Brad L. Hutchings |
FPGA | 2 |
| 2015 | Using source-to-source compilation to instrument circuits for debug with High Level SynthesisabstractC-based High Level Synthesis (HLS)-compatible circuit descriptions from the CHStone benchmark suite are instrumented for debugging purposes using a source-to-source compiler. The debug instrumentation connects C expressions to top-level ports that can be observed during the debugging process. Approximately 50,000 different experiments are conducted to determine the impact on the final circuit caused by the debug instrumentation. Experimental data indicate initial feasibility of the instrumentation approach; all assignment expressions in a program can be instrumented for an average increase in LUT count of about 24%. Increases in FF count and clock period were in the range of 5% to 10%. Joshua S. Monson, Brad L. Hutchings |
FPT | 2 |
| 2014 | Rapid Post-Map Insertion of Embedded Logic Analyzers for Xilinx FPGAsabstractA rapid post-map insertion of an embedded logic analyzer is discussed. The proposed technique makes use of otherwise unused resources in an already-mapped circuit and does not disturb the original placement and routing of the circuit. Using this technique, designers can add debugging circuitry to existing circuits and quickly modify the set of of observed signals in just a few minutes instead of waiting for a recompile of their circuit. All tests were performed on a Xilinx Virtex-5 FPGA. Brad L. Hutchings, Jared Keeley |
FCCM | 1 |
| 2014 | A power side-channel-based digital to analog converterfor Xilinx FPGAsabstractA novel Digital to Analog Converter (DAC) modulates the overall power consumption of an FPGA by disabling/enabling short circuits programmed into the interconnect. The power pin of the FPGA serves as the output of the DAC. The DAC achieves high linearity and can be used to implement applications in communications, security, etc. The shortcircuit-based DAC consumes 1/3 the area of an alternative shift-register-based DAC that is presented for the sake of comparison. Brad L. Hutchings, Joshua S. Monson, Danny Savory, Jared Keeley |
FPGA | 1 |
| 2014 | New approaches for in-system debug of behaviorally-synthesized FPGA circuitsabstractThis paper present new approaches for in-system, trace-based debug of High-Level Synthesis-generated hardware. These approaches include the use of Event Observability Ports (EOP) that provide observability of source-level events in the final hardware. We also propose the use of small, independent trace buffers called Event Observability Buffers (EOB) for tracing events through EOPs. EOBs include a data storage enable signal that allows cycle-by-cycle storage decisions to be made on an EOB-by-EOB basis. This approach causes the timing relationships of events captured in different trace buffers to be lost. Two methods are presented for recovering these relationships. Finally, we present a case study that demonstrates the feasibility and effectiveness of an EOB trace strategy. Joshua S. Monson, Brad L. Hutchings |
FPL | 2 |
| 2013 | Implementing high-performance, low-power FPGA-based optical flow accelerators in CabstractRecent developments in High-Level Synthesis (HLS) for FPGAs are making it possible to “run” C code on FPGAs thereby making modern programming environments available to FPGA developers. In this paper, C code for a complex optical-flow algorithm is optimized for both a desktop PC and for an FPGA-based system, the Xilinx Zynq-7000, a device containing both a programmable fabric and two ARM cores. The paper discusses how the code is optimized and restructured to execute effectively on the programmable fabric and the ARM cores. The resulting Zynq version of the C code is competitive with the desktop PC but only consumes 1/7th as much energy. Joshua S. Monson, Michael J. Wirthlin, Brad L. Hutchings |
ASAP | 3 |
| 2013 | Impact of hard macro size on FPGA clock rate and place/route timeabstractHard macros are completely placed/routed elements that are treated as primitives and that are relatively placed as a single element. A system composed of such macros consists of many fewer effective primitives and nets and as such can be placed and routed much more quickly. Prior work in this research area dealt with small, general-purpose macros such as 16-bit registers, adders, etc., and demonstrated that place/route time could be reduced by an order of magnitude with a corresponding 3-4X reduction in clock rate. In this work, much larger hard macros are developed such as mixers, softcore processors, FFTs, etc., and the use of these larger macros is shown to further reduce place/route time by an additional 2.5-4X, for a total of a 30-40X reduction in compile time. Clock rate is also improved, relative to earlier work, by an additional 60-70%. Chris Lavin, Brent E. Nelson, Brad L. Hutchings |
FPL | 3 |
| 2013 | Improving clock-rate of hard-macro designsabstractHMFlow reuses precompiled circuit modules (hard macros) and other techniques to rapidly compile large designs in a few seconds - many times faster than standard Xilinx flows. However, the clock rates of designs rapidly compiled by HMFlow are often significantly lower than those compiled by the Xilinx flow. To improve clock rates, HMFlow algorithms were modified as follows: (1) the router was modified to take advantage of longer routing wires in the FPGA devices, (2) the original greedy placer was replaced with an annealing-based placer, and (3) certain registers were removed from the hard-macro and moved into the fabric to reduce critical-path delays. Benchmark circuits compiled with these modifications can achieve clock rates that are about 75% as fast as those achieved by Xilinx, on average. Fast run-times are also preserved; the improved algorithms only increase HMFlow run-times by about 50% across the benchmark suite so that HMFlow remains more than 30× faster than the standard Xilinx flow for the benchmarks tested in this paper. Chris Lavin, Brent E. Nelson, Brad L. Hutchings |
FPT | 3 |
| 2012 | Profiling FPGA floor-planning effects on timing closureabstractThe impact of shape, area allocation and timing constraints on partitions was determined by selecting a standard set of submodules and performing over 1,000,000 place/route experiments. Place/route experiments used different area and timing constraints and their resulting trace reports provided timing results. These results suggest that the best results are obtained when about 20% additional area (above synthesis estimates) is allocated for each submodule. The aspect ratio of submodules is largely a non-issue (there was one exception in the data). In some cases, carefully constraining area dramatically improves results. Jaren Lamprecht, Brad L. Hutchings |
FPL | 2 |
| 2011 | HMFlow: Accelerating FPGA Compilation with Hard Macros for Rapid PrototypingabstractThe FPGA compilation process (synthesis, map, place, and route) is a time consuming task that severely limits designer productivity. Compilation time can be reduced by saving implementation data in the form of hard macros. Hard macros consist of previously synthesized, placed and routed circuits that enable rapid design assembly because of the native FPGA circuitry (primitives and nets)which they encapsulate. This work presents results from creating a new FPGA design flow based on hard macros called HMF low. HMF low has shown speedups of 10-50X over the fastest configuration of the Xilinx tools. Designed for rapid prototyping, HMF low achieves these speedups by only utilizing up to 50 percent of the resources on an FPGA and produces implementations that run 2-4X slower than those produced by Xilinx. These speedups are obtained on a wide range of benchmark designs with some exceeding 18,000 slices on a Virtex 4 LX200. Chris Lavin, Marc Padilla, Jaren Lamprecht, Philip Lundrigan, Brent E. Nelson, Brad L. Hutchings |
FCCM | 6 |
| 2011 | FPGA Communication FrameworkabstractFPGA-CF is an open-source, portable, extensible communications package that consists of a small hardware core (less than 600 slices) and and a host-software library/API. It enables a host PC to transmit data at 120 Mb/s to Xilinx-based FPGA boards via Ethernet using standard internet protocols. The hardware core is directly connected to the Xilinx internal configuration port (ICAP) and supports all ICAP functionality. The core also provides an extensible user-channel interface and provides up to 15, 8-bit user-data channels. The host software API supports both Java and C++ and provides high-level functionality for making connections and transmitting data. The utility of the system is demonstrated by implementing an on-chip test/debug system. Peter Lieber, Brad L. Hutchings |
FCCM | 2 |
| 2011 | RapidSmith: Do-It-Yourself CAD Tools for Xilinx FPGAsabstractCreating CAD tools for commercial FPGAs is a difficult task. Closed proprietary device databases and unsupported interfaces are largely to blame for the lack of CAD research found on commercial architectures versus hypothetical architectures. This paper formally introduces RapidSmith, a new set of tools and APIs that enable CAD tool creation for Xilinx FPGAs. Based on the Xilinx Design Language (XDL), RapidSmith provides a compact, yet, fast device database with hundreds of APIs that enable the creation of placers, routers and several other tools for Xilinx devices. RapidSmith alleviates several of the difficulties of using XDL and this work demonstrates the kinds of research facilitated by removing such challenges. Chris Lavin, Marc Padilla, Jaren Lamprecht, Philip Lundrigan, Brent E. Nelson, Brad L. Hutchings |
FPL | 6 |
| 2010 | Using Hard Macros to Reduce FPGA Compilation TimeabstractThe FPGA compilation process (synthesis, map, placement, routing) is a time-consuming process that limits designer productivity. Compilation time can be reduced by using pre-compiled circuit blocks (hard macros). Hard macros consist of previously synthesized, mapped, placed and routed circuitry that can be relatively placed with short tool runtimes and that make it possible to reuse previous computational effort. Two experiments were performed to demonstrate feasibility that hard macros can reduce compilation time. These experiments demonstrated that an augmented Xilinx flow designed specifically to support hard macros can reduce overall compilation time by 3x. Though the process of incorporating hard macros in designs is currently manual and error-prone, it can be automated to create compilation flows with much lower compilation time. Chris Lavin, Marc Padilla, Subhrashankha Ghosh, Brent E. Nelson, Brad L. Hutchings, Michael J. Wirthlin |
FPL | 5 |
| 2010 | Rapid prototyping tools for FPGA designs: RapidSmithabstractDesigner productivity for FPGA design is significantly limited by the time-consuming nature of the FPGA compilation process (synthesis, map, placement, and routing). However, experimentation on alternative CAD tools for this purpose for Xilinx devices has been somewhat limited. This paper describes the development and distribution of RapidSmith, a software library to facilitate the manipulation of XDL designs and upon which a complete CAD system can be based. The demonstration portion of this paper will show prototypes of representative CAD tools which can be easily built on top of the RapidSmith system. Chris Lavin, Marc Padilla, Philip Lundrigan, Brent E. Nelson, Brad L. Hutchings |
FPT | 5 |
| 2009 | Optical Flow on the Ambric Massively Parallel Processor Array (MPPA)abstractThe Ambric Massively Parallel Processor Array (MPPA) is a device that contains 336 32-bit RISC processors and is appropriate for embedded systems due to its relatively small physical and power footprint. Optical flow is a computationally-demanding and highly parallelizeable image-processing algorithm with applications in embedded systems such as robotics and autonomous vehicles. An optical flow algorithm is implemented on the Ambric device and is shown to achieve near FPGA performance at similar levels of power consumption while requiring many fewer lines of code (Java) than its FPGA counterpart (VHDL). Brad L. Hutchings, Brent E. Nelson, Stephen West, Reed Curtis |
FCCM | 1 |
| 2009 | Comparing fine-grained performance on the Ambric MPPA against an FPGAabstractA simple image-processing application is implemented on the Ambric MPPA and an FPGA, using a similar implementation for both devices. FPGAs perform extremely well on this kind of application and provide a good benchmark for comparison. The Ambric implementation starts out with a naive implementation and proceeds through several design optimizations until it reaches a maximum frame rate of 164 FPS (512 times 512 images) which turns out to be approximately 7times slower than the FPGA. The final Ambric implementation uses only 18 of 336 available processors, achieves more than sufficient performance for realtime embedded applications, and has excess processors to use for implementing additional algorithms. After introducing the image processing application and its implementation on both devices, the paper compares and contrasts the intrinsic, general characteristics of Ambric MPPA and FPGA devices. Brad L. Hutchings, Brent E. Nelson, Stephen West, Reed Curtis |
FPL | 1 |
| 2004 | What is the right model for programming and using modern FPGAs?abstractTraditionally, FPGAs have been the bastard step-brother of ASICs. They have been forced to act like ASICs and fit themselves into the ASIC development model. This has meant ignoring their unique strengths: reprogrammability, late-binding and run-time reconfiguration. Today, however, FPGAs are becoming more acceptable for their own merits. The majority of new design starts are FPGA designs. As FPGAs rise from under the shadow of their aging brother, should they continue to try to wear his hand-me-downs? Or is it time to develop more suitable models that lets them shine? At the same time, the old ASIC model is not even serving ASICs well, and new models for developing ASICs are emerging. All of this may encourage us to rethink how we should be programming FPGA-based systems. Possibilities include: André DeHon, Brad L. Hutchings, Daryl Rudusky, Nikhil, Salil Raje, Adrian Stoica |
FPGA | 2 |
| 2003 | Issues in debugging highly parallel FPGA-based applications derived from source codeabstractUsing high-level synthesis tools to map programs written in general-purpose languages to FPGA hardware has grown in popularity and it is becoming necessary to provide comprehensive debugging tools in order to verify the correctness of the synthesized hardware. Currently, post-synthesis debugging is done at the circuit level. This paper discusses the issues, as well as some early results, of creating a source level debugger for hardware synthesized from source code. This study is meant to provide some insight into what needs to be added or built into synthesizing compilers in order to allow debug of a synthesized circuit at the source level, which will provide the programmer with a familiar view of the program being debugged. Karl S. Hemmert, Brad L. Hutchings |
ASP-DAC | 2 |
| 2003 | Adaptive computing: what can it do, where can it go?abstractThe Adaptive Computing Systems (ACS) program was initiated by Defense Advanced Research Projects Agency (DARPA) of the United States in 1996. With the advent of FPGAs, has emerged a new class of computing systems that contain configurable hardware. This session begins by a presentation by the first ACS program manager of its motivation, original goals, and objectives. It is then followed by presentations of four specific projects under the ACS program. Future activities surrounding the ACS community will be discussed at the end. Robert Reuss, Jose L. Muñoz, Toshiaki Miyazaki, Nader Bagherzadeh, Prithviraj Banerjee, Brad L. Hutchings, Brian Schott |
ASP-DAC | 6 |
| 2003 | Source Level Debugger for the Sea Cucumber Synthesizing CompilerabstractWith the growing popularity of using high-level synthesis tools to map programs written in general-purpose programming languages to FPGA (field programmable gate array) hardware, it has become necessary to provide comprehensive, intuitive debugging tools in order to verify the correctness of the synthesized hardware. The difficulty in creating these tools lies in the fact that typical synthesizing compilers provide no information about how the source code is mapped to hardware. This paper discusses the creation of a debugger for the Sea Cucumber synthesizing compiler used to explore the issues associated with providing information about a circuit in the context of the original source code, thus making the debugging process more intuitive. Karl S. Hemmert, Justin L. Tripp, Brad L. Hutchings, Preston A. Jackson |
FCCM | 3 |
| 2003 | Simulation and Synthesis of CSP-based Interprocess CommunicationabstractThe Sea Cucumber project synthesizes parallel software threads into hardware on FPGAs (field programmable gate arrays). A communication system is needed which can be used to communicate between parallel software threads, and which can be synthesized into hardware with matching behavior. This paper describes a satisfactory communication system and presents three main contributions. First, it proposes a concise Java API for CSP (communication sequential processes) communication. Second, it presents a general hardware solution, which establishes interfaces and protocols for a various hardware implementations. Finally, it describes a hardware implementation created for the Xilinx Virtex II FPGA for performance analysis. Preston A. Jackson, Brad L. Hutchings, Justin L. Tripp |
FCCM | 2 |
| 2003 | Reconfigurable Computing Application FrameworksabstractFPGA-based (field programmable gate array) configurable computing machines (CCMs) offer powerful and flexible general-purpose computing platforms. However, development for FPGA-based designs using modern CAD (computer aided design) tools is geared mainly toward an ASIC-like process. This is inadequate for the needs of CCM application development. This paper discusses an application framework for developing CCM-based applications beyond just the hardware configuration. This framework leverages the advantages of CCMs (availability, programmability, visibility, and controllability) to help create CCM-based applications throughout the entire development process (i.e. design, debug, and deploy). The framework itself is deployed with the final application, thus permitting dynamic circuit configurations that include data folding optimizations based on user input. The resulting system aids in creating applications that are potentially more intuitive, easier to develop, and better performing. An example application demonstrates the use of the application framework and the potential benefits. Anthony L. Slade, Brent E. Nelson, Brad L. Hutchings |
FCCM | 3 |
| 2002 | Assisting Network Intrusion Detection with Reconfigurable HardwareabstractString matching is used by Network Intrusion Detection Systems (NIDS) to inspect incoming packet payloads for hostile data. String-matching speed is often the main factor limiting NIDS performance. String-matching performance can be dramatically improved by using Field-Programmable Gate Arrays (FPGAs); accordingly, a "regular-expression to FPGA circuit" module generator has been developed. The module generator extracts strings from the Snort NIDS rule-set, generates a regular expression that matches all extracted strings, synthesizes a FPGA-based string matching circuit, and generates an EDIF netlist that can be processed by Xilinx software to create an FPGA bitstream. The feasibility of this approach is demonstrated by comparing the performance of the FPGA-based string matcher against the software-based GNU regex program. The FPGA-based string matcher exceeds the performance of the software-based system by 600x for large patterns. Brad L. Hutchings, R. Franklin, D. Carver |
FCCM | 1 |
| 2002 | Multitasking Hardware on the SLAAC1-V Reconfigurable Computing System
Wesley J. Landaker, Michael J. Wirthlin, Brad L. Hutchings |
FPL | 3 |
| 2002 | Sea Cucumber: A Synthesizing Compiler for FPGAs
Justin L. Tripp, Preston A. Jackson, Brad L. Hutchings |
FPL | 3 |
| 2002 | Algorithms for Coloring Quadtrees
David Eppstein, Marshall W. Bern, Brad L. Hutchings |
Algorithmica | 3 |
| 2001 | Instrumenting Bitstreams for Debugging FPGA Circuits
Paul S. Graham, Brent E. Nelson, Brad L. Hutchings |
FCCM | 3 |
| 2001 | An Application-Specific Compiler for High-Speed Binary Image Morphology
Karl S. Hemmert, Brad L. Hutchings, Anshul Malvi |
FCCM | 2 |
| 2001 | Using Design-Level Scan to Improve FPGA Design Observability and Controllability for Functional Verification
Timothy Wheeler, Paul S. Graham, Brent E. Nelson, Brad L. Hutchings |
FPL | 4 |
| 2001 | Synthesizing RTL Hardware from Java Byte Codes
Michael J. Wirthlin, Brad L. Hutchings, Carl D. Worth |
FPL | 2 |
| 2001 | Gigaop DSP on FPGAabstractDSP algorithms such as sonar beamforming and automated target recognition, are a good match for FPGA technology due to their regular structure, available parallelism, pipeline-ability, and modest data word sizes. FPGA implementations of these applications outperformed their DSP and microprocessor counterparts by factors ranging from 10X on up with an equivalent sustained computational rate of more than 2 GOps/second per FPGA. This paper first describes each application and derives its computational requirements. The mapping process for each is then described followed by an analysis of the relative contributions to performance from pipelining, data parallelism, and memory usage. Brad L. Hutchings, Brent E. Nelson |
ICASSP | 1 |
| 2001 | Unifying simulation and execution in a design environment for FPGA systemsabstractField programmable gate array (FPGA)-based systems provide advantages over conventional hardware including: (1) availability of the hardware during design and debug; (2) programmability; and (3) visibility. These three advantages can greatly shorten the design and verification cycle. This paper discusses a design environment that exploits these three FPGA-specific advantages to create a unified simulation/execution debug environment implemented in the JHDL design system. The described system provides a hardware debugging environment with the functionality of a simulator but up to 10000/spl times/ faster. In addition, testbenches and other typical verification software used in simulators can be used to verify running hardware. Brad L. Hutchings, Brent E. Nelson |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2000 | Using general-purpose programming languages for FPGA designabstractGeneral-purpose programming languages (GPL) are effective vehicles for FPGA design because they are easy to use, extensible, widely available, and can be used to describe both the hardware and software aspects of a design. The strengths of the GPL approach to circuit design have been demonstrated by JHDL, a Java-based circuit design environment used to develop several large FPGA-based applications at several institutions. Major strengths of the JHDL environment include a common run-time for both simulation and hardware execution, and the overall extensibility of the parent Java environment. The common run-time environment means that all validation and support software (testbenches, application-specific interfaces, graphical user interfaces, etc.) can be used without modification with the built-in simulator or with the executing application as it runs in hardware. Extensibility also plays a big role because designers can easily add new capability to the environment by writing additional tools in the parent language, Java, using the wide variety of available libraries. This paper gives a brief introduction to JHDL syntax and demonstrates its features with an end-to-end application example. Brad L. Hutchings, Brent E. Nelson |
DAC | 1 |
| 2000 | Improving the FPGA Design Process through Determining and Applying Logical-to-Physical Design MappingsabstractWhile creating CCM-platform-independent, device-specific readback support for hardware debugging in the JHDL design environment, we have found that knowing how design elements from the user's logical design were mapped to their counterparts in the FPGA physical implementation can be very useful and important. With only a partial mapping from the logical to the physical, we would not be able to provide users of JHDL with a complete view of what their circuit is doing during hardware execution via FPGA readback mechanisms. As an example of how to determine logical-to-physical mappings of FPGA circuits, we outline the process of supporting readback for Xilinx XC4000 and Virtex designs under the JHDL environment. This same process should apply to other structural design methodologies and for other purposes. Synthesis methodologies require some additional steps to relate how the high-level HDL design mapped to the FPGA vendors' library elements. Paul S. Graham, Brad L. Hutchings, Brent E. Nelson |
FCCM | 2 |
| 1999 | A CAD Suite for High-Performance FPGA DesignabstractThis paper describes the current status of a suite of CAD tools designed specifically for use by designers who are developing high-performance configurable-computing applications. The basis of this tool suite is JHDL, a design tool originally conceived as a way to experiment with Run-Time Reconfigured (RTR) designs. However, what began as a limited experiment to model RTR designs with Java has evolved into a comprehensive suite of design tools and verification aids, with these tools being used successfully to implement high-performance applications in Automated Target Recognition (ATR), sonar beamforming, and general image processing on configurable-computing systems. Brad L. Hutchings, Peter Bellows, Joseph Hawkins, Karl S. Hemmert, Brent E. Nelson, Mike Rytting |
FCCM | 1 |
| 1999 | A Reconfigurable Arithmetic Array for Multimedia ApplicationabstractIn this paper we describe a recontigurable architecture optimised for media processing, and based on 4-bit ALUs and interconnect. Tony Stansfield, Igor Kostarnov, Jean Vuillemin, Brad L. Hutchings |
FPGA | 5 |
| 1998 | JHDL - An HDL for Reconfigurable SystemsabstractJHDL is a design tool for reconfigurable systems that allows designers to express circuit organizations that dynamically change over time in a natural way, using only standard programming abstractions found in object-oriented languages. JHDL manages FPGA resources in a manner that is similar to the way object-oriented languages manage memory: circuits are treated as distinct objects and a circuit is configured onto a configurable computing machine (CCM) by invoking its constructor effectively "constructing " an instance of the circuit onto the reconfigurable platform just as object instances are allocated in memory with conventional object-oriented languages. This approach of using object constructors/destructors to control the circuit lifetime on a CCM is a powerful technique that naturally leads to a dual simulation/execution environment where a designer can easily switch between either software simulation or hardware execution on a CCM with a single application description. Moreover JHDL supports dual hardware/software execution; parts of the application described using JHDL circuit constructs can be executed on the CCM while the remainder of the application the-GUI for example-can run on the CCM host. Based on an existing programming language (Java), JHDL requires no language extensions and can be used with any standard Java 1.1 distribution. Peter Bellows, Brad L. Hutchings |
FCCM | 2 |
| 1998 | Improving functional density using run-time circuit reconfiguration [FPGAs]abstractThe ability to provide flexibility and allow fine-grain circuit specialization make field programmable gate arrays (FPGA's) ideal candidates for computing elements within application-specific architectures. The benefits of gate-level specialization and reconfigurability can be extended by reconfiguring circuit resources at run-time. This technique, termed run-time reconfiguration (RTR), allows the exploitation of dynamic conditions or temporal locality within application-specific problems. For several applications, this technique has been shown to reduce the hardware resources required for computation. The use of this technique on conventional FPGA's, however, requires additional time for circuit reconfiguration. A functional density metric is introduced that balances the advantages of RTR against its associated reconfiguration costs. This metric is used to justify run-time reconfiguration against other more conventional approaches. Several run-time reconfigured applications are presented and analyzed using this approach. Michael J. Wirthlin, Brad L. Hutchings |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1997 | Automated target recognition on SPLASH 2abstractAutomated target recognition is an application area that requires special-purpose hardware to achieve reasonable performance. FPGA-based platforms can provide a high level of performance for ATR systems if the implementation can be adapted to the limited FPGA and routing resources of these architectures. The paper discusses a mapping experiment where a linear-systolic implementation of an ATR algorithm is mapped to the SPLASH 2 platform. Simple column oriented processors were used throughout the design to achieve high performance with limited nearest neighbor communication. The distributed SPLASH 2 memories are also exploited to achieve a high degree of parallelism. The resulting design is scalable and can be spread across multiple SPLASH 2 boards with a linear increase in performance. Michael Rencher, Brad L. Hutchings |
FCCM | 2 |
| 1997 | Improving Functional Density Through Run-Time Constant PropagationabstractCircuit specialization techniques such as constant propagation are commonly used to reduce both the hardware resources and cycle time of digital circuits. When reconfigurable FPGAs are used, these advantages can be extended by dynamically specializing circuits using run-time reconfiguration (RTR). For systems exploiting constant propagation, hardware resources can be reduced by folding constants within the circuit and dynamically changing the constants using circuit reconfiguration. To measure the benefits of circuit specialization, a functional density metric is presented. This metric allows the analysis of both static and run-time reconfigured circuits by including the cost of circuit reconfiguration. This metric will be used to justify runtime constant propagation as well as analyze the effects of reconfiguration time on run-time reconfigured systems. Michael J. Wirthlin, Brad L. Hutchings |
FPGA | 2 |
| 1996 | Mixing fixed and reconfigurable logic for array processingabstractThis paper describes the architecture of the MIX system that was designed to investigate the trade-off between the use of reconfigurable and fixed logic. The calculation of the dot-product of two vectors of 32 bit floating point numbers, that forms the basis of array processing in many engineering applications, is used as the basic algorithm for the investigation. The results indicate that fixed logic is more suited for floating point units and memories while reconfigurable logic is useful for implementing control logic providing significant flexibility. It was also found that the additional delay in reconfigurable logic can effectively overlap with the operating time of the fixed logic subsystems. The advantage of reconfigurability of the control is therefore combined with the high bandwidth properties of the fixed logic. Pieter J. Bakkes, Jan J. Du Plessis, Brad L. Hutchings |
FCCM | 3 |
| 1996 | Supporting FPGA microprocessors through retargetable software toolsabstractFPGA systems outperform many ASIC and supercomputer systems through effective use of the reconfigurable resource. Reusing design effort across different applications requires a standard, flexible software environment. Driving FPGA systems from ANSI C is possible using 1 cc (an ANSI C compiler) targeted at an FPGA system and dasm (a retargetable, flexible assembler). The compiler supports custom hardware capabilities of FPGA systems, as well as all constructs of C. The assembler reads instruction definitions at assemble time, allowing the user to add new custom hardware functions which dasm can assemble correctly to an instruction stream the hardware executes. A source code debugger has been implemented for this system. David A. Clark, Brad L. Hutchings |
FCCM | 2 |
| 1996 | Sequencing Run-Time Reconfigured Hardware with SoftwareabstractNo abstract available. Michael J. Wirthlin, Brad L. Hutchings |
FPGA | 2 |
| 1995 | Design methodologies for partially reconfigured systemsabstractRun time reconfiguration (RTR) as an implementation approach that divides an application into a series of sequentially executed stages with each stage implemented as a separate circuit module. Partial RTR extends this approach by partitioning these stages and designing their circuit modules such that they exhibit a high degree of functional and physical commonality. Transitioning between configurations can then be accomplished by updating only the differences between configurations. This reduces the amount of time that an RTR application spends configuring and significantly enhances overall performance. The paper presents the design methodology for partial RTR in the context of RRANN2, a partial RTR artificial neural network. James D. Hadley, Brad L. Hutchings |
FCCM | 2 |
| 1995 | A dynamic instruction set computerabstractA dynamic instruction set computer (DISC) has been developed that supports demand-driven modification of its instruction set. Implemented with partially reconfigurable FPGAs, DISC treats instructions as removable modules paged in and out through partial reconfiguration as demanded by the executing program. Instructions occupy FPGA resources only when needed and FPGA resources can be reused to implement an arbitrary number of performance-enhancing application-specific instructions. DISC further enhances the functional density of FPGAs by physically relocating instruction modules to available FPGA space. Michael J. Wirthlin, Brad L. Hutchings |
FCCM | 2 |
| 1994 | Multiple-Layer Cross-Field Ultrasonic Tactile SensorabstractThis paper presents an ultrasonic "cross-field" tactile sensor that is robust enough for use in industrial robotic applications. This sensor uses ultrasound to measure the thickness of an elastomer pad and a "cross-field" arrangement to eliminate the spurious coupling commonly found in systems that employ cross-point switching. A tactile sensor with 256 elements (arranged as a 16/spl times/16 array of 1/16"-square elements) was shown to linearly measure thickness of the elastomer pad across a range of 900 microns with a resolution of 6.8 microns. In addition, the sensor is shown to accurately measure surface normal information.> Brad L. Hutchings, A. R. Grahn, Russell J. Petersen |
ICRA | 1 |
| 1989 | VLSI design and implementation of a real-time image segmentation processor
Bir Bhanu, Brad L. Hutchings, Kent F. Smith |
Mach. Vis. Appl. | 2 |