Brent E. Nelson

dblp:90/2444 · DBLP profile ↗
← Back
52ranked-venue papers
2as first author
0since 2021 · last 2019
0000-0002-7523-3269ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 42 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-authorArtificial intelligence and machine learning · 3Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Electronic design automation · 47% Reconfigurable computing and FPGAs · 38% Emerging computing paradigms · 6%
Software engineering, system software, and programming languages
1 paper
Programming languages and type systems · 100%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
FPGA design flow
0.622019
Maverick: A Stand-alone CAD Flow for Xilinx 7-Series FPGAs · FPGA 2019
RapidSmith 2: A Framework for BEL-level CAD Exploration on Xilinx FPGAs · FPGA 2015
Electronic design automation
physical design
0.412019
Maverick: A Stand-alone CAD Flow for Xilinx 7-Series FPGAs · FPGA 2019
Electronic design automation › physical design
placement and routing
0.412019
Maverick: A Stand-alone CAD Flow for Xilinx 7-Series FPGAs · FPGA 2019
Electronic design automation › design automation tools › FPGA CAD
FPGA design tools
0.212015
RapidSmith 2: A Framework for BEL-level CAD Exploration on Xilinx FPGAs · FPGA 2015
Emerging computing paradigms › approximate and stochastic computing › stochastic computing
bit-stream generation
0.112019
Maverick: A Stand-alone CAD Flow for Xilinx 7-Series FPGAs · FPGA 2019
Reconfigurable computing and FPGAs › dynamic reconfiguration
partial reconfiguration
0.112019
Maverick: A Stand-alone CAD Flow for Xilinx 7-Series FPGAs · FPGA 2019
Integrated circuit design
digital circuit design
0.112005
Choice of base revisited: higher radices for FPGA-based floating-point computation (abstract only) · FPGA 2005
Processor architecture and microarchitecture › computer arithmetic
floating-point representation
0.112005
Choice of base revisited: higher radices for FPGA-based floating-point computation (abstract only) · FPGA 2005
Reconfigurable computing and FPGAs
FPGA arithmetic
0.112005
Choice of base revisited: higher radices for FPGA-based floating-point computation (abstract only) · FPGA 2005
Reconfigurable computing and FPGAs
FPGA-based signal processing
0.011998
FPGA-Based Sonar Processing · FPGA 1998
Performance modeling and evaluation › benchmarking
benchmark evaluation
0.011996
Locality as a Visualization Tool · IEEE Trans. Computers 1996
Performance modeling and evaluation
benchmarking
0.011996
Locality as a Visualization Tool · IEEE Trans. Computers 1996
Performance modeling and evaluation
workload characterization
0.011996
Locality as a Visualization Tool · IEEE Trans. Computers 1996
Storage systems
adaptive prefetching
0.011993
Multiple Prefetch Adaptive Disk Caching · IEEE Trans. Knowl. Data Eng. 1993
Memory systems › cache management › storage caching
disk cache
0.011993
Multiple Prefetch Adaptive Disk Caching · IEEE Trans. Knowl. Data Eng. 1993
Performance modeling and evaluation › simulation › discrete-event simulation
trace-driven simulation
0.011993
Multiple Prefetch Adaptive Disk Caching · IEEE Trans. Knowl. Data Eng. 1993
Embedded and real-time systems
real-time signal processing
0.011998
FPGA-Based Sonar Processing · FPGA 1998
Electronic design automation › CAD framework
integrated circuit design system
0.011984
The structure and operation of a relational database system in a cell-oriented integrated circuit design system · DAC 1984
Database system architecture and tuning
relational database system
0.011984
The structure and operation of a relational database system in a cell-oriented integrated circuit design system · DAC 1984

Methods — techniques the papers use, named apart from their topics

packing · 0.6synthesis · 0.4routing · 0.4placement · 0.4clustering · 0.2denormalized number support · 0.1reference sequence analysis · 0.0time-delay beamforming · 0.0trace-driven simulation · 0.0
YearPublicationVenuePosition
2019 Maverick: A Stand-Alone CAD Flow for Partially Reconfigurable FPGA Modules
abstract
This paper presents Maverick, a proof-of-concept computer-aided design (CAD) flow for generating reconfigurable modules (RMs) which target partial reconfiguration (PR) regions in field-programmable gate array (FPGA) designs. After an initial static design and PR region are created with Xilinx's Vivado PR flow, the Maverick flow can then compile and configure RMs onto that PR region—without the use of vendor tools. Maverick builds upon existing open source tools (Yosys, RapidSmith2, and Project X-Ray) to form an end-to-end compilation flow. This paper describes the Maverick flow and shows the results of it running on a PYNQ-Z1's ARM processor to compile a set of HDL designs to partial bitstreams. The resulting bitstreams were configured onto the PYNQ-Z1's FPGA fabric, demonstrating the feasibility of a single-chip embedded system which can both compile HDL designs to bitstreams and then configure them onto its own programmable fabric.
Dallon Glick, Jesse Grigg, Brent E. Nelson, Michael J. Wirthlin
FCCM3
2019 Maverick: A Stand-alone CAD Flow for Xilinx 7-Series FPGAs
abstract
Traditionally, an FPGA vendor's own set of computer-aided design (CAD) tools are used to generate circuits for a given vendor's FPGAs. However, numerous non-vendor CAD tools have been introduced to supplement the vendor-provided tools, allowing novel ideas to be explored and a variety of technical challenges to be addressed. This poster presents Maverick, a stand-alone CAD flow for compiling Verilog to bitstreams for Xilinx 7-Series devices. After an initial configuration design is created with Xilinx's Vivado partial reconfiguration (PR) flow to define a static design and a PR region, the Maverick flow can then compile and map Verilog design into that PR region - without the use of vendor tools. The Maverick flow combines two existing open source projects (Yosys and Project X-Ray) with our own RapidSmith2 tools to form an end-to-end compilation flow. It uses Yosys (synthesis), RapidSmith2 (pack, place, route), and the Project X-Ray tools (bitstream generation), taking Verilog designs as input and generating partial bitstreams as output. Several modifications were made to these existing tools and completely new tools were created, including a new RapidSmith2-based router, as a part of this work. This poster details these CAD steps and shows the results of the CAD flow running on a PYNQ-Z1 SoC's ARM processor to compile a set of HDL designs to partial bitstreams. The resulting bitstreams were configured onto the PYNQ-Z1's FPGA fabric, demonstrating the feasibility of a single-chip system which can both compile HDL designs to bitstreams and then configure them onto its own fabric.
Dallon Glick, Jesse Grigg, Brent E. Nelson, Michael J. Wirthlin
FPGA3
2017 Vivado design interface: An export/import capability for Vivado FPGA designs
abstract
Research tools targeting commercial FPGAs have most commonly been based on the Xilinx Design Language (XDL). Vivado, however, does not support XDL, preventing similar tools from being created for next-generation devices. Instead, Vivado includes a Tcl interface that exposes Xilinx's internal design and device data structures. Considerable challenges still remain to users attempting to leverage this Tcl interface to develop external CAD tools. This paper presents the Vivado Design Interface (VDI), a set of file formats and Tcl functions that address the challenges of exporting and importing designs to and from Vivado. To demonstrate its use, VDI has been integrated with RapidSmith2, an external FPGA CAD framework. To our knowledge this work is the first successful attempt to provide an open-source tool-flow that can export designs from Vivado, manipulate them with external CAD tools, and re-import an equivalent representation back into Vivado.
Thomas Townsend, Brent E. Nelson
FPL2
2016 Efficient processing of phased array radar in sense and avoid application using heterogeneous computing
abstract
This paper describes a tightly integrated phased-array radar platform for a UAV-based sense and avoid application. The system leverages heterogeneous computing on a Zynq 7000 SoC processor to efficiently process the large streams of radar data required to guide a small Unmanned Air Vehicle (UAV) away from obstacles and intruding air vehicles. The system is mounted on a small UAV which requires stringent size, weight, and power constraints, and thus data processing must be done efficiently. The system is implemented on the commercial MicroZed development board. Low-level radar data-stream processing is performed in the FPGA fabric on the Zynq, which accelerates processing and improves energy efficiency. The Zynq ARM CPUs are used to perform higher-level radar processing and avoidance algorithms. The final proof of concept system is 2.25 × 4 × 1.5 in3weighing only 120 g (0.26 lbs) and consumes 8 watts of power. As an example, the improvements demonstrated by executing the FFT algorithm in hardware will be highlighted where hardware FFT processing provides a 17× improvement in processing time and a 49× improvement in energy efficiency over a CPU only implementation on the same SoC.
Luke Newmeyer, Doran Wilde, Brent E. Nelson, Michael J. Wirthlin
FPL3
2016 An XDL alternative for interfacing RapidSmith and Vivado
abstract
In recent years, the RapidSmith CAD tool [1] has been used with ISE to create custom CAD tools targeting Xilinx FPGAs. This tool flow was based on the Xilinx Design Language (XDL), a human-readable representation of a netlist that contains placement and routing information. The XDL interface also provided device representation files (XDLRC files), detailing the available resources of a given FPGA part. Using RapidSmith, a Xilinx design could be exported out of ISE at any stage of the design flow, manipulated in RapidSmith (logic modification, place, or route), and imported back into ISE to complete the remainder of implementation.
Thomas Townsend, Brent E. Nelson, Michael J. Wirthlin
FPL2
2015 RapidSmith 2: A Framework for BEL-level CAD Exploration on Xilinx FPGAs
abstract
RapidSmith is an open-source framework that allows for the exploration of novel approaches to the FPGA CAD flow for Xilinx devices. However, RapidSmith has poor support for manipulating designs below the slice level. In this paper, we highlight many of the projects RapidSmith enables and present extensions incorporated into "RapidSmith 2" that expose LUTs and flip-flops for direct manipulation in custom-built CAD tools. To demonstrate the utility of RapidSmith 2 we present the results of work to identify BELs in a design which must be clustered together and a tool that does pre-packing clustering accordingly.
Travis Haroldsen, Brent E. Nelson, Brad L. Hutchings
FPGA2
2013 Rapid FPGA design prototyping through preservation of system logic: A case study
abstract
FPGA designs often contain significant amounts of logic such as a board support package that remains unaltered throughout the design process. However, during normal operation, standard FPGA implementation tools re-implement the entire system, including the unchanged logic, adding to the turn around time of design iterations. Recently, FPGA implementation flows have appeared that allow preserving parts of a previously implemented design. In this study, we evaluate the potential speedups in implementation time achievable through preserving the unchanging portion of a design's implementation. We perform these evaluations using Xilinx Partitions, Xilinx SmartGuide, and the HMFlow rapid implementation tool.
Travis Haroldsen, Brent E. Nelson, Brad White
FPL2
2013 Impact of hard macro size on FPGA clock rate and place/route time
abstract
Hard macros are completely placed/routed elements that are treated as primitives and that are relatively placed as a single element. A system composed of such macros consists of many fewer effective primitives and nets and as such can be placed and routed much more quickly. Prior work in this research area dealt with small, general-purpose macros such as 16-bit registers, adders, etc., and demonstrated that place/route time could be reduced by an order of magnitude with a corresponding 3-4X reduction in clock rate. In this work, much larger hard macros are developed such as mixers, softcore processors, FFTs, etc., and the use of these larger macros is shown to further reduce place/route time by an additional 2.5-4X, for a total of a 30-40X reduction in compile time. Clock rate is also improved, relative to earlier work, by an additional 60-70%.
Chris Lavin, Brent E. Nelson, Brad L. Hutchings
FPL2
2013 Improving clock-rate of hard-macro designs
abstract
HMFlow reuses precompiled circuit modules (hard macros) and other techniques to rapidly compile large designs in a few seconds - many times faster than standard Xilinx flows. However, the clock rates of designs rapidly compiled by HMFlow are often significantly lower than those compiled by the Xilinx flow. To improve clock rates, HMFlow algorithms were modified as follows: (1) the router was modified to take advantage of longer routing wires in the FPGA devices, (2) the original greedy placer was replaced with an annealing-based placer, and (3) certain registers were removed from the hard-macro and moved into the fabric to reduce critical-path delays. Benchmark circuits compiled with these modifications can achieve clock rates that are about 75% as fast as those achieved by Xilinx, on average. Fast run-times are also preserved; the improved algorithms only increase HMFlow run-times by about 50% across the benchmark suite so that HMFlow remains more than 30× faster than the standard Xilinx flow for the benchmarks tested in this paper.
Chris Lavin, Brent E. Nelson, Brad L. Hutchings
FPT2
2011 HMFlow: Accelerating FPGA Compilation with Hard Macros for Rapid Prototyping
abstract
The FPGA compilation process (synthesis, map, place, and route) is a time consuming task that severely limits designer productivity. Compilation time can be reduced by saving implementation data in the form of hard macros. Hard macros consist of previously synthesized, placed and routed circuits that enable rapid design assembly because of the native FPGA circuitry (primitives and nets)which they encapsulate. This work presents results from creating a new FPGA design flow based on hard macros called HMF low. HMF low has shown speedups of 10-50X over the fastest configuration of the Xilinx tools. Designed for rapid prototyping, HMF low achieves these speedups by only utilizing up to 50 percent of the resources on an FPGA and produces implementations that run 2-4X slower than those produced by Xilinx. These speedups are obtained on a wide range of benchmark designs with some exceeding 18,000 slices on a Virtex 4 LX200.
Chris Lavin, Marc Padilla, Jaren Lamprecht, Philip Lundrigan, Brent E. Nelson, Brad L. Hutchings
FCCM5
2011 XDL-Based Module Generators for Rapid FPGA Design Implementation
abstract
XDLCoreGen is described, a module generator framework which directly generates placed and routed hard macros in XDL. XDLCoreGen is intended to be used in a rapid prototyping flow such as HM Flow, which achieves short FPGA implementation times by bypassing the conventional Xilinx tool flow and directly assembling designs from pre-built hard macros. The structure of XDLCoreGen is described and its unique cache-based router is highlighted as a key component to achieving extremely fast module generation times. Testing results are provided to demonstrate XDLCoreGen's ability to generate fully placed and routed hard macros in milliseconds.
Subhrashankha Ghosh, Brent E. Nelson
FPL2
2011 RapidSmith: Do-It-Yourself CAD Tools for Xilinx FPGAs
abstract
Creating CAD tools for commercial FPGAs is a difficult task. Closed proprietary device databases and unsupported interfaces are largely to blame for the lack of CAD research found on commercial architectures versus hypothetical architectures. This paper formally introduces RapidSmith, a new set of tools and APIs that enable CAD tool creation for Xilinx FPGAs. Based on the Xilinx Design Language (XDL), RapidSmith provides a compact, yet, fast device database with hundreds of APIs that enable the creation of placers, routers and several other tools for Xilinx devices. RapidSmith alleviates several of the difficulties of using XDL and this work demonstrates the kinds of research facilitated by removing such challenges.
Chris Lavin, Marc Padilla, Jaren Lamprecht, Philip Lundrigan, Brent E. Nelson, Brad L. Hutchings
FPL5
2010 Increasing Design Productivity through Core Reuse, Meta-data Encapsulation, and Synthesis
abstract
This paper presents a novel IP core reuse strategy which reduces design time from days to hours for communication circuits such as digital radio receivers. This design productivity is obtained by leveraging a highly parameterized library of communication specific cores. These cores are described in IP-XACT XML with vendor extensions describing the timing behavior of their communication interfaces. A synthesis tool, called Ogre, was created that generates the communication interfaces between cores described in IP-XACT and synthesizes full designs from structural synchronous dataflow specifications. Design productivity improvements are demonstrated with several radio receiver designs.
Adam Arnesen, Kevin Ellsworth, Derrick Gibelyou, Travis Haroldsen, Jared Havican, Marc Padilla, Brent E. Nelson, Michael Rice, Michael J. Wirthlin
FPL7
2010 Using Hard Macros to Reduce FPGA Compilation Time
abstract
The FPGA compilation process (synthesis, map, placement, routing) is a time-consuming process that limits designer productivity. Compilation time can be reduced by using pre-compiled circuit blocks (hard macros). Hard macros consist of previously synthesized, mapped, placed and routed circuitry that can be relatively placed with short tool runtimes and that make it possible to reuse previous computational effort. Two experiments were performed to demonstrate feasibility that hard macros can reduce compilation time. These experiments demonstrated that an augmented Xilinx flow designed specifically to support hard macros can reduce overall compilation time by 3x. Though the process of incorporating hard macros in designs is currently manual and error-prone, it can be automated to create compilation flows with much lower compilation time.
Chris Lavin, Marc Padilla, Subhrashankha Ghosh, Brent E. Nelson, Brad L. Hutchings, Michael J. Wirthlin
FPL4
2010 Rapid prototyping tools for FPGA designs: RapidSmith
abstract
Designer productivity for FPGA design is significantly limited by the time-consuming nature of the FPGA compilation process (synthesis, map, placement, and routing). However, experimentation on alternative CAD tools for this purpose for Xilinx devices has been somewhat limited. This paper describes the development and distribution of RapidSmith, a software library to facilitate the manipulation of XDL designs and upon which a complete CAD system can be based. The demonstration portion of this paper will show prototypes of representative CAD tools which can be easily built on top of the RapidSmith system.
Chris Lavin, Marc Padilla, Philip Lundrigan, Brent E. Nelson, Brad L. Hutchings
FPT4
2010 Two-frame structure from motion using optical flow probability distributions for unmanned air vehicle obstacle avoidance
Dah-Jye Lee, Paul C. Merrell, Zhaoyi Wei, Brent E. Nelson
Mach. Vis. Appl.4
2010 Hardware-Friendly Vision Algorithms for Embedded Obstacle Detection Applications
abstract
Accurate optical flow estimation is a crucial task for many computer vision applications. However, because of its computational power and processing speed requirements, it is rarely used for real-time obstacle detection, especially for small unmanned vehicle and embedded applications. Two hardware-friendly vision algorithms are proposed in this paper to address this challenge. A ridge regression-based optical flow algorithm is developed to cope with the existing collinear problem in traditional least-squares approaches for calculating optical flow. Additionally, taking advantage of hardware parallelism, spatial and temporal smoothing operations are applied to image sequence derivatives to improve accuracy. An efficient motion field analysis algorithm using the optical flow values and based on a simplified motion model is also developed and implemented in hardware. The resulting obstacle detection algorithm is specifically designed for ground vehicles moving on planar surfaces. Results from the software simulations and hardware execution of the two proposed algorithms prove that with adequate hardware, a low power, compact obstacle detection sensor can be realized for small unmanned vehicles and embedded applications.
Zhaoyi Wei, Dah-Jye Lee, Brent E. Nelson, James K. Archibald
IEEE Trans. Circuits Syst. Video Technol.3
2010 A Comparison Study on Implementing Optical Flow and Digital Communications on FPGAs and GPUs
abstract
FPGA devices have often found use as higher-performance alternatives to programmable processors for implementing computations. Applications successfully implemented on FPGAs typically contain high levels of parallelism and often use simple statically scheduled control and modest arithmetic. Recently introduced computing devices such as coarse-grain reconfigurable arrays, multi-core processors, and graphical processing units promise to significantly change the computational landscape and take advantage of many of the same application characteristics that fit well on FPGAs. One real-time computing task, optical flow, is difficult to apply in robotic vision applications because of its high computational and data rate requirements, and so is a good candidate for implementation on FPGAs and other custom computing architectures. This article reports on a series of experiments mapping a collection of different algorithms onto both an FPGA and a GPU. For two different optical flow algorithms the GPU had better performance, while for a set of digital comm MIMO computations, they had similar performance. In all cases the FPGA implementations required 10x the development time. Finally, a discussion of the two technology’s characteristics is given to show they achieve high performance in different ways.
John Bodily, Brent E. Nelson, Zhaoyi Wei, Dah-Jye Lee, Jeff Chase
ACM Trans. Reconfigurable Technol. Syst.2
2009 Optical Flow on the Ambric Massively Parallel Processor Array (MPPA)
abstract
The Ambric Massively Parallel Processor Array (MPPA) is a device that contains 336 32-bit RISC processors and is appropriate for embedded systems due to its relatively small physical and power footprint. Optical flow is a computationally-demanding and highly parallelizeable image-processing algorithm with applications in embedded systems such as robotics and autonomous vehicles. An optical flow algorithm is implemented on the Ambric device and is shown to achieve near FPGA performance at similar levels of power consumption while requiring many fewer lines of code (Java) than its FPGA counterpart (VHDL).
Brad L. Hutchings, Brent E. Nelson, Stephen West, Reed Curtis
FCCM2
2009 Comparing fine-grained performance on the Ambric MPPA against an FPGA
abstract
A simple image-processing application is implemented on the Ambric MPPA and an FPGA, using a similar implementation for both devices. FPGAs perform extremely well on this kind of application and provide a good benchmark for comparison. The Ambric implementation starts out with a naive implementation and proceeds through several design optimizations until it reaches a maximum frame rate of 164 FPS (512 times 512 images) which turns out to be approximately 7times slower than the FPGA. The final Ambric implementation uses only 18 of 336 available processors, achieves more than sufficient performance for realtime embedded applications, and has excess processors to use for implementing additional algorithms. After introducing the image processing application and its implementation on both devices, the paper compares and contrasts the intrinsic, general characteristics of Ambric MPPA and FPGA devices.
Brad L. Hutchings, Brent E. Nelson, Stephen West, Reed Curtis
FPL2
2009 Multi-frame structure from motion using optical flow probability distributions
Dah-Jye Lee, Paul C. Merrell, Brent E. Nelson, Zhaoyi Wei
Neurocomputing3
2008 Real-Time Optical Flow Calculations on FPGA and GPU Architectures: A Comparison Study
abstract
FPGA devices have often found use as higher-performance alternatives to programmable processors for implementing a variety of computations. Applications successfully implemented on FPGAs have typically contained high levels of parallelism and have often used simple statically-scheduled control and modest arithmetic. Recently introduced computing devices such as coarse grain reconfigurable arrays, multi-core processors, and graphical processing units (GPUs) promise to significantly change the computational landscape for the implementation of high-speed real-time computing tasks. One reason for this is that these architectures take advantage of many of the same application characteristics that fit well on FPGAs. One real-time computing task, optical flow, is difficult to apply in robotic vision applications in practice because of its high computational and data rate requirements, and so is a good candidate for implementation on FPGAs and other custom computing architectures. In this paper, a tensor-based optical flow algorithm is implemented on both an FPGA and a GPU and the two implementations discussed. The two implementations had similar performance, but with the FPGA implementation requiring 12x more development time. Other comparison data for these two technologies is then given for three additional applications taken from a MIMO digital communication system design, providing additional examples of the relative capabilities of these two technologies.
Jeff Chase, Brent E. Nelson, John Bodily, Zhaoyi Wei, Dah-Jye Lee
FCCM2
2008 Real-time accurate optical flow-based motion sensor
abstract
An accurate real-time motion sensor implemented in an FPGA is introduced in this paper. This sensor applies an optical flow algorithm based on ridge regression to solve the collinear problem existing in traditional least squares methods. It additionally applies extensive temporal smoothing of the image sequence derivatives to improve the accuracy of its optical flow estimates. Implemented on a customized embedded FPGA platform, it is capable of processing 60 320 × 240 images or 15 640 × 480 images per second. By evaluating its accuracy on synthetic sequences, it is shown here that the proposed design achieves very high accuracy compared to other known hardware-based designs.
Zhaoyi Wei, Dah-Jye Lee, Brent E. Nelson, James K. Archibald
ICPR3
2007 A Fast and Accurate Tensor-based Optical Flow Algorithm Implemented in FPGA
abstract
Many computer vision applications require real-time processing of image data. This requirement is especially critical for autonomous vehicles performing obstacle avoidance, path planning, and target tracking tasks. A quickly calculated and relatively rough motion estimate is more useful for autonomous navigation than a more accurate, but slowly calculated estimate. Recent technology advancements in small unmanned air and ground vehicles make many low-cost surveillance and military applications possible. Most of these applications demand a low power, compact, light weight, and high speed computation platform for processing image data in real time. In most cases, the traditional general purpose processor and sequentially executed software approach does not meet these requirements. In this paper, a tensor-based optical flow algorithm is modified and implemented using field programmable gate array (FPGA) for small unmanned vehicle obstacle avoidance and navigation
Zhaoyi Wei, Dah-Jye Lee, Brent E. Nelson, Michael Martineau
WACV3
2006 The Mythical CCM: In Search of Usable (and Resuable) FPGA-Based General Computing Machines
abstract
Early FPGA researchers understood that FPGAs made possible the creation of a new, flexible, and powerful class of machine - the configurable computing machine (CCM). The earliest CCMs featured rudimentary but significant integrated design, debug, and runtime environments. This paper reviews those environments as well as more recent work using JHDL, designed to investigate how a symbolic hardware debugging environment for CCMs can be created with many of the features normally associated with software debug systems. The paper reviews lessons learned from that work and concludes by discussing the role integrated design, debug, and runtime environments may play in future CCM-based systems.
Brent E. Nelson
ASAP1
2005 Higher Radix Floating-Point Representations for FPGA-Based Arithmetic
abstract
FPGA implementations of floating-point operators have historically been designed to use binary floating-point representations. The general computing world settled on binary floating-point representations over three decades ago, and more recently, the FPGA community followed their example. Binary representations were chosen to maximize numerical accuracy per bit of data, however, the unique nature of FPGA-based computation makes numerical accuracy per unit of FPGA resources a more important measure of the usefulness of a given floating-point representation. In this paper, we show that higher radix floating-point representations are well suited to FPGA-based computations, especially high precision calculations which require the support of denormalized numbers. Higher radix representations use FPGA resources more efficiently. For example, a hexadecimal floating-point adder has a 30% smaller area-time product than its binary counterpart, while still delivering equal worst-case and better average-case numerical accuracy. Contrary to established belief, higher radix representations are useful for FPGA applications requiring IEEE 754 compliance, since they can deliver superior numerical performance while still using less FPGA resources.
Bryan Catanzaro, Brent E. Nelson
FCCM2
2005 Choice of base revisited: higher radices for FPGA-based floating-point computation (abstract only)
abstract
The general computing world settled on radix 2 floating point representations over three decades ago. The analyses which led to this choice were all based on the underlying premise that the goal of a floating-point representation is to maximize numerical accuracy per bit of data. However, the unique nature of FPGA-based computations makes numerical accuracy per unit of FPGA resources a more important measure by which to judge the usefulness of a given floating point representation. Due to the high cost of shifters as implemented on FPGAs, higher radix floating-point representations are uniquely suited to FPGA-based computations, especially high precision calculations which require the support of denormalized numbers. Higher radix representations use FPGA resources more efficiently. For example, a radix 16 adder requires 20% less LUTs than its radix 2 counterpart, while delivering equal worst-case and better average case numerical accuracy.
Bryan Catanzaro, Brent E. Nelson
FPGA2
2005 A Flexible Circuit-Switched NOC for FPGA-Based Systems
abstract
Increases in chip density due to Moore's Law allow for the implementation of ever larger and more complex systems on a single chip. The communication mechanisms employed in such SOC's is an important contribution to their overall performance. Networks on chip promise to overcome the scalability problems found in bus-based interconnect. Most work to this point has focused on packet-switched NOC's. Circuit-switched networks are a lightweight alternative that promise high communication rates and predictable communication latencies. This paper describes PNoC, a very flexible circuit-switched NOC suitable for use in FPGA-based systems. Implementation results on a Virtex-il Pro device are given using an image binarization demonstration which resulted in a 2 - 23x speedup with a 29% area overhead compared to a shared bus implementation.
Clint Hilton, Brent E. Nelson
FPL2
2004 A Parallel FFT Architecture for FPGAs
Joseph M. Palmer, Brent E. Nelson
FPL2
2004 JHDLBits: The Merging of Two Worlds
Alexandra Poetter, Jesse Hunter, Cameron D. Patterson, Peter M. Athanas, Brent E. Nelson, Neil Steiner
FPL5
2003 Reconfigurable Computing Application Frameworks
abstract
FPGA-based (field programmable gate array) configurable computing machines (CCMs) offer powerful and flexible general-purpose computing platforms. However, development for FPGA-based designs using modern CAD (computer aided design) tools is geared mainly toward an ASIC-like process. This is inadequate for the needs of CCM application development. This paper discusses an application framework for developing CCM-based applications beyond just the hardware configuration. This framework leverages the advantages of CCMs (availability, programmability, visibility, and controllability) to help create CCM-based applications throughout the entire development process (i.e. design, debug, and deploy). The framework itself is deployed with the final application, thus permitting dynamic circuit configurations that include data folding optimizations based on user input. The resulting system aids in creating applications that are potentially more intuitive, easier to develop, and better performing. An example application demonstrates the use of the application framework and the potential benefits.
Anthony L. Slade, Brent E. Nelson, Brad L. Hutchings
FCCM2
2003 Tradeoffs of Designing Floating-Point Division and Square Root on Virtex FPGAs
abstract
Low latency, high throughput and small area are three major design considerations of an FPGA (field programmable gate array) design. In this paper, we present a high radix SRT division algorithm and a binary restoring square root algorithm. We describe three implementations of floating-point division operations with a variable width and precision based on Virtex-2 FPGAs. One is a low cost iterative implementation; another is a low latency array implementation; and the third is a high throughput pipelined implementation. The implementations of floating-point square root operations are presented as well. In addition to the design of modules, we also analyze the tradeoffs among the cost, latency and throughput with strategies on how to reduce the cost or improve the performance.
Brent E. Nelson
FCCM2
2002 Novel Optimizations for Hardware Floating-Point Units in a Modern FPGA Architecture
Eric Roesler, Brent E. Nelson
FPL2
2002 Debug methods for hybrid CPU/FPGA systems
abstract
The combining of one or more CPU's and an FPGA fabric on the same die is growing in popularity. Such programmable system-on-chip (PSOC) systems promise performance and development time advantages over conventional technology. In a PSOC design, design errors may occur in many different places - the design of the CPU, the embedded software, the FPGA-based parts of the design, or the interfaces between these various parts. This paper presents a flexible tool that allows the user to dynamically adapt the CAD tool's behavior to the level of detail needed to track down design errors in PSOC designs. It provides for the creation of and coexistence of software source debuggers and gate level debug tools, all incorporated into the same debugging environment. A prototype PSOC debugging system based on a derivative of the JHDL CAD tool is presented which illustrates the range of debug support in both simulation and hardware execution modes that can be provided for PSOC debug. Extensions to the tool to support a wider range of embedded processors and GNU compiler tools are discussed along with conclusions and future work.
Eric Roesler, Brent E. Nelson
FPT2
2001 Instrumenting Bitstreams for Debugging FPGA Circuits
Paul S. Graham, Brent E. Nelson, Brad L. Hutchings
FCCM2
2001 Using Design-Level Scan to Improve FPGA Design Observability and Controllability for Functional Verification
Timothy Wheeler, Paul S. Graham, Brent E. Nelson, Brad L. Hutchings
FPL3
2001 Gigaop DSP on FPGA
abstract
DSP algorithms such as sonar beamforming and automated target recognition, are a good match for FPGA technology due to their regular structure, available parallelism, pipeline-ability, and modest data word sizes. FPGA implementations of these applications outperformed their DSP and microprocessor counterparts by factors ranging from 10X on up with an equivalent sustained computational rate of more than 2 GOps/second per FPGA. This paper first describes each application and derives its computational requirements. The mapping process for each is then described followed by an analysis of the relative contributions to performance from pipelining, data parallelism, and memory usage.
Brad L. Hutchings, Brent E. Nelson
ICASSP2
2001 Unifying simulation and execution in a design environment for FPGA systems
abstract
Field programmable gate array (FPGA)-based systems provide advantages over conventional hardware including: (1) availability of the hardware during design and debug; (2) programmability; and (3) visibility. These three advantages can greatly shorten the design and verification cycle. This paper discusses a design environment that exploits these three FPGA-specific advantages to create a unified simulation/execution debug environment implemented in the JHDL design system. The described system provides a hardware debugging environment with the functionality of a simulator but up to 10000/spl times/ faster. In addition, testbenches and other typical verification software used in simulators can be used to verify running hardware.
Brad L. Hutchings, Brent E. Nelson
IEEE Trans. Very Large Scale Integr. Syst.2
2000 Using general-purpose programming languages for FPGA design
abstract
General-purpose programming languages (GPL) are effective vehicles for FPGA design because they are easy to use, extensible, widely available, and can be used to describe both the hardware and software aspects of a design. The strengths of the GPL approach to circuit design have been demonstrated by JHDL, a Java-based circuit design environment used to develop several large FPGA-based applications at several institutions. Major strengths of the JHDL environment include a common run-time for both simulation and hardware execution, and the overall extensibility of the parent Java environment. The common run-time environment means that all validation and support software (testbenches, application-specific interfaces, graphical user interfaces, etc.) can be used without modification with the built-in simulator or with the executing application as it runs in hardware. Extensibility also plays a big role because designers can easily add new capability to the environment by writing additional tools in the parent language, Java, using the wide variety of available libraries. This paper gives a brief introduction to JHDL syntax and demonstrates its features with an end-to-end application example.
Brad L. Hutchings, Brent E. Nelson
DAC2
2000 Improving the FPGA Design Process through Determining and Applying Logical-to-Physical Design Mappings
abstract
While creating CCM-platform-independent, device-specific readback support for hardware debugging in the JHDL design environment, we have found that knowing how design elements from the user's logical design were mapped to their counterparts in the FPGA physical implementation can be very useful and important. With only a partial mapping from the logical to the physical, we would not be able to provide users of JHDL with a complete view of what their circuit is doing during hardware execution via FPGA readback mechanisms. As an example of how to determine logical-to-physical mappings of FPGA circuits, we outline the process of supporting readback for Xilinx XC4000 and Virtex designs under the JHDL environment. This same process should apply to other structural design methodologies and for other purposes. Synthesis methodologies require some additional steps to relate how the high-level HDL design mapped to the FPGA vendors' library elements.
Paul S. Graham, Brad L. Hutchings, Brent E. Nelson
FCCM3
2000 A Reconfigurable Computing Architecture for Microsensors
abstract
Microsensor systems are described that support reconnaissance, surveillance and target acquisition (RSTA) operations. Since communication bandwidth on a microsensor is limited by the power constraints imposed by desired sensor lifespan, the amount of data that can be transmitted is minimal. Therefore, much of the signal processing needed to implement the desired functionality must be performed within the aggressive size, power and weight constraints of the microsensor itself. Furthermore, it is desired that these microsensors be inexpensive and have a very small logistics tail. In order to make the solution inexpensive, it is asserted that a common and open architecture for microsensors should be developed so that a wide range of sensor heads can be seamlessly interchanged utilizing a common piece of hardware. This not only allows the development cost to be shared among the widest possible range of applications but results in a generic sensor processor that can be configured at time of deployment. This paper describes a computing architecture developed by Sanders which employs FPGA technology married with a general purpose processor. In addition, this effort has demonstrated the applicability of FPGA technology to a widespread DoD application space and shown it to be a technology discriminator for future microsensor systems. The motivation for this common architecture for microsensors (CA/spl mu/S) the CA/spl mu/S architecture, the baseline acoustic algorithm implemented, and the results of the fielded system, which achieved more than four orders of magnitude reduction in size*weight*power over the ARL DUNES testbed, are discussed. Future work is also described.
Stephen M. Scalera, Mark Falco, Brent E. Nelson
FCCM3
1999 A CAD Suite for High-Performance FPGA Design
abstract
This paper describes the current status of a suite of CAD tools designed specifically for use by designers who are developing high-performance configurable-computing applications. The basis of this tool suite is JHDL, a design tool originally conceived as a way to experiment with Run-Time Reconfigured (RTR) designs. However, what began as a limited experiment to model RTR designs with Java has evolved into a comprehensive suite of design tools and verification aids, with these tools being used successfully to implement high-performance applications in Automated Target Recognition (ATR), sonar beamforming, and general image processing on configurable-computing systems.
Brad L. Hutchings, Peter Bellows, Joseph Hawkins, Karl S. Hemmert, Brent E. Nelson, Mike Rytting
FCCM5
1998 Frequency-Domain Sonar Processing in FPGAs and DSPs
abstract
Beamforming is a spatial filtering operation performed on data received by an array of sensors, such as antennas, microphones, or hydrophones. It provides a system with the ability to "listen" directionally even when the individual sensors in the array are omnidirectional. Over the past year we have been exploring the use of FPGA based custom computing machines for several sonar beamforming applications, including time domain beamforming (P. Graham nd B. Nelson, 1998), frequency domain beamforming, and matched field processing. In many ways sonar processing fits the criteria found by W. Mangione-Smith and B. Hutchings (1997) for good FPGA applications-the computations are data parallel, they require little control, the data sets are large (infinite streams), and the raw sensor data is at most 12 bits. However, they have three characteristics which make them challenging. First, they involve intensive arithmetic (multiply accumulates and trigonometric functions) on real and/or complex data. Second, they require significant memory support, far beyond that indicated in much previously published work. Third, the scale of the computation is large, requiring (possibly) hundreds of FPGAs and high bandwidth interconnections to meet real time constraints. We address the first issue.
Paul S. Graham, Brent E. Nelson
FCCM2
1998 FPGA-Based Sonar Processing
abstract
This paper presents the application of time-delay sonar beamforming and discusses a multi-board FPGA system for performing several variations of this beamforming method in real-time for realistic sonar arrays. Additionally, we show that our proposed FPGA system has a six to twelve times performance advantage over an equivalent system created using currently available, high-performance DSPs designed for multiprocessing systems. This performance advantage is due to the simplicity of the core calculation, the limitations of the the DSP's address calculation hardware, and the ability to customize the I/O of the FPGA to the application.
Paul S. Graham, Brent E. Nelson
FPGA2
1996 Genetic algorithms in software and in hardware-a performance analysis of workstation and custom computing machine implementations
abstract
The paper analyzes the performance differences found between the hardware and software versions of a genetic algorithm used to solve the travelling salesman problem. The hardware implementation requires 4 FPGA's on a Splash 2 board and runs at 11 MHz. The software implementation was written in C++ and executed on a 125 MHz HP PA-RISC workstation. The software run time was more than four times that of the hardware (up to 50 times as many cycles). The paper analyses the contribution made to this performance difference by the following hardware features: hard-wired control, custom address generation logic, memory hierarchy efficiency, and both fine- and course-grained parallelism. The results indicate that the major contributor to the hardware performance advantage is fine-grained parallelism-RTL-level parallelism due to operator pipelining. This alone accounts for as much as a 38X cycle-count reduction over the software in one section of the algorithm. The next major contributors include hard-wired control and custom address generation which account for as much as a 3X speedup in other sections of the algorithm. Finally, memory hierarchy inefficiencies in the software (cache misses and paging) and coarse-grained parallelism in the hardware are each shown to have lesser effect on the performance difference between the implementations.
Paul S. Graham, Brent E. Nelson
FCCM2
1996 Locality as a Visualization Tool
abstract
This brief contribution introduces a method of quantifying the locality in a given reference sequence and visually representing it as a three dimensional surface. We explore some of the properties of this formulation of locality and show the correlation between graphical features and specific reference patterns. The utility of our formulation of locality is demonstrated through two of its potential applications as a visualizations tool: characterizing and summarizing workload locality, and evaluating the effectiveness of benchmark programs in exercising memory hierarchies.
Knuth Stener Grimsrud, James K. Archibald, Richard L. Frost, Brent E. Nelson
IEEE Trans. Computers4
1993 Multiple Prefetch Adaptive Disk Caching
abstract
A new disk caching algorithm is presented that uses an adaptive prefetching scheme to reduce the average service time for disk references. Unlike schemes which simply prefetch the next sector or group of sectors, this method maintains information about the order of past disk accesses which is used to accurately predict future access sequences. The range of parameters of this scheme is explored, and its performance is evaluated through trace-driven simulation, using traces obtained from three different UNIX minicomputers. Unlike disk trace data previously described in the literature, the traces used include time stamps for each reference. With this timing information-essential for evaluating any prefetching scheme-it is shown that a cache with the adaptive prefetching mechanism can reduce the average time to service a disk request by a factor of up to three, relative to an identical disk cache without prefetching.>
Knuth Stener Grimsrud, James K. Archibald, Brent E. Nelson
IEEE Trans. Knowl. Data Eng.3
1989 Vector quantization codebook generation using simulated annealing
abstract
The authors present an algorithm for the generation of codebooks from a set of training vectors using simulated annealing. Convergence of the algorithm to the globally optimal codebook in finite time is proved, and experimental results comparing simulated annealing with Lloyd algorithms for image quantization are presented. The experimental results indicate that the proposed algorithm obtains the best known codebook for the experimental situation described by R.M. Gray and E.D. Karnin (IEEE Trans. on Inf. Theory, vol.IT-28, no.2, p.256-61, Mar. 1982). It has also been demonstrated that this technique works well for the construction of codebooks from real image data.>
J. Kelly Flanagan, Darryl Morrell, Richard L. Frost, Christopher J. Read, Brent E. Nelson
ICASSP5
1988 Processor design using path programmable logic
abstract
Path programmable logic (PPL) is a VLSI design methodology that is very efficient in the implementation of systems consisting of random logic, counters, and finite-state machines. Previous designs have shown that PPL does not efficiently allow the implementation of bus-oriented designs, such as microprocessors, arithmetic processors, and DSP (digital signal processor) chips. The authors present solutions that enable these architectures to be implemented more efficiently. The solutions are the introduction of automated routing and placement tools as well as the description of sophisticated PPL library cells. The implementation of a RISC (reduced-instruction-set computer) processor is used as a case study to verify the effectiveness of these changes to the existing PPL methodology.>
J. Kelly Flanagan, Brent E. Nelson
ICCD2
1987 A multiprogrammed parallel architecture for digital signal processing
abstract
A parallel architecture for DSP algorithms is presented along with a corresponding computation method known as address-directed computing. This approach overcomes the large amount of addressing overhead in conventional DSP programs, especially those coded in high level languages. The approach used is to decompose a computation into address and data streams. The address stream is computed using a set of n concurrently operating address processors. An architecture to implement this method is presented as well as an example of programming it.
Brent E. Nelson, J. Kelly Flanagan, Christopher J. Read
ICASSP2
1986 A bit-serial VLSI vector quantizer
abstract
A VLSI architecture for a mean residual reflected vector quantizer (MRRVQ) is presented which uses bit and word level pipelining to increase throughput. The tradeoffs between silicon area and processing speed are described along with two different architectures for implementing the vector quantizer algorithm. Throughputs for the architectures range from 450,000 vectors per second down to 28,000 vectors per second assuming a 10MHz clock rate. The chips were designed and implemented using a VLSI design methodology known as Structured Tiling (ST) which reduced the required design time to under three man-months including only two weeks for layout.
Brent E. Nelson, Christopher J. Read
ICASSP1
1984 The structure and operation of a relational database system in a cell-oriented integrated circuit design system
Lee A. Hollaar, Brent E. Nelson, Tony M. Carter, Raymond A. Lorie
DAC2