Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Karl S. Hemmert

dblp:49/4419 · also K. Scott Hemmert · DBLP profile ↗
← Back
30ranked-venue papers
9as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 8 first-author · 1 since 2021Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Reconfigurable computing and FPGAs · 51% GPUs and heterogeneous computing · 26% Processor architecture and microarchitecture · 15%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
FPGA accelerator
0.112006
Tools and techniques for performance - Architectures and APIs: assessing requirements for delivering FPGA performance to applications · SC 2006
Reconfigurable computing and FPGAs
FPGA architecture
0.112006
Embedded floating-point units in FPGAs · FPGA 2006
GPUs and heterogeneous computing › heterogeneous supercomputing
FPGA for HPC
0.112006
Tools and techniques for performance - Architectures and APIs: assessing requirements for delivering FPGA performance to applications · SC 2006
Processor architecture and microarchitecture › computer arithmetic
floating-point arithmetic
0.012006
Embedded floating-point units in FPGAs · FPGA 2006
Processor architecture and microarchitecture › arithmetic unit
multiply-accumulate unit
0.012006
Embedded floating-point units in FPGAs · FPGA 2006

Methods — techniques the papers use, named apart from their topics

island-style FPGA architecture · 0.1dense matrix multiplication · 0.1area and clock rate evaluation · 0.1FPGA · 0.1FFT · 0.1
YearPublicationVenuePosition
2022 "Smarter" NICs for faster molecular dynamics: a case study
abstract
This work evaluates the benefits of using a “smart” network interface card (SmartNIC) as a compute accelerator for the example of the MiniMD molecular dynamics proxy application. The accelerator is NVIDIA's BlueField-2 card, which includes an 8-core Arm processor along with a small amount of DRAM and storage. We test the networking and data movement performance of these cards compared to a standard Intel server host using microbenchmarks and MiniMD. In MiniMD, we identify two distinct classes of computation, namely core computation and maintenance computation, which are executed in sequence. We restructure the algorithm and code to weaken this dependence and increase task parallelism, thereby making it possible to increase utilization of the BlueField-2 concurrently with the host. We evaluate our implementation on a cluster consisting of 16 dual-socket Intel Broadwell host nodes with one BlueField-2 per host-node. Our results show that while the overall compute performance of BlueField-2 is limited, using them with a modified MiniMD algorithm allows for up to 20% speedup over the host CPU baseline with no loss in simulation accuracy.
Sara Karamati, Clay Hughes, Karl S. Hemmert, Ryan E. Grant, Whit Schonbein, Scott Levy, Thomas M. Conte, Jeffrey Young 0001, Richard W. Vuduc
IPDPS3
2017 Unveiling the Interplay Between Global Link Arrangements and Network Management Algorithms on Dragonfly Networks
abstract
Network messaging delay historically constitutes a large portion of the wall-clock time for High Performance Computing (HPC) applications, as these applications run on many nodes and involve intensive communication among their tasks. Dragonfly network topology has emerged as a promising solution for building exascale HPC systems owing to its low network diameter and large bisection bandwidth. Dragonfly includes local links that form groups and global links that connect these groups via high bandwidth optical links. Many aspects of the dragonfly network design are yet to be explored, such as the performance impact of the connectivity of the global links, i.e., global link arrangements, the bandwidth of the local and global links, or the job allocation algorithm. This paper first introduces a packet-level simulation framework to model the performance of HPC applications in detail. The proposed framework is able to simulate known MPI (message passing interface) routines as well as applications with custom-defined communication patterns for a given job placement algorithm and network topology. Using this simulation framework, we investigate the coupling between global link bandwidth and arrangements, communication pattern and intensity, job allocation and task mapping algorithms, and routing mechanisms in dragonfly topologies. We demonstrate that by choosing the right combination of system settings and workload allocation algorithms, communication overhead can be decreased by up to 44%. We also show that circulant arrangement provides up to 15% higher bisection bandwidth compared to the other arrangements, but for realistic workloads, the performance impact of link arrangements is less than 3%.
Fulya Kaplan, Ozan Tuncer, Vitus J. Leung, Karl S. Hemmert, Ayse K. Coskun
CCGrid4
2017 Two-level main memory co-design: Multi-threaded algorithmic primitives, analysis, and simulation
Michael A. Bender, Jonathan W. Berry, Simon D. Hammond, Karl S. Hemmert, Samuel McCauley, Branden Moore, Benjamin Moseley, Cynthia A. Phillips, David S. Resnick, Arun Rodrigues
J. Parallel Distributed Comput.4
2016 (SAI) Stalled, Active and Idle: Characterizing Power and Performance of Large-Scale Dragonfly Networks
abstract
Exascale networks are expected to comprise a significant part of the total monetary cost and 10-20% of the power budget allocated to exascale systems. Yet, our understanding of current and emerging workloads on these networks is limited. Left ignored, this knowledge gap likely will translate into missed opportunities for (1) improved application performance and (2) decreased power and monetary costs in next generation systems. This work targets a detailed understanding and analysis of the performance and utilization of the dragonfly network topology. Using the Structural Simulation Toolkit (SST) and a range of relevant workloads on a dragonfly topology of 110,592 nodes, we examine network design tradeoffs amongst execution time, power, bandwidth, and the number of global links. Our simulations report stalled, active and idle time on a per-port level of the fabric, in order to provide a detailed picture of future networks. The results of this work show potential savings of 3-10% of the exascale power budget and provide valuableinsights to researchers looking for new opportunities to improve performance and increase power efficiency of next generation HPC systems.
Taylor L. Groves, Ryan E. Grant, Karl S. Hemmert, Simon D. Hammond, Michael J. Levenhagen, Dorian C. Arnold
CLUSTER3
2015 Two-Level Main Memory Co-Design: Multi-threaded Algorithmic Primitives, Analysis, and Simulation
abstract
A fundamental challenge for supercomputer architecture is that processors cannot be fed data from DRAM as fast as CPUs can consume it. Therefore, many applications are memory-bandwidth bound. As the number of cores per chip increases, and traditional DDR DRAM speeds stagnate, the problem is only getting worse. A variety of non-DDR 3D memory technologies (Wide I/O 2, HBM) offer higher bandwidth and lower power by stacking DRAM chips on the processor or nearby on a silicon interposer. However, such a packaging scheme cannot contain sufficient memory capacity for a node. It seems likely that future systems will require at least two levels of main memory: high-bandwidth, low-power memory near the processor and low-bandwidth high-capacity memory further away. This near memory will probably not have significantly faster latency than the far memory. This, combined with the large size of the near memory (multiple GB) and power constraints, may make it difficult to treat it as a standard cache. In this paper, we explore some of the design space for a user-controlled multi-level main memory. We present algorithms designed for the heterogeneous bandwidth, using streaming to exploit data locality. We consider algorithms for the fundamental application of sorting. Our algorithms asymptotically reduce memory-block transfers under certain architectural parameter settings. We use and extend Sandia National Laboratories' SST simulation capability to demonstrate the relationship between increased bandwidth and improved algorithmic performance. Memory access counts from simulations corroborate predicted performance. This co-design effort suggests implementing two-level main memory systems may improve memory performance in fundamental applications.
Michael A. Bender, Jonathan W. Berry, Simon D. Hammond, Karl S. Hemmert, Samuel McCauley, Branden Moore, Benjamin Moseley, Cynthia A. Phillips, David S. Resnick, Arun Rodrigues
IPDPS4
2014 Exascale design space exploration and co-design
Sudip S. Dosanjh, Richard F. Barrett, Douglas Doerfler, Simon D. Hammond, Karl S. Hemmert, Michael A. Heroux, Paul T. Lin, Kevin T. Pedretti, Arun Rodrigues, Timothy G. Trucano, Justin Luitjens
Future Gener. Comput. Syst.5
2013 The impact of hybrid-core processors on MPI message rate
abstract
Power and energy concerns are motivating chip manufacturers to consider future hybrid-core processor designs that combine a small number of traditional cores optimized for single-thread performance with a large number of simpler cores optimized for throughput performance. This trend is likely to impact the way compute resources for network protocol processing functions are allocated and managed. In particular, the performance of MPI match processing is critical to achieving high message throughput. In this paper, we analyze the ability of simple and more complex cores to perform MPI matching operations for various scenarios in order to gain insight into how MPI implementations for future hybrid-core processors should be designed.
Brian W. Barrett, Simon D. Hammond, Ron Brightwell, Karl S. Hemmert
EuroMPI4
2012 Application-driven analysis of two generations of capability computing: the transition to multicore processors
abstract
SUMMARY Multicore processors form the basis of most traditional high performance parallel processing architectures. Early experiences with these computers showed significant performance problems, both with regard to computation and inter‐process communication. The transition from Purple, an IBM POWER5‐based machine, to Cielo, a Cray XE6, as the main capability computing platform for the United States Department of Energy's Advanced Simulation and Computing campaign provides an opportunity to reexamine these issues after experiences with a few generations of multicore‐based machines. Experiences with Purple identified some important characteristics that led to strong performance of complex scientific application programs at very large scales. Herein, we compare the performance of some Advanced Simulation and Computing mission critical applications at capability scale across this transition to multicore processors. Copyright © 2012 John Wiley & Sons, Ltd.
Mahesh Rajan, Courtenay T. Vaughan, Douglas Doerfler, Richard F. Barrett, Paul T. Lin, Kevin T. Pedretti, Karl S. Hemmert
Concurr. Comput. Pract. Exp.7
2011 Using Triggered Operations to Offload Rendezvous Messages
Brian W. Barrett, Ron Brightwell, Karl S. Hemmert, Kyle B. Wheeler, Keith D. Underwood
EuroMPI3
2011 The Impact of Injection Bandwidth Performance on Application Scalability
Kevin T. Pedretti, Ron Brightwell, Douglas Doerfler, Karl S. Hemmert, James H. Laros III
EuroMPI4
2010 Challenges for High-Performance Networking for Exascale Computing
abstract
Achieving the next three orders of magnitude performance increase to move from petascale to exascale computing will require a significant advancements in several fundamental areas. Recent studies have outlined many of the challenges in hardware and software that will be needed. In this paper, we examine these challenges with respect to high-performance networking. We describe the repercussions of anticipated changes to computing and networking hardware and discuss the impact that alternative parallel programming models will have on the network software stack. We also present some ideas on possible approaches that address some of these challenges.
Ron Brightwell, Brian W. Barrett, Karl S. Hemmert, Keith D. Underwood
ICCCN3
2010 Using Triggered Operations to Offload Collective Communication Operations
Karl S. Hemmert, Brian W. Barrett, Keith D. Underwood
EuroMPI1
2010 Fast, Efficient Floating-Point Adders and Multipliers for FPGAs
abstract
Floating-point applications are a growing trend in the FPGA community. As such, it has become critical to create floating-point units optimized for standard FPGA technology. Unfortunately, the FPGA design space is very different from the VLSI design space; thus, optimizations for FPGAs can differ significantly from optimizations for VLSI. In particular, the FPGA environment constrains the design space such that only limited parallelism can be effectively exploited to reduce latency. Obtaining the right balances between clock speed, latency, and area in FPGAs can be particularly challenging. This article presents implementation details for an IEEE-754 standard floating-point adder and multiplier for FPGAs. The designs presented here enable a Xilinx Virtex4 FPGA (-11 speed grade) to achieve 270 MHz IEEE compliant double precision floating-point performance with a 9-stage adder pipeline and 14-stage multiplier pipeline. The area requirement is approximately 500 slices for the adder and under 750 slices for the multiplier.
Karl S. Hemmert, Keith D. Underwood
ACM Trans. Reconfigurable Technol. Syst.1
2009 An application based MPI message throughput benchmark
abstract
Recent trends in high performance computing have renewed interest in the ability of platforms to sustain high message throughput rates. The continued growth in platform scale, combined with emerging application areas, are pushing platforms to support increasing message rates. Best-case message throughput has grown in previous hardware generations due to growing clock rates and software optimization techniques. However, previous work has shown that MPI receive queue length and cache hit rates can drastically impact message throughput, leading to a significantly lower worst-case message throughput. This paper introduces the Sandia message throughput benchmark which measures message throughput using a communication pattern which is neither best-case nor worst-case, but which mimics communication patterns found in real-world applications. Results on InfiniBand, Myrinet, and Cray XT platforms are presented, and suggest that message rates on some networks are greatly impacted by cache invalidation between communication phases, simultaneously sending and receiving, and by communicating with more than one peer simultaneously.
Brian W. Barrett, Karl S. Hemmert
CLUSTER2
2009 From Silicon to Science: The Long Road to Production Reconfigurable Supercomputing
abstract
The field of high performance computing (HPC) currently abounds with excitement about the potential of a broad class of things called accelerators . And, yet, few accelerator based systems are being deployed in general purpose HPC environments. Why is that? This article explores the challenges that accelerators face in the HPC world, with a specific focus on FPGA based systems. We begin with an overview of the characteristics and challenges of typical HPC systems and applications and discuss why FPGAs have the potential to have a significant impact. The bulk of the article is focused on twelve specific areas where FPGA researchers can make contributions to hasten the adoption of FPGAs in HPC environments.
Keith D. Underwood, Karl S. Hemmert, Craig D. Ulmer
ACM Trans. Reconfigurable Technol. Syst.2
2008 High message rate, NIC-based atomics: Design and performance considerations
abstract
Remote atomic memory operations are critical for achieving high-performance synchronization in tightly-coupled systems. Previous approaches to implementing atomic memory operations on high-performance networks have explored providing the primitives necessary to achieve low latency and low host processor overhead. In this paper, we explore the implementation of atomic memory operations with a focus on achieving high message rate. We believe that high message rate is a key performance characteristic that will determine the viability of a high-performance network to support future multi-petascale systems, especially those that expect to employ a partitioned global address space (PGAS) programming model. As an example, many have proposed using network interface level atomic operations to enhance the performance of the HPCC RandomAccess benchmark. This paper explores several issues relevant to the design of an atomic unit on the network interface. We explore the implications of the size of the cache as well as the associativity. Given the growing ratio of bandwidth to latency of modern host interfaces, we explore some of the interactions that impact the concurrency needed to saturate the interface.
Keith D. Underwood, Michael J. Levenhagen, Karl S. Hemmert, Ron Brightwell
CLUSTER3
2008 Architectural Modifications to Enhance the Floating-Point Performance of FPGAs
abstract
With the density of field-programmable gate arrays (FPGAs) steadily increasing, FPGAs have reached the point where they are capable of implementing complex floating-point applications. However, their general-purpose nature has limited the use of FPGAs in scientific applications that require floating-point arithmetic due to the large amount of FPGA resources that floating-point operations still require. This paper considers three architectural modifications that make floating-point operations more efficient on FPGAs. The first modification embeds floating-point multiply-add units in an island-style FPGA. While offering a dramatic reduction in area and improvement in clock rate, these embedded units are a significant change and may not be justified by the market. The next two modifications target a major component of IEEE compliant floating-point computations: variable length shifters. The first alternative to lookup tables (LUTs) for implementing the variable length shifters is a coarse-grained approach: embedded variable length shifters in the FPGA fabric. These shifters offer a significant reduction in area with a modest increase in clock rate and are smaller and more general than embedded floating-point units. The next alternative is a fine-grained approach: adding a 4:1 multiplexer unit inside a configurable logic block (CLB), in parallel to each 4-LUT. While this offers the smallest overall area improvement, it does offer a significant improvement in clock rate with only a trivial increase in the size of the CLB.
Michael J. Beauchamp, Scott Hauck, Keith D. Underwood, Karl S. Hemmert
IEEE Trans. Very Large Scale Integr. Syst.4
2007 An architecture to perform NIC based MPI matching
abstract
Modern supercomputers aggregate thousands of microprocessors through a high performance network. Many of these systems place a processor on the network interface controller (NIC) to handle some portion of the MPI processing. This processing involves traversing a linked list and invoking a matching function for each item. Although this task is critical to the performance of the system, microprocessors perform it extremely poorly. Furthermore, the traditional network processor approaches of multicore and multithreading map poorly to the problem because the list is a shared data structure. While match processing can be implemented directly in hardware, hardware implementations can be extremely inflexible and lead to extremely high risk. This paper presents a novel, programmable architecture for a processor to handle the matching function. The matching engine approaches the performance of a direct hardware implementation while maintaining a high degree of flexibility and programmability. More importantly, it requires a dramatically smaller area than a conventional processor.
Karl S. Hemmert, Keith D. Underwood, Arun Rodrigues
CLUSTER1
2007 Floating-Point Divider Design for FPGAs
abstract
Growth in floating-point applications for field-programmable gate arrays (FPGAs) has made it critical to optimize floating-point units for FPGA technology. The divider is of particular interest because the design space is large and divider usage in applications varies widely. Obtaining the right balance between clock speed, latency, throughput, and area in FPGAs can be challenging. The designs presented here cover a range of performance, throughput, and area constraints. On a Xilinx Virtex4-11 FPGA, the range includes 250-MHz IEEE compliant double precision divides that are fully pipelined to 187-MHz iterative cores. Similarly, area requirements range from 4100 slices down to a mere 334 slices
Karl S. Hemmert, Keith D. Underwood
IEEE Trans. Very Large Scale Integr. Syst.1
2006 Open Source High Performance Floating-Point Modules
abstract
Given the logic density of modern FPGAs, it is feasible to use FPGAs for floating-point applications. However, it is important that any floating-point units that are used be highly optimized. This paper introduces an open source library of highly optimized floating-point units for Xilinx FPGAs. The units are fully IEEE compliant and acheive approximately 230 MHz operation frequency for double-precision add and multiply in a Xilinx Virtex-2-Pro FPGA (-7 speed grade). This speed is acheived with a 10 stage adder pipeline and a 12 stage multiplier pipeline. The area requirement is 571 slices for the adder and 905 slices for the multiplier
Karl S. Hemmert, Keith D. Underwood
FCCM1
2006 Embedded floating-point units in FPGAs
abstract
Due to their generic and highly programmable nature, FPGAs provide the ability to implement a wide range of applications. However, it is this nonspecific nature that has limited the use of FPGAs in scientific applications that require floating-point arithmetic. Even simple floating-point operations consume a large amount of computational resources. In this paper, we introduce embedding floating-point multiply-add units in an island style FPGA. This has shown to have an average area savings of 55.0% and an average increase of 40.7% in clock rate over existing architectures.
Michael J. Beauchamp, Scott Hauck, Keith D. Underwood, Karl S. Hemmert
FPGA4
2006 Architectural Modifications to Improve Floating-Point Unit Efficiency in FPGAs
abstract
FPGAs have reached densities that can implement floating point applications, but floating-point operations still require a large amount of FPGA resources. One major component of IEEE compliant floating-point computations is variable length shifters. They account for over 30% of a double-precision floating-point adder and 25% of a double-precision multiplier. This paper introduces two alternatives for implementing these shifters. One alternative is a coarse-grained approach: embedding variable length shifters in the FPGA fabric. These units provide significant area savings with a modest clock rate improvement over existing architectures. Another alternative is a fine-grained approach: adding a 4:1 multiplexer inside the slices, in parallel to the LUTs. While providing a more modest area savings, these multiplexers provide a significant boost in clock rate with a small impact on the FPGA fabric
Michael J. Beauchamp, Scott Hauck, Keith D. Underwood, Karl S. Hemmert
FPL4
2006 Tools and techniques for performance - Architectures and APIs: assessing requirements for delivering FPGA performance to applications
abstract
Reconfigurable computing leveraging field programmable gate arrays (FPGAs) is one of many accelerator technologies that are being investigated for application to high performance computing (HPC). Like most accelerators, FPGAs are very efficient at both dense matrix multiplication and FFT computations, but two important aspects of how to deliver that performance to applications have received too little attention. First, the standard API for important compute kernels hides parallelism from the system. Second, the issue of system architecture is virtually never addressed. This paper explores both issues and their implications for applications. We find that high bandwidth, low latency connectivity can be important, but the right API can be even more important.
Keith D. Underwood, Karl S. Hemmert, Craig D. Ulmer
SC2
2005 A Comparison of Floating Point and Logarithmic Number Systems for FPGAs
abstract
There have been many papers proposing the use of logarithmic numbers (LNS) as an alternative to floating point because of simpler multiplication, division and exponentiation computations. However, this advantage comes at the cost of complicated, inexact addition and subtraction, as well as the need to convert between the formats. In this work, we created a parameterized LNS library of computational units and compared them to an existing floating point library. Specifically, we considered multiplication, division, addition, subtraction, and format conversion to determine when one format should be used over the other and when it is advantageous to change formats during a calculation.
Michael Haselman, Michael J. Beauchamp, Aaron Wood, Scott Hauck, Keith D. Underwood, Karl S. Hemmert
FCCM6
2005 An Analysis of the Double-Precision Floating-Point FFT on FPGAs
abstract
Advances in FPGA technology have led to dramatic improvements in double precision floating-point performance. Modern FPGAs boast several GigaFLOPs of raw computing power. Unfortunately, this computing power is distributed across 30 floating-point units with over 10 cycles of latency each. The user must find two orders of magnitude more parallelism than is typically exploited in a single microprocessor; thus, it is not clear that the computational power of FPGAs can be exploited across a wide range of algorithms. This paper explores three implementation alternatives for the fast Fourier transform (FFT) on FPGAs. The algorithms are compared in terms of sustained performance and memory requirements for various FFT sizes and FPGA sizes. The results indicate that FPGAs are competitive with microprocessors in terms of performance and that the "correct" FFT implementation varies based on the size of the transform and the size of the FPGA.
Karl S. Hemmert, Keith D. Underwood
FCCM1
2004 Closing the Gap: CPU and FPGA Trends in Sustainable Floating-Point BLAS Performance
abstract
Field programmable gate arrays (FPGAs) have long been an attractive alternative to microprocessors for computing tasks - as long as floating-point arithmetic is not required. Fueled by the advance of Moore's law, FPGAs are rapidly reaching sufficient densities to enhance peak floating-point performance as well. The question, however, is how much of this peak performance can be sustained. This paper examines three of the basic linear algebra subroutine (BLAS) functions: vector dot product, matrix-vector multiply, and matrix multiply. A comparison of microprocessors, FPGAs, and reconfigurable computing platforms is performed for each operation. The analysis highlights the amount of memory bandwidth and internal storage needed to sustain peak performance with FPGAs. This analysis considers the historical context of the last six years and is extrapolated for the next six years.
Keith D. Underwood, Karl S. Hemmert
FCCM2
2003 Issues in debugging highly parallel FPGA-based applications derived from source code
abstract
Using high-level synthesis tools to map programs written in general-purpose languages to FPGA hardware has grown in popularity and it is becoming necessary to provide comprehensive debugging tools in order to verify the correctness of the synthesized hardware. Currently, post-synthesis debugging is done at the circuit level. This paper discusses the issues, as well as some early results, of creating a source level debugger for hardware synthesized from source code. This study is meant to provide some insight into what needs to be added or built into synthesizing compilers in order to allow debug of a synthesized circuit at the source level, which will provide the programmer with a familiar view of the program being debugged.
Karl S. Hemmert, Brad L. Hutchings
ASP-DAC1
2003 Source Level Debugger for the Sea Cucumber Synthesizing Compiler
abstract
With the growing popularity of using high-level synthesis tools to map programs written in general-purpose programming languages to FPGA (field programmable gate array) hardware, it has become necessary to provide comprehensive, intuitive debugging tools in order to verify the correctness of the synthesized hardware. The difficulty in creating these tools lies in the fact that typical synthesizing compilers provide no information about how the source code is mapped to hardware. This paper discusses the creation of a debugger for the Sea Cucumber synthesizing compiler used to explore the issues associated with providing information about a circuit in the context of the original source code, thus making the debugging process more intuitive.
Karl S. Hemmert, Justin L. Tripp, Brad L. Hutchings, Preston A. Jackson
FCCM1
2001 An Application-Specific Compiler for High-Speed Binary Image Morphology
Karl S. Hemmert, Brad L. Hutchings, Anshul Malvi
FCCM1
1999 A CAD Suite for High-Performance FPGA Design
abstract
This paper describes the current status of a suite of CAD tools designed specifically for use by designers who are developing high-performance configurable-computing applications. The basis of this tool suite is JHDL, a design tool originally conceived as a way to experiment with Run-Time Reconfigured (RTR) designs. However, what began as a limited experiment to model RTR designs with Java has evolved into a comprehensive suite of design tools and verification aids, with these tools being used successfully to implement high-performance applications in Automated Target Recognition (ATR), sonar beamforming, and general image processing on configurable-computing systems.
Brad L. Hutchings, Peter Bellows, Joseph Hawkins, Karl S. Hemmert, Brent E. Nelson, Mike Rytting
FCCM4