VLDB 2026 Research / reviewers in the wild / expert
Keith D. Underwood
dblp:52/5042
· DBLP profile ↗
54ranked-venue papers
16as first author
1since 2021 · last 2023
0009-0001-0078-9959ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 50 · 16 first-author · 1 since 2021Computer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
9 papers |
Reconfigurable computing and FPGAs · 31% Parallel and multicore computing · 14% Performance modeling and evaluation · 13% |
Topics — the 17 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2007 | Evaluating NIC hardware requirements to achieve high message rate PGAS support on multi-core processors · SC 2007 |
Parallel and multicore computing › parallel programming models › distributed memory programming models
partitioned global address space |
0.1 | 1 | 2007 | Evaluating NIC hardware requirements to achieve high message rate PGAS support on multi-core processors · SC 2007 |
Performance modeling and evaluation › simulation
architectural simulation |
0.1 | 1 | 2006 | Poster reception - The structural simulation toolkit: exploring novel architectures · SC 2006 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.1 | 1 | 2006 | Tools and techniques for performance - Architectures and APIs: assessing requirements for delivering FPGA performance to applications · SC 2006 |
Reconfigurable computing and FPGAs
FPGA architecture |
0.1 | 1 | 2006 | Embedded floating-point units in FPGAs · FPGA 2006 |
GPUs and heterogeneous computing › heterogeneous supercomputing
FPGA for HPC |
0.1 | 1 | 2006 | Tools and techniques for performance - Architectures and APIs: assessing requirements for delivering FPGA performance to applications · SC 2006 |
Reconfigurable computing and FPGAs
high-performance reconfigurable computing |
0.1 | 1 | 2006 | Reconfigurable supercomputing - Is high-performance reconfigurable computing the next supercomputing paradigm? · SC 2006 |
High-performance computing
scientific computing |
0.1 | 1 | 2006 | Reconfigurable supercomputing - Is high-performance reconfigurable computing the next supercomputing paradigm? · SC 2006 |
Performance modeling and evaluation
simulation |
0.1 | 1 | 2006 | Poster reception - The structural simulation toolkit: exploring novel architectures · SC 2006 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.0 | 1 | 2001 | Cost effectiveness of an adaptable computing cluster · SC 2001 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 1 | 2007 | Evaluating NIC hardware requirements to achieve high message rate PGAS support on multi-core processors · SC 2007 |
Reconfigurable computing and FPGAs › FPGA arithmetic
FPGA floating-point arithmetic |
0.0 | 1 | 1998 | Implementation of IEEE Single-Precision Floating-Point Operations on FPGAs (Abstract) · FPGA 1998 |
Processor architecture and microarchitecture › computer arithmetic
floating-point arithmetic |
0.0 | 1 | 2006 | Embedded floating-point units in FPGAs · FPGA 2006 |
Memory systems
memory system modeling |
0.0 | 1 | 2006 | Poster reception - The structural simulation toolkit: exploring novel architectures · SC 2006 |
Processor architecture and microarchitecture › arithmetic unit
multiply-accumulate unit |
0.0 | 1 | 2006 | Embedded floating-point units in FPGAs · FPGA 2006 |
Distributed systems › distributed computing theory
network model |
0.0 | 1 | 2006 | Poster reception - The structural simulation toolkit: exploring novel architectures · SC 2006 |
High-performance computing
performance optimization at scale |
0.0 | 1 | 2001 | Cost effectiveness of an adaptable computing cluster · SC 2001 |
Methods — techniques the papers use, named apart from their topics
randomaccess benchmark · 0.1SHMEM · 0.1PGAS · 0.1island-style FPGA architecture · 0.1discrete-event simulation · 0.1dense matrix multiplication · 0.1area and clock rate evaluation · 0.1FPGA · 0.1FFT · 0.1instruction mix benchmarking · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Not all applications have boring communication patterns: Profiling message matching with BMMabstractSummary Message matching within MPI is an important performance consideration for applications that utilize two‐sided semantics. In this work, we present an instrumentation of the CrayMPI library that allows the collection of detailed message‐matching statistics as well as an implementation of hashed matching in software. We use this functionality to profile key DOE applications with complex communication patterns to determine under what circumstances an application might benefit from hardware offload capabilities within the NIC to accelerate message matching. We find that there are several applications and libraries that exhibit sufficiently long match list lengths to motivate a Binned Message Matching approach. Taylor L. Groves, Naveen Ravichandrasekaran, Brandon Cook 0001, Noel Keen, David Trebotich, Nicholas J. Wright, Robert Alverson, Duncan Roweth, Keith D. Underwood |
Concurr. Comput. Pract. Exp. | 9 |
| 2013 | Evaluating on-die interconnects for a 4 TB/s routerabstractFuture high performance computing networks will exploit routers with both high port counts and high port bandwidth. Scalable on-die interconnects will be needed to insure that the router can sustain its full bandwidth for a variety of traffic patterns. Otherwise, blocking behavior within a router can be encountered by a variety of challenging HPC traffic patterns. We examine the router on-die interconnect problem in the context of a hypothetical 4 TB/s router, including throughput on various traffic patterns and die area considerations. The results indicate that the on-die topologies that have been used in the past require either too much area, or achieve too little performance. We present three topologies (two adaptations of existing topologies, and one new topology) that can deliver area-efficient sustained performance. Keith D. Underwood, Eric Borch, John Sizer, Timothy Stremcha, Michael Strom |
ICS | 1 |
| 2012 | Exploiting communication and packaging locality for cost-effective large scale networksabstractAs processing power increases, maintaining the balance between network and computing is becoming increasingly difficult. The two major contributors to this imbalance are the cost and the power of high bandwidth networks, and both network cost and power are heavily impacted by the type of signaling used. Reducing the length of a network link leads to both lower cost and lower power. Unfortunately, the low dimension mesh and torus topologies that enable the shortest physical links also scale poorly in terms of hop count and global bandwidth. In contrast, topologies with low hop count and high global bandwidth have a large fraction of physical links that are several meters long. We propose the cube collective topology --- a hierarchical topology that uses a mesh topology locally to minimize link length and an all-to-all topology globally to minimize global hops. The result is that over 80% of the links can be very short (under 1 meter). This enables significant reductions in both network cost and network power, while still providing a balance of high global and high local bandwidth. Keith D. Underwood, Eric Borch |
ICS | 1 |
| 2012 | A Low Impact Flow Control Implementation for Offload Communication Interfaces
Brian W. Barrett, Ron Brightwell, Keith D. Underwood |
EuroMPI | 3 |
| 2011 | Using Triggered Operations to Offload Rendezvous Messages
Brian W. Barrett, Ron Brightwell, Karl S. Hemmert, Kyle B. Wheeler, Keith D. Underwood |
EuroMPI | 5 |
| 2011 | Scientific Application Demands on a Reconfigurable Functional Unit InterfaceabstractModern scientific applications are large, complex, and highly parallel they are commonly executed on supercomputers with tens of thousands of processors. Yet these applications still commonly require weeks or even months to execute. Thus, single-thread performance remains a concern for highly parallel scientific applications. Adding a reconfigurable accelerator to each CPU could improve system performance; however, scientific applications have design constraints that differ from most application domains commonly accelerated by reconfigurable logic. In this article, we discuss the constraints imposed by scientific applications on the computation model, the accelerator architecture, and the accelerator’s communication interface with the CPU. Based on these constraints and application analysis, we have previously proposed adding a Reconfigurable Functional Unit (RFU) to accelerate integer graphs that calculate complex memory addresses. In this work, we now propose a flexible multi-instruction interface technique that allows dataflow graphs implemented on the RFU to access a large number of inputs and outputs with minor CPU datapath modifications. We present an in-depth examination of the performance effects of different communication interfaces that use this technique, and select one that best matches the needs of Sandia’s scientific applications. Although RFU execution overall improves performance, we also isolate two key negative performance effects introduced by aggregating CPU instructions into dataflow graphs: delayed issue and graph serialization. Finally, to demonstrate the marketability of an RFU beyond scientific applications, we reanalyze the proposed interfaces using the SPEC-fp benchmark suite. We show that although choosing an interface based on SPEC-fp needs is detrimental to Sandia application performance, choosing an interface based on Sandia demands works well for more general-purpose applications. Kyle Rupnow, Keith D. Underwood, Katherine Compton |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2010 | Challenges for High-Performance Networking for Exascale ComputingabstractAchieving the next three orders of magnitude performance increase to move from petascale to exascale computing will require a significant advancements in several fundamental areas. Recent studies have outlined many of the challenges in hardware and software that will be needed. In this paper, we examine these challenges with respect to high-performance networking. We describe the repercussions of anticipated changes to computing and networking hardware and discuss the impact that alternative parallel programming models will have on the network software stack. We also present some ideas on possible approaches that address some of these challenges. Ron Brightwell, Brian W. Barrett, Karl S. Hemmert, Keith D. Underwood |
ICCCN | 4 |
| 2010 | Using Triggered Operations to Offload Collective Communication Operations
Karl S. Hemmert, Brian W. Barrett, Keith D. Underwood |
EuroMPI | 3 |
| 2010 | Performance evaluation of the Red Storm dual-core upgradeabstractAbstract In 2007, the Cray Red Storm system at Sandia National Laboratories completed an upgrade of the processor and network hardware. Single‐core 2.0 GHz AMD Opteron processors were replaced with dual‐core 2.4 GHz AMD Opterons, while the network interface hardware was upgraded from a sustained rate of 1.1 GBps to 2.0 GBps (without changing the router link rates). These changes more than doubled the theoretical peak floating‐point performance of the compute nodes and doubled the bandwidth performance of the network. This paper provides an analysis of the impact of this upgrade on the performance of several applications and micro‐benchmarks. Performance results show that the additional core provides a performance boost of 20–50% for real applications on a fixed problem size per‐socket basis on up to 2048 cores and that scalability is impacted relatively little by the upgrade. Copyright © 2009 John Wiley & Sons, Ltd. Ron Brightwell, Keith D. Underwood, Courtenay T. Vaughan, Joel Stevenson |
Concurr. Comput. Pract. Exp. | 2 |
| 2010 | Fast, Efficient Floating-Point Adders and Multipliers for FPGAsabstractFloating-point applications are a growing trend in the FPGA community. As such, it has become critical to create floating-point units optimized for standard FPGA technology. Unfortunately, the FPGA design space is very different from the VLSI design space; thus, optimizations for FPGAs can differ significantly from optimizations for VLSI. In particular, the FPGA environment constrains the design space such that only limited parallelism can be effectively exploited to reduce latency. Obtaining the right balances between clock speed, latency, and area in FPGAs can be particularly challenging. This article presents implementation details for an IEEE-754 standard floating-point adder and multiplier for FPGAs. The designs presented here enable a Xilinx Virtex4 FPGA (-11 speed grade) to achieve 270 MHz IEEE compliant double precision floating-point performance with a 9-stage adder pipeline and 14-stage multiplier pipeline. The area requirement is approximately 500 slices for the adder and under 750 slices for the multiplier. Karl S. Hemmert, Keith D. Underwood |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2009 | From Silicon to Science: The Long Road to Production Reconfigurable SupercomputingabstractThe field of high performance computing (HPC) currently abounds with excitement about the potential of a broad class of things called accelerators . And, yet, few accelerator based systems are being deployed in general purpose HPC environments. Why is that? This article explores the challenges that accelerators face in the HPC world, with a specific focus on FPGA based systems. We begin with an overview of the characteristics and challenges of typical HPC systems and applications and discuss why FPGAs have the potential to have a significant impact. The bulk of the article is focused on twelve specific areas where FPGA researchers can make contributions to hasten the adoption of FPGAs in HPC environments. Keith D. Underwood, Karl S. Hemmert, Craig D. Ulmer |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2008 | High message rate, NIC-based atomics: Design and performance considerationsabstractRemote atomic memory operations are critical for achieving high-performance synchronization in tightly-coupled systems. Previous approaches to implementing atomic memory operations on high-performance networks have explored providing the primitives necessary to achieve low latency and low host processor overhead. In this paper, we explore the implementation of atomic memory operations with a focus on achieving high message rate. We believe that high message rate is a key performance characteristic that will determine the viability of a high-performance network to support future multi-petascale systems, especially those that expect to employ a partitioned global address space (PGAS) programming model. As an example, many have proposed using network interface level atomic operations to enhance the performance of the HPCC RandomAccess benchmark. This paper explores several issues relevant to the design of an atomic unit on the network interface. We explore the implications of the size of the cache as well as the associativity. Given the growing ratio of bandwidth to latency of modern host interfaces, we explore some of the interactions that impact the concurrency needed to saturate the interface. Keith D. Underwood, Michael J. Levenhagen, Karl S. Hemmert, Ron Brightwell |
CLUSTER | 1 |
| 2008 | Architectural Modifications to Enhance the Floating-Point Performance of FPGAsabstractWith the density of field-programmable gate arrays (FPGAs) steadily increasing, FPGAs have reached the point where they are capable of implementing complex floating-point applications. However, their general-purpose nature has limited the use of FPGAs in scientific applications that require floating-point arithmetic due to the large amount of FPGA resources that floating-point operations still require. This paper considers three architectural modifications that make floating-point operations more efficient on FPGAs. The first modification embeds floating-point multiply-add units in an island-style FPGA. While offering a dramatic reduction in area and improvement in clock rate, these embedded units are a significant change and may not be justified by the market. The next two modifications target a major component of IEEE compliant floating-point computations: variable length shifters. The first alternative to lookup tables (LUTs) for implementing the variable length shifters is a coarse-grained approach: embedded variable length shifters in the FPGA fabric. These shifters offer a significant reduction in area with a modest increase in clock rate and are smaller and more general than embedded floating-point units. The next alternative is a fine-grained approach: adding a 4:1 multiplexer unit inside a configurable logic block (CLB), in parallel to each 4-LUT. While this offers the smallest overall area improvement, it does offer a significant improvement in clock rate with only a trivial increase in the size of the CLB. Michael J. Beauchamp, Scott Hauck, Keith D. Underwood, Karl S. Hemmert |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2007 | An architecture to perform NIC based MPI matchingabstractModern supercomputers aggregate thousands of microprocessors through a high performance network. Many of these systems place a processor on the network interface controller (NIC) to handle some portion of the MPI processing. This processing involves traversing a linked list and invoking a matching function for each item. Although this task is critical to the performance of the system, microprocessors perform it extremely poorly. Furthermore, the traditional network processor approaches of multicore and multithreading map poorly to the problem because the list is a shared data structure. While match processing can be implemented directly in hardware, hardware implementations can be extremely inflexible and lead to extremely high risk. This paper presents a novel, programmable architecture for a processor to handle the matching function. The matching engine approaches the performance of a direct hardware implementation while maintaining a high degree of flexibility and programmability. More importantly, it requires a dramatically smaller area than a conventional processor. Karl S. Hemmert, Keith D. Underwood, Arun Rodrigues |
CLUSTER | 2 |
| 2007 | Scientific Application Acceleration with Reconfigurable Functional UnitsabstractWhile scientific applications in the past were limited by floating point computations, modern scientific applications use more unstructured formulations. These applications have a significant percentage of integer computation - increasingly a limiting factor in scientific application performance. In real scientific applications employed at Sandia National Labs, integer computations constitute on average 37% of the application operations, forming large and complex dataflow graphs. Reconfigurable functional units (RFUs) are a particularly attractive accelerator for these graphs because they can potentially accelerate many unique graphs with a small amount of additional hardware. In this study, we analyze application traces of Sandia's scientific applications and the SPEC-FP benchmark suite. First we select a set of dataflow graphs to accelerate using the RFU, then we use execution-based simulation to determine the acceleration potential of the applications when using an RFU. On average, a set of 32 or fewer graphs is sufficient to capture the dataflow behavior of 30% of the integer computation, and more than half of Sandia applications show an improvement of 5% or more. Kyle Rupnow, Keith D. Underwood, Katherine Compton |
FCCM | 2 |
| 2007 | Simulating Red Storm: Challenges and Successes in Building a System SimulationabstractSupercomputers are increasingly complex systems merging conventional microprocessors with system on a chip level designs that provide the network interface and router. At Sandia National Labs, we are developing a simulator to explore the complex interactions that occur at the system level This paper presents an overview of the simulation framework with a focus on the enhancements needed to transform traditional simulation tools into a simulator capable of modeling system level hardware interactions and running native software. Initial validation results demonstrate simulated performance that matches the Cray Red Storm system installed at Sandia. In addition, we include a "what if" study of performance implications on the Red Storm network interface. Keith D. Underwood, Michael J. Levenhagen, Arun Rodrigues |
IPDPS | 1 |
| 2007 | Analyzing the Scalability of Graph Algorithms on EldoradoabstractThe Cray MTA-2 system provides exceptional performance on a variety of sparse graph algorithms. Unfortunately, it was an extremely expensive platform. Cray is preparing an Eldorado platform that leverages the Cray XT3 network and system infrastructure while integrating a new revision of the MTA-2 processors that is pin compatible with the AMD Opteron socket. Unlike the MTA-2, this platform will have a more constrained network bisection bandwidth and will pay a high penalty for random memory accesses. This work assesses the hardware level scalability of the Eldorado platform on several graph algorithms. Keith D. Underwood, Megan Vance, Jonathan W. Berry, Bruce Hendrickson |
IPDPS | 1 |
| 2007 | Evaluating NIC hardware requirements to achieve high message rate PGAS support on multi-core processorsabstractPartitioned global address space (PGAS) programming models have been identified as one of the few viable approaches for dealing with emerging many-core systems. These models tend to generate many small messages, which requires specific support from the network interface hardware to enable efficient execution. In the past, Cray included E-registers on the Cray T3E to support the SHMEM API; however, with the advent of multi-core processors, the balance of computation to communication capabilities has shifted toward computation. This paper explores the message rates that are achievable with multi-core processors and simplified PGAS support on a more conventional network interface. For message rate tests, we find that simple network interface hardware is more than sufficient. We also find that even typical data distributions, such as cyclic or block-cyclic, do not need specialized hardware support. Finally, we assess the impact of such support on the well known RandomAccess benchmark. 1. Keith D. Underwood, Michael J. Levenhagen, Ron Brightwell |
SC | 1 |
| 2007 | Floating-Point Divider Design for FPGAsabstractGrowth in floating-point applications for field-programmable gate arrays (FPGAs) has made it critical to optimize floating-point units for FPGA technology. The divider is of particular interest because the design space is large and divider usage in applications varies widely. Obtaining the right balance between clock speed, latency, throughput, and area in FPGAs can be challenging. The designs presented here cover a range of performance, throughput, and area constraints. On a Xilinx Virtex4-11 FPGA, the range includes 250-MHz IEEE compliant double precision divides that are fully pipelined to 187-MHz iterative cores. Similarly, area requirements range from 4100 slices down to a mere 334 slices Karl S. Hemmert, Keith D. Underwood |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2006 | A Simple Synchronous Distributed-Memory Algorithm for the HPCC RandomAccess BenchmarkabstractThe RandomAccess benchmark as defined by the High Performance Computing Challenge (HPCC) tests the speed at which a machine can update the elements of a table spread across global system memory, as measured in billions (giga) of updates per second (GUPS). The parallel implementation provided by HPCC typically performs poorly on distributed-memory machines, due to updates requiring numerous small point-to-point messages between processors. We present an alternative algorithm which treats the collection of P processors as a hypercube, aggregating data so that larger messages are sent, and routing individual datums through dimensions of the hypercube to their destination processor. The algorithm's computation (the GUP count) scales linearly with P while its communication overhead scales as log2(P), thus enabling better performance on large numbers of processors. The new algorithm achieves a GUPS rate of 19.98 on 8192 processors of Sandia's Red Storm machine, compared to 1.02 for the HPCCprovided algorithm on 10350 processors. We also illustrate how GUPS performance varies with the benchmark's specification of its "look-ahead" parameter. As expected, parallel performance degrades for small look-ahead values, and improves dramatically for large values. Steven J. Plimpton, Ron Brightwell, Courtenay T. Vaughan, Keith D. Underwood |
CLUSTER | 4 |
| 2006 | Fine-Grained Message Pipelining for Improved MPI PerformanceabstractBy its nature, MPI leads to coarse grained communications. This is because all current MPI implementations deliver two orders of magnitude more bandwidth for large message sizes (kilobytes) than small message sizes (bytes). This translates into applications that bundle their small communications into larger communications whenever possible. In modern implementations, this sacrifice in the granularity of communication translates directly into a sacrifice in the granularity of synchronization. MPI requires that the entire message arrive before any of the data can be delivered to the application, because message completion is the only synchronization semantic the network can expose to the processor. This paper explores the implications of providing synchronization between the network and the processor at the memory word level using a mechanism such as Full/Empty Bits. This enables the application to begin computing as soon as the data for the first memory referenced has arrived without having to wait for all of the data in the message Arun Rodrigues, Kyle B. Wheeler, Peter M. Kogge, Keith D. Underwood |
CLUSTER | 4 |
| 2006 | Open Source High Performance Floating-Point ModulesabstractGiven the logic density of modern FPGAs, it is feasible to use FPGAs for floating-point applications. However, it is important that any floating-point units that are used be highly optimized. This paper introduces an open source library of highly optimized floating-point units for Xilinx FPGAs. The units are fully IEEE compliant and acheive approximately 230 MHz operation frequency for double-precision add and multiply in a Xilinx Virtex-2-Pro FPGA (-7 speed grade). This speed is acheived with a 10 stage adder pipeline and a 12 stage multiplier pipeline. The area requirement is 571 slices for the adder and 905 slices for the multiplier Karl S. Hemmert, Keith D. Underwood |
FCCM | 2 |
| 2006 | Embedded floating-point units in FPGAsabstractDue to their generic and highly programmable nature, FPGAs provide the ability to implement a wide range of applications. However, it is this nonspecific nature that has limited the use of FPGAs in scientific applications that require floating-point arithmetic. Even simple floating-point operations consume a large amount of computational resources. In this paper, we introduce embedding floating-point multiply-add units in an island style FPGA. This has shown to have an average area savings of 55.0% and an average increase of 40.7% in clock rate over existing architectures. Michael J. Beauchamp, Scott Hauck, Keith D. Underwood, Karl S. Hemmert |
FPGA | 3 |
| 2006 | Architectural Modifications to Improve Floating-Point Unit Efficiency in FPGAsabstractFPGAs have reached densities that can implement floating point applications, but floating-point operations still require a large amount of FPGA resources. One major component of IEEE compliant floating-point computations is variable length shifters. They account for over 30% of a double-precision floating-point adder and 25% of a double-precision multiplier. This paper introduces two alternatives for implementing these shifters. One alternative is a coarse-grained approach: embedding variable length shifters in the FPGA fabric. These units provide significant area savings with a modest clock rate improvement over existing architectures. Another alternative is a fine-grained approach: adding a 4:1 multiplexer inside the slices, in parallel to the LUTs. While providing a more modest area savings, these multiplexers provide a significant boost in clock rate with a small impact on the FPGA fabric Michael J. Beauchamp, Scott Hauck, Keith D. Underwood, Karl S. Hemmert |
FPL | 3 |
| 2006 | Scientific applications vs. SPEC-FP: a comparison of program behaviorabstractMany modern scientific applications execute on massively parallel collections of microprocessors. Supercomputers such as the Cray XT3 (Red Storm) and Blue Gene/L support thousands to tens of thousands of processors per parallel job. However, individual microprocessor performance remains a critical component of overall performance. Traditional approaches to improve scientific application performance concentrate on floating-point (FP) instructions; however, our studies show that in the scientific applications used at Sandia National Labs, integer instructions constitute a large and critical part of the instruction mix. Although the SPEC-FP benchmark suite is considered representative of FP workloads, it has a much smaller proportion of integer computation instructions than the Sandia scientific applications, with 22.9% as compared to 36.9%. Integer instructions in Sandia applications also behave differently than in SPEC-FP. Integer instruction outputs are reused 8.8x to 13.1x more often in SPEC-FP benchmarks, and integer dataflow in Sandia applications is more complex than in the SPEC-FP suite. In this work, we examine common dataflow and usage patterns of integer instructions---information essential to develop hardware techniques to accelerate critical scientific applications. We present statistics for SPEC-FP and Sandia applications, summarizing integer computation usage and the size, shape and interface (number of inputs/outputs) of dataflow graphs. Kyle Rupnow, Arun Rodrigues, Keith D. Underwood, Katherine Compton |
ICS | 3 |
| 2006 | A preliminary analysis of the InfiniPath and XD1 network interfacesabstractTwo recently delivered systems have begun a new trend in cluster interconnects. Both the InfiniPath network from PathScale, Inc., and the rapidarray fabric in the XDI system from Cray, Inc., leverage commodity network fabrics while customizing the network interface in an attempt to add value specifically for the high performance computing (HPC) cluster market. Both network interfaces are compatible with standard InfiniBand (IB) switches, but neither use the traditional programming interfaces to support MPI. Another fundamental difference between these networks and other modern network adapters is that much of the processing needed for the network protocol stack is performed on the host processor(s) rather than by the network interface itself. This approach stands in stark contrast to the current direction of most high-performance networking activities, which is to offload as much protocol processing as possible to the network interface. In this paper, we provide an initial performance comparison of the two partially custom networks (PathScale's InfiniPath and Cray's XDI) with a more commodity network (standard IB) and a more custom network (Quadrics Elan4). Our evaluation includes several micro-benchmark results as well as some initial application performance data Ron Brightwell, Douglas Doerfler, Keith D. Underwood |
IPDPS | 3 |
| 2006 | Reconfigurable supercomputing - Is high-performance reconfigurable computing the next supercomputing paradigm?abstractHigh-Performance Reconfigurable Computers (HPRCs) based on the combination of conventional processors and FPGAs have been gaining attention in the past few years. Their benefits were particularly harnessed in compute-intensive integer applications. However, there has been doubt that the same benefits can be attained for general scientific applications. Fortunately, the trend in reconfigurable chip sizes and diversity of resources may be relieving some of those concerns. Yet, with the hardware reconfigurability, it is feared that domain scientists have to learn how to design hardware if they were to use such machines effectively. In order to address the overarching question, this panel will address the following questions: Can FPGAs deliver order-of-magnitude performance gains in scientific floating-point applications in the foreseeable future? Can programming HPRCs programmability become similar to that of HPCs in its level of difficulty? What are the major developments in the industry or the community that make all this possible? Tarek A. El-Ghazawi, Dave Bennett, Daniel S. Poznanovic, Allan Cantle, Keith D. Underwood, Rob Pennington, Duncan A. Buell, Alan D. George, Volodymyr V. Kindratenko |
SC | 5 |
| 2006 | Poster reception - The structural simulation toolkit: exploring novel architecturesabstractExploring novel computer system designs requires modeling the complex interactions between processor, memory, and network. The Structural Simulation Toolkit (SST) has been developed to explore innovations in both the programming models and hardware implementation of highly concurrent systems. The Toolkit's modular design allows extensive exploration of system parameters while maximizing code reuse and provides an explicit separation of instruction interpretation from microarchitectural timing. This is built upon a high performance hybrid discrete event framework. The SST has modeled a variety of systems, from processor-in-memory to CMP and MPP. It has examined a variety of hardware and software issues in the context of HPC.This poster presents an overview of the SST. Several of its models for processors, memory systems, and networks will be detailed. Its software stack, including support for MPI and OpenMP, will also be covered. Performance results and current directions for the SST will also be shown. Arun Rodrigues, Richard C. Murphy, Peter M. Kogge, Keith D. Underwood |
SC | 4 |
| 2006 | Tools and techniques for performance - Architectures and APIs: assessing requirements for delivering FPGA performance to applicationsabstractReconfigurable computing leveraging field programmable gate arrays (FPGAs) is one of many accelerator technologies that are being investigated for application to high performance computing (HPC). Like most accelerators, FPGAs are very efficient at both dense matrix multiplication and FFT computations, but two important aspects of how to deliver that performance to applications have received too little attention. First, the standard API for important compute kernels hides parallelism from the system. Second, the issue of system architecture is virtually never addressed. This paper explores both issues and their implications for applications. We find that high bandwidth, low latency connectivity can be important, but the right API can be even more important. Keith D. Underwood, Karl S. Hemmert, Craig D. Ulmer |
SC | 1 |
| 2005 | Implementation and Performance of Portals 3.3 on the Cray XT3abstractThe Portals data movement interface was developed at Sandia National Laboratories in collaboration with the University of New Mexico over the last ten years. Portals is intended to provide the functionality necessary to scale a distributed memory parallel computing system to thousands of nodes. Previous versions of Portals ran on several large-scale machines, including a 1024-node nCUBE-2, a 1800-node Intel Paragon, and the 4500-node Intel ASCI Red machine. The latest version of Portals was initially developed for an 1800-node Linux/Myrinet cluster and has since been adopted by Cray as the lowest-level network programming interface for their XT3 platform. In this paper, we describe the implementation of Portals 3.3 on the Cray XT3 and present some initial performance results from several micro-benchmark tests. Despite some limitations, the implementation of Portals is able to achieve a zero-length one-way latency of under six microseconds and a uni-directional bandwidth of more than 1.1 GB/s Ron Brightwell, Trammell Hudson, Kevin T. Pedretti, Rolf Riesen, Keith D. Underwood |
CLUSTER | 5 |
| 2005 | Accelerating List Management for MPIabstractThe latency and throughput of MPI messages are critically important to a range of parallel scientific applications. In many modern networks, both of these performance characteristics are largely driven by the performance of a processor on the network interface. Because of the semantics of MPI, this embedded processor is forced to traverse a linked list of posted receives each time a messages is received. As this list grows long, the latency of message reception grows and the throughput of MPI messages decreases. This paper presents a novel hardware feature to handle list management functions on a network interface. By moving functions such as list insertion, list traversal, and list deletion to the hardware unit, latencies are decreased by up to 20% in the zero length queue case with dramatic improvements in the presence of long queues. Similarly, the throughput is increased by up to 10% in the zero length queue case and by nearly 100% in the presence queues of 30 messages Keith D. Underwood, Arun Rodrigues, K. S. Hemmeit |
CLUSTER | 1 |
| 2005 | A Comparison of Floating Point and Logarithmic Number Systems for FPGAsabstractThere have been many papers proposing the use of logarithmic numbers (LNS) as an alternative to floating point because of simpler multiplication, division and exponentiation computations. However, this advantage comes at the cost of complicated, inexact addition and subtraction, as well as the need to convert between the formats. In this work, we created a parameterized LNS library of computational units and compared them to an existing floating point library. Specifically, we considered multiplication, division, addition, subtraction, and format conversion to determine when one format should be used over the other and when it is advantageous to change formats during a calculation. Michael Haselman, Michael J. Beauchamp, Aaron Wood, Scott Hauck, Keith D. Underwood, Karl S. Hemmert |
FCCM | 5 |
| 2005 | An Analysis of the Double-Precision Floating-Point FFT on FPGAsabstractAdvances in FPGA technology have led to dramatic improvements in double precision floating-point performance. Modern FPGAs boast several GigaFLOPs of raw computing power. Unfortunately, this computing power is distributed across 30 floating-point units with over 10 cycles of latency each. The user must find two orders of magnitude more parallelism than is typically exploited in a single microprocessor; thus, it is not clear that the computational power of FPGAs can be exploited across a wide range of algorithms. This paper explores three implementation alternatives for the fast Fourier transform (FFT) on FPGAs. The algorithms are compared in terms of sustained performance and memory requirements for various FFT sizes and FPGA sizes. The results indicate that FPGAs are competitive with microprocessors in terms of performance and that the "correct" FFT implementation varies based on the size of the transform and the size of the FPGA. Karl S. Hemmert, Keith D. Underwood |
FCCM | 2 |
| 2005 | A Preliminary Analysis of the MPI Queue Characteristics of Several ApplicationsabstractUnderstanding the message passing behavior and network resource usage of distributed-memory message-passing parallel applications is critical to achieving high performance and scalability. While much research has focused on how applications use critical compute related resources, relatively little attention has been devoted to characterizing the usage of network resources, specifically those needed by the network interface. This paper discusses the importance of understanding network interface resource usage requirements for parallel applications and describes an initial attempt to gather network resource usage data for several real-world codes. The results show widely varying usage patterns between processes in the same parallel job and indicate that resource requirements can change dramatically as process counts increase and input data changes. This suggests that general network resource management strategies may not be widely applicable, and that adaptive strategies or more fine-grained controls may be necessary for environments where network interface resources are severely constrained. Ron Brightwell, Sue Goudy, Keith D. Underwood |
ICPP | 3 |
| 2005 | Considering the Relative Importance of Network Performance and Network FeaturesabstractLatency and bandwidth are usually considered to be the dominant factor in parallel application performance; however, recent studies have indicated that support for independent progress in MPI can also have a significant impact on application performance. This paper leverages the Cplant system at Sandia National Labs to compare a faster, vendor provided MPI library without independent progress to an internally developed MPI library that sacrifices some performance to provide independent progress. The results are surprising. Although some applications see significant negative impacts from the reduced network performance, others are more sensitive to the presence of independent progress. William Lawry, Keith D. Underwood |
ICPP | 2 |
| 2005 | The implications of working set analysis on supercomputing memory hierarchy designabstractSupercomputer architects strive to maximize the performance of scientific applications. Unfortunately, the large, unwieldy nature of most scientific applications has lead to the creation of artificial benchmarks, such as SPEC-FP, for architecture research. Given the impact that these benchmarks have on architecture research, this paper seeks an understanding of how they relate to real-world applications within the Department of Energy. Since the memory system has been found to be a particularly key issue for many applications, the focus of the paper is on the relationship between how the SPEC-FP benchmarks and DOE applications use the memory system. The results indicate that while the SPEC-FP suite is a well balanced suite, supercomputing applications typically demand more from the memory system and must perform more "other work" (in the form of integer computations) along with the floating point operations. The SPEC-FP suite generally demonstrates slightly more temporal locality leading to somewhat lower bandwidth demands. The most striking result is the cumulative difference between the benchmarks and the applications in terms of the requirements to sustain the floating-point operation rate: the DOE applications require significantly more data from main memory (not cache) per FLOP and dramatically more integer instructions per FLOP. Richard C. Murphy, Arun Rodrigues, Peter M. Kogge, Keith D. Underwood |
ICS | 4 |
| 2004 | A comparison of 4X InfiniBand and Quadrics Elan-4 technologiesabstractQuadrics Elan-4 and 4X InfiniBand have comparable performance in terms of peak bandwidth and ping-pong latency. In contrast, the two network architectures differ dramatically in details ranging from signaling technologies to programming interface design to software stacks. Both networks compete in the high performance computing marketplace, and InfiniBand is currently receiving a significant amount of attention, due mostly to its potential cost/performance advantage. This work compares 4X InfiniBand and Quadrics Elan-4 on identical compute hardware using application benchmarks of importance to the DOE community. We use scaling efficiency as the main performance metric, and we also provide a cost analysis for different network configurations. Although our 32-node test platform is relatively small, some scaling issues are evident. In general, the Quadrics hardware scales slightly better on most of the applications tested. Ron Brightwell, Douglas Doerfler, Keith D. Underwood |
CLUSTER | 3 |
| 2004 | Closing the Gap: CPU and FPGA Trends in Sustainable Floating-Point BLAS PerformanceabstractField programmable gate arrays (FPGAs) have long been an attractive alternative to microprocessors for computing tasks - as long as floating-point arithmetic is not required. Fueled by the advance of Moore's law, FPGAs are rapidly reaching sufficient densities to enhance peak floating-point performance as well. The question, however, is how much of this peak performance can be sustained. This paper examines three of the basic linear algebra subroutine (BLAS) functions: vector dot product, matrix-vector multiply, and matrix multiply. A comparison of microprocessors, FPGAs, and reconfigurable computing platforms is performed for each operation. The analysis highlights the amount of memory bandwidth and internal storage needed to sustain peak performance with FPGAs. This analysis considers the historical context of the last six years and is extrapolated for the next six years. Keith D. Underwood, Karl S. Hemmert |
FCCM | 1 |
| 2004 | FPGAs vs. CPUs: trends in peak floating-point performanceabstractMoore's Law states that the number of transistors on a device doubles every two years; however, it is often (mis)quoted based on its impact on CPU performance. This important corollary of Moore's Law states that improved clock frequency plus improved architecture yields a doubling of CPU performance every 18 months. This paper examines the impact of Moore's Law on the peak floating-point performance of FPGAs. Performance trends for individual operations are analyzed as well as the performance trend of a common instruction mix (multiply accumulate). The important result is that peak FPGA floating-point performance is growing significantly faster than peak floating-point performance for a CPU. Keith D. Underwood |
FPGA | 1 |
| 2004 | The Impact of MPI Queue Usage on Message LatencyabstractIt is well known that traditional microbenchmarks do not fully capture the salient architectural features that impact application performance. Even worse, microbenchmarks that target MPI and the communications subsystem do not accurately represent the way that applications use MPI. For example, traditional MPI latency benchmarks time a ping-pong communication with one send and one receive on each of two nodes. The time to post the receive is never counted as part of the latency. This scenario is not even marginally representative of most applications. Two new microbenchmarks are presented here that analyze network latency in a way that more realistically represents the way that MPI is typically used. These benchmarks are used to evaluate modern high-performance networks, including Quadrics, InfiniBand, and Myrinet. Keith D. Underwood, Ron Brightwell |
ICPP | 1 |
| 2004 | An analysis of the impact of MPI overlap and independent progressabstractThe overlap of computation and communication has long been considered to be a significant performance benefit for applications. Similarly, the ability of MPI to make independent progress (that is, to make progress on outstanding communication operations while not in the MPI library) is also believed to yield performance benefits. Using an intelligent network interface to offload the work required to support overlap and independent progress is thought to be an ideal solution, but the benefits of this approach have been poorly studied at the application level. This lack of analysis is complicated by the fact that most MPI implementations do not sufficiently support overlap or independent progress. Recent work has demonstrated a quantifiable advantage for an MPI implementation that uses offload to provide overlap and independent progress. This paper extends this previous work by further qualifying the source of the performance advantage (offload, overlap, or independent progress). Ron Brightwell, Keith D. Underwood |
ICS | 2 |
| 2004 | Characterizing a new class of threads in scientific applications for high end supercomputersabstractChip level multithreading is growing in use throughout the microprocessor world as evidenced in the Intel Pentium 4 and the upcoming innovations in the POWER architecture. These processors typically use a few coarse grain threads that can be difficult for the programmer or compiler to exploit; however, Processing in Memory (PIM) is a technology that has been explored through a long series of supercomputer projects as a facilitator for a different multithreaded execution models. In the multithreading model explored by PIMs, the threads can have radically different characteristics. Specifically, PIMs seek to exploit a large number of very fine grained threads to hide memory access latency and increase parallelism. PIM supports these small threads, or "threadlets", by providing a fast hardware synchronization mechanism, support for harware managment of creation and destruction of threads, and a "shared register" approach which extends the shared memory thread model. This paper discusses some analysis of some very large scientific codes in terms of how they might be mapped onto such a multithreading model with a focus on extremely fine grain threads. Arun Rodrigues, Richard C. Murphy, Peter M. Kogge, Keith D. Underwood |
ICS | 4 |
| 2004 | An Analysis of NIC Resource Usage for Offloading MPIabstractSummary form only given. Modern cluster interconnection networks rely on processing on the network interface to deliver higher bandwidth and lower latency than what could be achieved otherwise. These processors are relatively slow, but they provide adequate capabilities to accelerate some portion of the protocol stack in a cluster computing environment. This offload capability is conceptually appealing, but the standard evaluation of NIC-based protocol implementations relies on simplistic microbenchmarks that create idealized usage scenarios. We evaluate characteristics of MPI usage scenarios using application benchmarks to help define the parameter space that protocol offload implementations should target. Specifically, we analyze characteristics that we expect to have an impact on NIC resource allocation and management strategies, including the length of the MPI posted receive and unexpected message queues, the number of entries in these queues that are examined for a typical operation, and the number of unexpected and expected messages. Ron Brightwell, Keith D. Underwood |
IPDPS | 2 |
| 2003 | A Performance Comparison of Linux and a Lightweight KernelabstractIn this paper, we compare running the Linux operating system on the compute nodes of ASCI Red hardware to running a specialized, highly-optimized lightweight kernel (LWK) operating system. We have ported Linux to the compute and service nodes of the ASCI Red supercomputer, and have run several benchmarks. We present performance and scalability results for Linux compared with the LWK environment. To our knowledge, this is the first direct comparison on identical hardware of Linux and an operating system designed specifically for large-scale supercomputers. In addition to presenting these results, we discuss the limitations of both operating systems, in terms of the empirical evidence as well as other important factors. Ron Brightwell, Rolf Riesen, Keith D. Underwood, Trammell Hudson, Patrick G. Bridges, Arthur B. Maccabe |
CLUSTER | 3 |
| 2003 | Implications of a PIM Architectural Model for MPIabstractMemory may be the only system component that is more commoditized than a microprocessor. To simultaneously exploit this and address the impending memory wall, processing in memory (PIM) research efforts are considering ways to move processing into memory without significantly increasing the cost of the memory. As such, PIM devices may become the basis for future commodity clusters. Although these PIM devices may leverage new computational paradigms such as hardware support for multi-threading and traveling threads, they must provide support for legacy programming models if they are to supplant commodity clusters. This paper presents a prototype implementation of MPI over a traveling thread mechanism called parcels. A performance analysis indicates that the direct hardware support of a traveling thread model can lead to an efficient, lightweight MPI implementation. Arun Rodrigues, Richard C. Murphy, Peter M. Kogge, Jay B. Brockman, Ron Brightwell, Keith D. Underwood |
CLUSTER | 6 |
| 2003 | A Configurable Network Protocol for Cluster Based Communications using Modular Hardware Primitives on an Intelligent NICabstractThe high overhead of generic protocols like TCP/IP provides strong motivation for the development of a better protocol architecture for cluster-based parallel computers. Reconfigurable computing has a unique opportunity to contribute hardware level protocol acceleration while retaining the flexibility to adapt to changing needs. Thus, it is possible to provide application-specific protocol processing to improve performance and to reduce space utilization. Reducing space utilization permits the use of a greater portion of the FPGA for other application-specific processing. This paper focuses on work to create a set of components that can be put together as needed to obtain a customized protocol for each application. The components are parameterizable, increasing the protocol's flexibility. To study the feasibility of such an architecture, hardware components for the reconfigurable logic on the NIC were built such that they can be stitched together as required to provide the required functionality. Feasibility is demonstrated using four different protocol configurations in this paper. The different configurations illustrate trade-offs between chip space and functionality. Ranjesh G. Jaganathan, Keith D. Underwood, Ron Sass |
FCCM | 2 |
| 2003 | A Configurable Network Protocol for Cluster Based Communications using Modular Hardware Primitives on an Intelligent NICabstractThe high overhead of generic protocols like TCP/IP provides strong motivation for the development of a better protocol architecture for cluster-based parallel computers. Reconfigurable computing has a unique opportunity to contribute hardware level protocol acceleration while retaining the flexibility to adapt to changing needs. Specifically, applications on a cluster have various quality of service needs. In addition, these applications typically run for a long time relative to the reconfiguration time of an FPGA. Thus, it is possible to provide application-specific protocol processing to improve performance and reduce space utilization. Reducing space utilization permits the use of a greater portion of the FPGA for other application-specific processing. This paper focuses on work to create a set of parameterizable components that can be put together as needed to obtain a customized protocol for each application. To study the feasibility of such an architecture, hardware components were built that can be stitched together as needed to provide the required functionality. Feasibility is demonstrated using four different protocol configurations, namely: (1) unreliable packet transfer; (2) reliable, unordered message transfer without duplicate elimination; (3) reliable, unordered message transfer with duplicate elimination; and (4) reliable, ordered message transfer with duplicate elimination. The different configurations illustrate trade-offs between chip space and functionality. Ranjesh G. Jaganathan, Keith D. Underwood, Ron Sass |
SC | 2 |
| 2003 | Analysis of a prototype intelligent network interfaceabstractAbstract With a focus on commodity PC systems, Beowulf clusters traditionally lack the cutting edge network architectures, memory subsystems, and processor technologies found in their more expensive supercomputer counterparts. Many users find that what Beowulf clusters lack in technology, they more than make up for with their significant cost advantage. In this paper, an architectural extension that adds reconfigurable computing to the network interface of Beowulf clusters is proposed. The proposed extension, called an intelligent network interface card (or INIC), enhances both the network and processor capabilities of the cluster, which has a significant impact on the performance of a crucial class of applications. Furthermore, for some applications, the proposed extension partially compensates for weaknesses in the PC memory subsystem. A prototype of the proposed architecture was constructed and analyzed. In addition, two applications, the 2D Fast Fourier Transform (2D‐FFT) and integer sorting, which benefit from the resulting architecture, are discussed and analyzed on the proposed architecture. Specifically, results indicate that the 2D‐FFT is performed 15–50% faster when using the prototype INIC rather than the comparable Gigabit Ethernet. Early simulation results also indicate that integer sorting will receive a 21–32% performance boost from the prototype. The significant improvements seen with a relatively limited prototype lead to the conclusion that cluster network interfaces enhanced with reconfigurable computing could significantly improve the Beowulf architecture. Copyright © 2003 John Wiley & Sons, Ltd. Keith D. Underwood, Walter B. Ligon III, Ron Sass |
Concurr. Comput. Pract. Exp. | 1 |
| 2002 | GRIP: A Reconfigurable Architecture for Host-Based Gigabit-Rate Packet ProcessingabstractOne of the fundamental challenges for modern high-performance network interfaces is the processing capabilities required to process packets at high speeds. Simply transmitting or receiving data at gigabit speeds fully utilizes the CPU on a standard workstation. Any processing that must be done to the data, whether at the application layer or the network layer, decreases the achievable throughput. This paper presents an architecture for offloading a significant portion of the network, processing from the host CPU onto the network interface. A prototype, called the GRIP (Gigabit Rate IPSec) card, has been constructed based on an FPGA coupled with a commodity Gigabit Ethernet MAC. Experimental results based on the prototype are presented and analyzed. In addition, a second generation design is presented in the context of lessons learned from the prototype. Peter Bellows, Jaroslav Flidr, Tom Lehman, Brian Schott, Keith D. Underwood |
FCCM | 5 |
| 2001 | A Reconfigurable Extension to the Network Interface of Beowulf ClustersabstractWith a focus on commodity PC systems, Beowulf clusters traditionally lack the cutting edge network architectures, memory subsystems, and processor technologies found in their more expensive supercomputer counterparts. What Beowulf clusters lack in technology, they more than make up for with their significant cost advantage over traditional supercomputers. We propose an architectural extension that adds reconfigurable computing to the network interface of Beowulf clusters. This enhances both the network and processor capabilities of the cluster. Furthermore, for some applications, the proposed extension partially compensates for weaknesses in the PC memory subsystem. We discuss two applications, the 2D Fast Fourier Transform (FFT) and integer sorting, which benefit from the resulting architecture. Keith D. Underwood, Ron Sass, Walter B. Ligon III |
CLUSTER | 1 |
| 2001 | Acceleration of a 2D-FFT on an Adaptable Computing Cluster
Keith D. Underwood, Ron Sass, Walter B. Ligon III |
FCCM | 1 |
| 2001 | Cost effectiveness of an adaptable computing clusterabstractWith a focus on commodity PC systems, Beowulf clusters traditionally lack the cutting edge network architectures, memory subsystems, and processor technologies found in their more expensive supercomputer counterparts. What Beowulf clusters lack in technology, they more than make up for with their significant cost advantage over traditional supercomputers. This paper presents the cost implications of an architectural extension that adds reconfigurable computing to the network interface of Beowulf clusters. A quantitative idea of cost-effectiveness is formulated to evaluate computing technologies. Here, cost-effectiveness is considered in the context of two applications: the 2D Fast Fourier Transform (2D-FFT) and integer sorting. Keith D. Underwood, Ron Sass, Walter B. Ligon III |
SC | 1 |
| 1998 | A Re-evaluation of the Practicality of Floating-Point Operations on FPGAsabstractThe use of reconfigurable hardware to perform high precision operations such as IEEE floating point operations has been limited in the past by FPGA resources. We discuss the implementation of IEEE single precision floating-point multiplication and addition. Then, we assess the practical implications of using these operations in the Xilinx 4000 series FPGAs considering densities available now and scheduled for the near future. For each operation, we present space requirements and performance information. This is followed by a discussion of an algorithm, matrix multiplication, based on these operations, which achieves performance comparable to conventional microprocessors. Algorithm implementation options and their performance implications are discussed and corresponding measured results are given. Walter B. Ligon III, Scott McMillan, Greg Monn, Kevin Schoonover, Fred Stivers, Keith D. Underwood |
FCCM | 6 |
| 1998 | Implementation of IEEE Single-Precision Floating-Point Operations on FPGAs (Abstract)
Walter B. Ligon III, Greg Monn, S. P. McMillan, Kevin Schoonover, Fred Stivers, Keith D. Underwood |
FPGA | 6 |