VLDB 2026 Research / reviewers in the wild / expert
Michael J. Schulte
dblp:s/MichaelJSchulte · also Mike Schulte 0001
· DBLP profile ↗
91ranked-venue papers
16as first author
2since 2021 · last 2024
0000-0003-1305-406XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 72 · 12 first-author · 2 since 2021Theory of computation · 14 · 4 first-authorSoftware engineering, systems software and programming languages · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Realizing the AMD Exascale Heterogeneous Processor Vision : Industry ProductabstractAMD had previously detailed its exascale research journey from initial targets and requirements to the development and evolution of its vision of a high-performance computing (HPC) accelerated processing unit (APU), dubbed the Exascale Heterogeneous Processor or EHP. At the conclusion of that work, the learnings were integrated into the design of the node architecture that went into the Frontier supercomputer, the world’s first exascale machine. However, while the Frontier node architecture embodied many of the attributes of the EHP concept, advanced heterogeneous integration capabilities at the time were not yet sufficiently mature to realize our vision of a fully-integrated APU for HPC and AI. In this paper, we finish the EHP’s story by digging deeper into why an APU was not the right solution at the time of our first exascale architecture, what the shortcomings were of previous EHP concepts, and how AMD further evolved the concept into the AMD Instinct™ MI300A APU. MI300A is the culmination of years of AMD developments in advanced packaging technologies, its APU hardware and software, and the next step in our highly effective chiplet strategy to not only deliver a groundbreaking design for exascale computing, but to also meet the demands of new large-language model and generative AI applications. Alan Smith 0003, Gabriel H. Loh, Michael J. Schulte, Mike Ignatowski, Samuel Naffziger, Mike Mantor, Nathan Kalyanasundharam, Vamsi Alla, Nicholas Malaya, Joseph L. Greathouse, Eric Chapman, Raja Swaminathan |
ISCA | 3 |
| 2023 | A Research Retrospective on AMD's Exascale Computing JourneyabstractThe pace of advancement of the top-end supercomputers historically followed an exponential curve similar to (and driven in part by) Moore's Law. Shortly after hitting the petaflop mark, the community started looking ahead to the next milestone: Exascale. However, many obstacles were already looming on the horizon, such as the slowing of Moore's Law, and others like the end of Dennard Scaling had already arrived. Anticipating significant challenges for the overall high-performance computing (HPC) community to achieve the next 1000x improvement, the U.S. Department of Energy (DOE) launched the Exascale Computing Program to enable and accelerate fundamental research across the many technologies needed to achieve exascale computing. Gabriel H. Loh, Michael J. Schulte, Mike Ignatowski, Vignesh Adhinarayanan, Shaizeen Aga, Derrick Aguren, Varun Agrawal, Ashwin M. Aji, Johnathan Alsop, Paul T. Bauman, Bradford M. Beckmann, Majed Valad Beigi, Sergey Blagodurov, Travis Boraten, Michael Boyer, William C. Brantley, Noel Chalmers, Shaoming Chen, Michael L. Chu, David Cownie, Nicholas Curtis, Joris Del Pino, Nam Duong, Alexandru Dutu, Yasuko Eckert, Christopher Erb, Chip Freitag, Joseph L. Greathouse, Sudhanva Gurumurthi, Anthony Gutierrez, Khaled Hamidouche, Sachin Hossamani, Wei Huang 0004, Mahzabeen Islam, Nuwan Jayasena, John Kalamatianos, Onur Kayiran, Jagadish Kotra, Alan Lee, Daniel Lowell, Niti Madan, Abhinandan Majumdar, Nicholas Malaya, Srilatha Manne, Susumu Mashimo, Damon McDougall, Elliot Mednick, Michael Mishkin, Mark Nutter, Indrani Paul, Matthew Poremba, Brandon Potter, Kishore Punniyamurthy, Sooraj Puthoor, Steven E. Raasch, Karthik Rao, Gregory Rodgers, Marko Scrbak, Mohammad Seyedzadeh, John Slice, Vilas Sridharan, René van Oostrum, Eric Van Tassell, Abhinav Vishnu, Samuel Wasmundt, Mark Wilkening, Noah Wolfe, Mark Wyse, Adithya Yalavarti, Dmitri Yudanov |
ISCA | 2 |
| 2020 | Approximate Computing: From Circuits to Applications [Scanning the Issue]abstractThis special issue explores the technological contributions and developments of approximate computing at disparate levels and provides insight into exciting directions for the future. Weiqiang Liu 0001, Fabrizio Lombardi, Michael J. Schulte |
Proc. IEEE | 3 |
| 2017 | Accelerating Matrix Processing with GPUsabstractMatrix operations are common and expensive computations in a variety of applications. They occur frequently in high-performance computing, graphics, graph processing, and machine learning applications. This paper discusses how to map a variety of important matrix computations, including sparse matrix-vector multiplication (SpMV), sparse triangle solve (SpTS), graph processing, and dense matrix-matrix multiplication, to GPUs. Since many emerging systems will use heterogeneous architectures (e.g. CPUs and GPUs) to attain the desired performance targets under strict power constraints, this paper discusses implications and future research for matrix processing with heterogeneous designs. Conclusions common to the matrix operations discussed in this paper are: (1) Future algorithms should be written to ensure that the essential computations fit into local memory, which may require direct programmer management. (2) Algorithms are needed that expose high levels of parallelism. (3) While the scale of computation is often sufficient to support algorithms with superior asymptotic order, additional considerations, such as memory capacity and bandwidth, must also be carefully managed. (4) Libraries should be used to provide portable performance. Nicholas Malaya, Shuai Che, Joseph L. Greathouse, René van Oostrum, Michael J. Schulte |
ARITH | 5 |
| 2017 | Design and Analysis of an APU for Exascale ComputingabstractThe challenges to push computing to exaflop levels are difficult given desired targets for memory capacity, memory bandwidth, power efficiency, reliability, and cost. This paper presents a vision for an architecture that can be used to construct exascale systems. We describe a conceptual Exascale Node Architecture (ENA), which is the computational building block for an exascale supercomputer. The ENA consists of an Exascale Heterogeneous Processor (EHP) coupled with an advanced memory system. The EHP provides a high-performance accelerated processing unit (CPU+GPU), in-package high-bandwidth 3D memory, and aggressive use of die-stacking and chiplet technologies to meet the requirements for exascale computing in a balanced manner. We present initial experimental analysis to demonstrate the promise of our approach, and we discuss remaining open research challenges for the community. Thiruvengadam Vijayaraghavan, Yasuko Eckert, Gabriel H. Loh, Michael J. Schulte, Mike Ignatowski, Bradford M. Beckmann, William C. Brantley, Joseph L. Greathouse, Wei Huang 0004, Arun Karunanithi, Onur Kayiran, Mitesh R. Meswani, Indrani Paul, Matthew Poremba, Steven E. Raasch, Steven K. Reinhardt, Greg Sadowski, Vilas Sridharan |
HPCA | 4 |
| 2014 | Memory scheduling towards high-throughput cooperative heterogeneous computingabstractTechnology scaling enables the integration of both the CPU and the GPU into a single chip for higher throughput and energy efficiency. In such a single-chip heterogeneous processor (SCHP), its memory bandwidth is the most critically shared resource, requiring judicious management to maximize the throughput. Previous studies on memory scheduling for SCHPs have focused on the scenario where multiple applications are running on the CPU and the GPU respectively, which we denote as a multi-tasking scenario. However, another increasingly important usage scenario for SCHPs is cooperative heterogeneous computing, where a single parallel application is partitioned between the CPU and the GPU such that the overall throughput is maximized. Hao Wang 0011, Ripudaman Singh, Michael J. Schulte, Nam Sung Kim |
PACT | 3 |
| 2014 | Process variation-aware workload partitioning algorithms for GPUs supporting spatial-multitaskingabstractHigh-level programming languages have transformed graphics processing units (GPUs) from domain-restricted devices into powerful compute platforms. Yet many “generalpurpose GPU” (GPGPU) applications fail to fully utilize the GPU resources. Executing multiple applications simultaneously on different regions of the GPU (spatial multitasking) thus improves system performance. However, within-die process variations lead to significantly different maximum operating frequencies (Fmax) of the streaming multiprocessors (SMs) within a GPU. As the chip size and number of SMs per chip increase, the frequency variation is also expected to increase, exacerbating the problem. The increased number of SMs also provides a unique opportunity: we can allocate resources to concurrently-executing applications based on how those applications are affected by the different available Fmaxvalues. In this paper, we study the effects of per-SM clocking on spatial multitasking-capable GPUs. We demonstrate two factors that affect the performance of simultaneously-running applications: (i) the SM partitioning algorithm that decides how many resources to assign to each application, and (ii) the assignment of SMs to applications based on the operating frequencies of those SMs and the applications characteristics. Our experimental results show that spatial multitasking that partitions SMs based on application characteristics, when combined with per-SM clocking, can greatly improve application performance by up to 46% on average compared to cooperative multitasking with global clocking. Paula Aguilera, Jungseob Lee, Amin Farmahini Farahani, Katherine Morrow, Michael J. Schulte, Nam Sung Kim |
DATE | 5 |
| 2014 | Energy-Efficient Pixel-ArithmeticabstractWith the advent of pervasive computing, the performance requirements of visual applications have increased significantly. On the other hand, the energy budget for future devices may decrease due to reduced form factors and thus smaller battery sizes. It is thus imperative to improve energy efficiency of visual applications to meet their stringent demands in energy-constrained devices. This paper presents a novel energy-efficient floating-point unit (E2FPU) that dynamically detects multiplications in which at least one operand is an integer power of two or the sum of consecutive integer powers of two (POW2 operands) and executes them on the E2FPU. We show that POW2 operands are extremely common in visual applications. For non-POW2 operands, we propose dynamic approximation of suitable operands as the closest POW2 operand with the same exponent while ensuring a small and limited approximation error. Finally, we exploit the limited dynamic range of pixel values to enhance the E2FPU to support low-energy FP addition for a sub-set of the single-precision floating-point dynamic range. The proposed E2FPU can reduce execution energy by 35 percent with a negligible impact on the peak signal to noise ratio. Overall, at the chip-level, our approaches yield a 12 percent dynamic energy reduction for a graphics processing unit (GPU). Syed Zohaib Gilani, Nam Sung Kim, Michael J. Schulte |
IEEE Trans. Computers | 3 |
| 2014 | Low-Cost Per-Core Voltage Domain Support for Power-Constrained High-Performance ProcessorsabstractPer-core voltage domains can improve performance under a power constraint. Most commercial processors, however, only have a single voltage domain for all processor cores. This is because splitting the single voltage domain into per-core voltage domains and powering them with multiple off-chip voltage regulators (VRs) incur a high cost for the platform and package designs. Although using on-chip switching VRs can be an alternative solution, integrating high-quality inductors for VRs with cores has been a technical challenge. In this paper, we propose a cost-effective power delivery technique to support per-core voltage domains. Our technique is based on the observations that: 1) core-to-core (C2C) voltage variations are relatively small for most execution intervals when the voltages/frequencies are optimized to maximize performance under a power constraint and 2) per-core power-gating devices augmented with feedback control circuitry can serve as low-cost VRs that can provide high efficiency in situations like 1). Our experimental results show that processors using our technique can achieve power efficiency as high as those using the per-core on-chip switching VRs at a much lower cost. Abhishek A. Sinkar, Hamid Reza Ghasemi, Michael J. Schulte, Ulya R. Karpuzcu, Nam Sung Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | Reevaluating the latency claims of 3D stacked memoriesabstractIn recent years, 3D technology has been a popular area of study that has allowed researchers to explore a number of novel computer architectures. One of the more popular topics is that of integrating 3D main memory dies below the computing die and connecting them with through-silicon vias (TSVs). This is assumed to reduce off-chip main memory access latencies by roughly 45% to 60%. Our detailed circuit-level models, however, demonstrate that this latency reduction from the TSVs is significantly less. In this paper, we present these models, compare 2D and 3D main memory latencies, and show that the reduction in latency from using 3D main memory to be no more than 2.4 ns. We also show that although the wider I/O bus width enabled by using TSVs increases performance, it may do so with an increase in power consumption. Although TSVs consume less power per bit transfer than off-chip metal interconnects (11.2 times less power per bit transfer), TSVs typically use considerably more bits and may result in a net increase in power due to the large number of bits in the memory I/O bus. Our analysis shows that although a 3D memory hierarchy exploiting a wider memory bus can increase performance, this performance increase may not justify the net increase in power consumption. Daniel W. Chang, Gyungsu Byun, Hoyoung Kim, Minwook Ahn, Soojung Ryu, Nam Sung Kim, Michael J. Schulte |
ASP-DAC | 7 |
| 2013 | Power-efficient computing for compute-intensive GPGPU applicationsabstractThe peak compute performance of GPUs has been increased by integrating more compute resources and operating them at higher frequency. However, such approaches significantly increase power consumption of GPUs, limiting further performance increase due to the power constraint. Facing such a challenge, we propose three techniques to improve power efficiency and performance of GPUs in this paper. First, we observe that many GPGPU applications are integer-intensive. For such applications, we combine a pair of dependent integer instructions into a composite instruction that can be executed by an enhanced fused multiply-add unit. Second, we observe that computations for many instructions are duplicated across multiple threads. We dynamically detect such instructions and execute them in a separate scalar unit. Finally, we observe that 16 or fewer bits are sufficient for accurate representation of operands and results of many instructions. Thus, we split the 32-bit datapath into two 16-bit datapath slices that can concurrently issue and execute up to two such instructions per cycle. All three proposed techniques can considerably increase utilization of compute resources, improving power efficiency and performance by 20% and 15%, respectively. Syed Zohaib Gilani, Nam Sung Kim, Michael J. Schulte |
HPCA | 3 |
| 2013 | Dynamic bandwidth scaling for embedded DSPs with 3D-stacked DRAM and wide I/Osabstract3D main memory is an emerging technology that stacks DRAM dies underneath the processor die using through-silicon vias (TSVs). Prior studies assumed that such technology would decrease main memory access latency by 45% to 60%, while also allowing designers to increase main memory bandwidth. Although the latter is true, it was recently shown that the latency savings of 3D main memory is only 6.3%. In this paper, we first analyze memory latency reduction opportunities in a 3D main memory system with Wide I/O by taking better advantage of 3D integration technology and quantify their benefit. Specifically, redesigning the DRAM to memory controller synchronizers and placing the address, command, and data pads closer to the DRAM banks can decrease 3D main memory latency by 24.7%. We show that current 3D DRAM with Wide I/O can increase the geometric mean performance of an embedded processor that is similar to a Texas instrument C67x DSP by 9.7% (and up to 23.3%). Second, we observe that 3D DRAM with Wide IO can increase average system energy consumption of energy-constrained embedded DSPs by 2.6% (and up to 8.9%). To improve I/O energy efficiency, we propose to dynamically scale memory bandwidth (i.e. the I/O width) at runtime based on an application's program phases. Our dynamic bandwidth scaling algorithms increase average performance by 6.6% while increasing average energy consumption by only 0.5%. Daniel W. Chang, Young Hoon Son, Jung Ho Ahn, Hoyoung Kim, Minwook Ahn, Michael J. Schulte, Nam Sung Kim |
ICCAD | 6 |
| 2013 | Performance boosting under reliability and power constraintsabstractVoltage droops resulting from inductive noise are common in state-of-the-art processors. Many of the techniques used to reduce energy consumption - clock gating, power gating, process shrinks, and voltage reduction - lead to increased voltage droops or increased sensitivity to voltage variations. Designers use voltage guardbands to minimize errors due to voltage fluctuations and inductive noise; however, this leads to lower performance because the voltage and frequency points are set to deal with voltage droops from a worst-case benchmark or stressmark. Although most applications do not approach the voltage droop caused by the stressmark, there is no mechanism to guarantee correct operation outside the tested range. In this paper, we examine floating-point issue throttling (FP throttling), a hardware technique that reduces worst-case voltage droop. By lowering the issue rate in the FP scheduler, the processor can significantly reduce the maximum voltage droop in the system. We show the impact of FP throttling on voltage droop, and translate this reduction in voltage droop to an increase in operating frequency (and hence increased performance) because an additional guardband is no longer required to guard against droops resulting from heavy FP usage. We then examine the impact of FP throttling and guardband reduction on the SPEC CPU2006 benchmarks and show that some benchmarks benefit from the frequency improvements with FP throttling while others suffer due to reduced FP throughput. Finally, we present two techniques to determine dynamically when to trade FP throughput for reduced voltage margin and increased frequency, and show performance improvements of up to 15% for CINT2006 benchmarks and up to 8% for CFP2006 benchmarks. Our studies are done on hardware in which FP units generate the worst-case voltage droop. The technique can be modified for architectures in which other units cause the worst droop. Youngtaek Kim, Lizy Kurian John, Indrani Paul, Srilatha Manne, Michael J. Schulte |
ICCAD | 5 |
| 2013 | REEL: Reducing effective execution latency of floating point operationsabstractThe height of the dynamic dependence graph of a program, as executed by a processor, determines the minimum bound on the execution time. This height can be decreased by reducing the effective execution latency of operations that form dependence chains in the graph. In this paper, we propose a technique called REEL to reduce overall latency of chains of dependent floating point (FP) operations by increasing the throughput of computation. REEL comprises of a high-throughput floating point unit (HFP) that allows early issue of an FP Add that is dependent on another FP Add or FP Multiply. This is complemented by instruction scheduler modifications that allow early issue of dependent FP Adds, and a novel checker logic that corrects any precision errors. Unlike conventional static operation fusion, like fused Multiply-Add (FMA), there are no changes to the instruction set to enable utilization of the new hardware, and no recompilation is necessary. Furthermore, unlike ISA-level FMA, our technique produces results that are bit compatible while boosting performance of Add-Add dependence pairs in addition to Multiply-Add pairs. Our evaluation of REEL using CFP2006 benchmarks shows an average performance gain of 7.6% and maximum performance gain of 17% while consuming 1.2% lower energy. Vignyan Reddy Kothinti Naresh, Syed Zohaib Gilani, Erika Gunadi, Nam Sung Kim, Michael J. Schulte, Mikko H. Lipasti |
ISLPED | 5 |
| 2013 | Exploiting GPU peak-power and performance tradeoffs through reduced effective pipeline latencyabstractModern GPUs share limited hardware resources, such as register files, among a large number of concurrently executing threads. For efficient resource sharing, several buffering and collision avoidance stages are inserted in the GPU pipeline. These additional stages increase the read-after-write (RAW) latencies of instructions. Since GPUs are often architected to hide RAW latencies through extensive multithreading, they typically do not employ power-hungry data-forwarding networks (DFNs). However, we observe that many GPGPU applications do not have enough active threads that are ready to issue instructions to hide these RAW latencies. In this paper, we first demonstrate that DFNs can considerably improve the performance of many compute-intensive GPGPU applications and then propose most recent result forwarding (MoRF) as a low-power alternative to the DFN. Second, for floating-point (FP) operations, we exploit a high-throughput fused multiply-add (HFMA) unit to further reduce both RAW latencies and the number of FMA units in the GPU without impacting instruction throughput. MoRF and HFMA together provide a geometric mean performance improvement of 18% and 29% for integer/single-precision and double-precision GPGPU applications, respectively. Finally, both MoRF and HFMA allow the GPU to effectively mimic a shallower pipeline for a large percentage of instructions. Exploiting such a benefit, we propose low-power pipelines that can reduce peak power consumption by 14% without affecting the performance or increasing the complexity of the forwarding network. The peak power reduction allows GPUs to operate more cores within the same power budget, achieving a geometric mean performance improvement of 33% for double-precision GPGPU applications. Syed Zohaib Gilani, Nam Sung Kim, Michael J. Schulte |
MICRO | 3 |
| 2013 | Modular Design of High-Throughput, Low-Latency Sorting UnitsabstractHigh-throughput and low-latency sorting is a key requirement in many applications that deal with large amounts of data. This paper presents efficient techniques for designing high-throughput, low-latency sorting units. Our sorting architectures utilize modular design techniques that hierarchically construct large sorting units from smaller building blocks. The sorting units are optimized for situations in which only the M largest numbers from N inputs are needed, because this situation commonly occurs in many applications for scientific computing, data mining, network processing, digital signal processing, and high-energy physics. We utilize our proposed techniques to design parameterized, pipelined, and modular sorting units. A detailed analysis of these sorting units indicates that as the number of inputs increases their resource requirements scale linearly, their latencies scale logarithmically, and their frequencies remain almost constant. When synthesized to a 65-nm TSMC technology, a pipelined 256-to-4 sorting unit with 19 stages can perform more than 2.7 billion sorts per second with a latency of about 7 ns per sort. We also propose iterative sorting techniques, in which a small sorting unit is used several times to find the largest values. Amin Farmahini Farahani, Henry Duwe, Michael J. Schulte, Katherine Compton |
IEEE Trans. Computers | 3 |
| 2013 | Binary Integer Decimal-Based Floating-Point MultiplicationabstractThis paper presents a multiplier that operates on binary integer decimal (BID) encoded decimal floating-point (DFP) numbers. It uses a single binary multiplier with carry-save feedback for both significand multiplication and rounding, and it is compliant with the IEEE 754-2008 Standard. Optimizations decrease the BID multiplier's area and critical path delay. Sonia Gonzalez-Navarro, Charles Tsen, Michael J. Schulte |
IEEE Trans. Computers | 3 |
| 2012 | Power-efficient computing for compute-intensive GPGPU applicationsabstractThe peak performance of graphics processing units (GPUs) has traditionally been increased by increasing the number of compute resources and/or their frequency. However, these approaches significantly increase the power consumption of GPUs. Consequently, modern high-performance GPUs are power constrained and must employ more power efficient approaches for performance improvements in future processors. In this paper we propose three power-efficient techniques for improving the performance of GPUs. First, we observe that many GPGPU applications are integer instruction intensive. For such applications, we propose to utilize the fused multiply-add (FMA) units to fuse dependent integer instructions into a composite instruction, improving power efficiency and performance by reducing the number of fetched/executed instructions. Secondly, GPUs often perform computations that are duplicated across multiple threads. We dynamically detect such instructions and execute them in a separate scalar pipeline. Finally, the register file bandwidth in GPUs is a critical resource that is optimized for 32-bit instruction operands. However, many operands require considerably fewer bits for accurate representation and computations. We propose a sliced GPU architecture that improves performance of the GPU by dual-issuing instructions to two 16-bit execution slices. Overall, our techniques result in more than a 25% (geometric mean) power efficiency improvement. Syed Zohaib Gilani, Nam Sung Kim, Michael J. Schulte |
PACT | 3 |
| 2012 | Lossless and lossy memory I/O link compression for improving performance of GPGPU workloadsabstractState-of-the-art graphic processing units (GPUs) provide very high memory bandwidth, but the performance of many general-purpose GPU (GPGPU) workloads is still bounded by memory bandwidth. Although compression techniques have been adopted by commercial GPUs, they are only used for compressing texture and color data, not data for GPGPU workloads. Furthermore, the microarchitectural details of GPU compression are proprietary and its performance benefits have not been previously published. In this paper, we first investigate required microarchitectural changes to support lossless compression techniques for data transferred between the GPU and its off-chip memory to provide higher effective bandwidth. Second, by exploiting some characteristics of floating-point numbers in many GPGPU workloads, we propose to apply lossless compression to floating-point numbers after truncating their least-significant bits (i.e., lossy compression). This can reduce the bandwidth usage even further with very little impact on overall computational accuracy. Finally, we demonstrate that a GPU with our lossless and lossy compression techniques can improve the performance of memory-bound GPGPU workloads by 26% and 41% on average. Vijay Sathish, Michael J. Schulte, Nam Sung Kim |
PACT | 2 |
| 2012 | Workload and power budget partitioning for single-chip heterogeneous processorsabstractWith technology scaling, manufacturers are integrating both CPU and GPU cores in a single chip to improve the throughput of emerging applications. To maximize the throughput of a single-chip heterogeneous processor (SCHP), the chip power budget shared between the CPU and GPU must be effectively utilized. At the same time, the CPU and GPU in an SCHP must each satisfy its own power constraint. Furthermore, the power budget allocated to the CPU and GPU impacts performance. In this paper, using a detailed cycle-level SCHP simulator, we first demonstrate that the joint optimization of workload and power budget partitioning between the CPU and GPU can provide 13% higher throughput than the optimization of workload partitioning alone under a fixed power budget allocation to the CPU and GPU. Second, we propose an effective runtime algorithm that can determine near-optimal or optimal combinations of workload and power budget partitioning. The algorithm exploits the runtime power efficiencies of the workload executed on the CPU and the GPU. Using the detailed cycle-level SCHP simulator, we show that within five to eight kernel invocations the algorithm can achieve 96% of the maximum throughput obtained by an exhaustive search algorithm. Finally, we demonstrate comparable throughput improvements when we apply the algorithm to a commercial computing system with an SCHP. Hao Wang 0011, Vijay Sathish, Ripudaman Singh, Michael J. Schulte, Nam Sung Kim |
PACT | 4 |
| 2012 | Virtual Floating-Point Units for Low-Power Embedded ProcessorsabstractFloating-point (FP) arithmetic is becoming increasingly common in many embedded applications. Typically these applications execute in battery-powered, energy-constrained environments. Due to their tight area and power constraints, however, embedded processors often do not incorporate dedicated FP hardware. Instead, they only support fixed-point (FxP) arithmetic at the expense of considerably increased programming complexity and longer runtimes. In this paper, we propose low-overhead approaches to support FP arithmetic (addition, subtraction, multiplication, fused multiply-add) without incurring the high area and power penalties of dedicated FP hardware. Our approaches utilize the existing FxP execution resources in processors plus a small amount of additional hardware to support FP operations. Compared to a baseline processor with dedicated FP hardware, a processor with our approaches can reduce the area and power consumption by 24% and 31%, respectively. We also demonstrate that a processor using our approaches improves energy efficiency and performance by nearly 30%. Syed Zohaib Gilani, Nam Sung Kim, Michael J. Schulte |
ASAP | 3 |
| 2012 | A Linear Algebra Core Design for Efficient Level-3 BLASabstractReducing power consumption and increasing efficiency is a key concern for many applications. It is well-accepted that specialization and heterogeneity are crucial strategies to improve both power and performance. Yet, how to design highly efficient processing elements while maintaining enough flexibility within a domain of applications is a fundamental question. In this paper, we present the design of a specialized Linear Algebra Core (LAC) for an important class of computational kernels, the level-3 Basic Linear Algebra Subprograms (BLAS). We demonstrate a detailed algorithm/architecture co-design for mapping a number of level-3 BLAS operations onto the LAC. Results show that our prototype LAC achieves a performance of around 64 GFLOPS (double precision) for these operations, while consuming less than 1.3 Watts in standard 45nm CMOS technology. This is on par with a full-custom design and up to 50× and 10× better in terms of power efficiency than CPUs and GPUs. Ardavan Pedram, Syed Zohaib Gilani, Nam Sung Kim, Robert A. van de Geijn, Michael J. Schulte, Andreas Gerstlauer |
ASAP | 5 |
| 2012 | Cost-effective power delivery to support per-core voltage domains for power-constrained processorsabstractPer-core voltage domains can improve performance under a power constraint. Most commercial processors, however, only have one chip-wide voltage domain because splitting the voltage domain into per-core voltage domains and powering them with multiple off-chip voltage regulators (VRs) incurs a high cost for the platform and package designs. Although using on-chip switching VRs can be an alternative solution, integrating high-quality inductors and cores on the same chip has been a technical challenge. In this paper, we propose a cost-effective power delivery technique to support per-core voltage domains. Our technique is based on the observations that (i) core-to-core voltage variations are relatively small for most execution intervals when the voltages/frequencies are optimized to maximize performance under a power constraint and (ii) per-core power-gating devices augmented with small circuits can serve as low-cost VRs that can provide high efficiency in situations like (i). Our experimental results show that processors using our technique can achieve power efficiency as high as those using per-core on-chip switching VRs at much lower cost. Hamid Reza Ghasemi, Abhishek A. Sinkar, Michael J. Schulte, Nam Sung Kim |
DAC | 3 |
| 2012 | The case for GPGPU spatial multitaskingabstractThe set-top and portable device market continues to grow, as does the demand for more performance under increasing cost, power, and thermal constraints. The integration of Graphics Processing Units (GPUs) into these devices and the emergence of general-purpose computations on graphics hardware enable a new set of highly parallel applications. In this paper, we propose and make the case for a GPU multitasking technique called spatial multitasking. Traditional GPU multitasking techniques, such as cooperative and preemptive multitasking, partition GPU time among applications, while spatial multitasking allows GPU resources to be partitioned among multiple applications simultaneously. We demonstrate the potential benefits of spatial multitasking with an analysis and characterization of General-Purpose GPU (GPGPU) applications. We find that many GPGPU applications fail to utilize available GPU resources fully, which suggests the potential for significant performance benefits using spatial multitasking instead of, or in combination with, preemptive or cooperative multitasking. We then implement spatial multitasking and compare it to cooperative multitasking using simulation. We evaluate several heuristics for partitioning GPU stream multiprocessors (SMs) among applications and find spatial multitasking shows an average speedup of up to 1.19 over cooperative multitasking when two applications are sharing the GPU. Speedups are even higher when more than two applications are sharing the GPU. Jacob Adriaens, Katherine Compton, Nam Sung Kim, Michael J. Schulte |
HPCA | 4 |
| 2012 | Something old and something new: P-states can borrow microarchitecture techniques tooabstractThe limited utility of voltage scaling in nano-scale technologies has led high-performance processors to rely increasingly on frequency scaling for power management. However, frequency scaling provides only a linear dynamic power reduction. Yasuko Eckert, Srilatha Manne, Michael J. Schulte, David A. Wood 0001 |
ISLPED | 3 |
| 2012 | AUDIT: Stress Testing the Automatic WayabstractSudden variations in current (large di/dt) can lead to significant power supply voltage droops and timing errors in modern microprocessors. Several papers discuss the complexity involved with developing test programs, also known as stress marks, to stress the system. Authors of these papers produced tools and methodologies to generate stress marks automatically using techniques such as integer linear programming or genetic algorithms. However, nearly all of the previous work took place in the context of single-core systems, and results were collected and analyzed using cycle-level simulators. In this paper, we measure and analyze di/dt issues on state-of-the-art multi-core x86 systems using real hardware rather than simulators. We build on an existing single-core stress mark generation tool to develop an Automated DI/dT stress mark generation framework, referred to as AUDIT, to generate di/dt stress marks quickly and effectively for multicore systems. We showcase AUDIT's capabilities to adjust to micro architectural and architectural changes. We also present a dithering algorithm to address thread alignment issues on multi-core processors. We compare standard benchmarks, existing di/dt stress marks, and AUDIT-generated stress marks executing on multi-threaded, multi-core systems with complex out-of-order pipelines. Finally, we show how stress analysis using simulators may lead to flawed insights about di/dt issues. Youngtaek Kim, Lizy Kurian John, Sanjay Pant, Srilatha Manne, Michael J. Schulte, William Lloyd Bircher, Madhu Saravana Sibi Govindan |
MICRO | 5 |
| 2012 | A study of decimal left shifters for binary numbers
Sonia Gonzalez-Navarro, Javier Hormigo, Michael J. Schulte |
Inf. Comput. | 3 |
| 2011 | Improving Throughput of Power-Constrained GPUs Using Dynamic Voltage/Frequency and Core ScalingabstractState-of-the-art graphic processing units (GPUs) can offer very high computational throughput for highly parallel applications using hundreds of integrated cores. In general, the peak throughput of a GPU is proportional to the product of the number of cores and their frequency. However, the product is often limited by a power constraint. Although the throughput can be increased with more cores for some applications, it cannot for others because parallelism of applications and/or bandwidth of on-chip interconnects/caches and off-chip memory are limited. In this paper, first, we demonstrate that adjusting the number of operating cores and the voltage/frequency of cores and/or on-chip interconnects/caches for different applications can improve the throughput of GPUs under a power constraint. Second, we show that dynamically scaling the number of operating cores and the voltages/frequencies of both cores and on-chip interconnects/caches at runtime can improve the throughput of application even further. Our experimental results show that a GPU adopting our runtime dynamic voltage/frequency and core scaling technique can provide up to 38% (and nearly 20% on average) higher throughput than the baseline GPU under the same power constraint. Jungseob Lee, Vijay Sathish, Michael J. Schulte, Katherine Compton, Nam Sung Kim |
PACT | 3 |
| 2011 | A decimal floating-point fused multiply-add unit with a novel decimal leading-zero anticipatorabstractThere is a significant demand for decimal arithmetic, especially in commercial and financial applications. Furthermore, specifications for decimal floating-point (DFP) formats and arithmetic operations have been added to the IEEE-754-2008 Standard. Hardware and software support for DFP arithmetic operations have been investigated, especially in the last decade. This paper presents a novel DFP fused multiply-add (DFP-FMA) unit and a new decimal leading-zero anticipator (LZA). The DFP-FMA design uses a previously published parallel fixed-point decimal multiplier for multiplication and a Kogge-Stone parallel prefix adder for decimal addition. The DFP-FMA unit is synthesized using LSI Logic's gflxp 65-nm CMOS standard cell library. Synthesis results show that a 1.07-ns cycle time is obtained when the unit is synthesized with five pipeline stages. Furthermore, a new decimal LZA presented in this paper predicts zero-digit positions exactly and reduces the latency of the DFP-FMA operation. The correctness of the new decimal LZA is fully tested. Ahmet Akkas, Michael J. Schulte |
ASAP | 2 |
| 2011 | Energy-efficient floating-point arithmetic for software-defined radio architecturesabstractThe lack of hardware support for floating-point arithmetic in low-power software-defined radio architectures can significantly increase their software design time due to a time-consuming process of converting floating-point code to fixed-point code. Moreover, emerging wireless communication protocols involve several matrix based algorithms that are extremely sensitive to round-off errors in computations. Using fixed-point arithmetic for these algorithms can significantly impact the accuracy of algorithm results and may incur additional energy overhead due to the extra instructions required for fixed-point arithmetic. In this paper, we demonstrate that supporting floating-point arithmetic in hardware can deliver nearly 30% higher performance and energy efficiency than supporting only fixed-point arithmetic for key kernels of modern wireless communication protocols. The improvements can be further enhanced by our proposed high-throughput floating-point fused-multiply-add unit. Applying our proposed fused-multiply-add unit to key kernels improves performance of the baseline floating-point unit by as much as 60%, while reducing energy consumption by 30% and area by 33%. Although our approach may cause execution stalls depending on data, we show the performance impact of these stalls is negligible. We also employ dynamic range-based dynamic voltage and frequency scaling to further reduce the energy consumption of the processor by 25% for the same worst-case performance as the baseline floating-point implementation. Syed Zohaib Gilani, Nam Sung Kim, Michael J. Schulte |
ASAP | 3 |
| 2011 | Scratchpad memory optimizations for digital signal processing applicationsabstractModern digital signal processors (DSPs) need to support a diverse array of applications ranging from digital filters to video decoding. Many of these applications have drastically different precision and on-chip memory requirements. Moreover, DSPs often employ aggressive dynamic voltage and frequency scaling (DVFS) techniques to minimize power consumption. However, at reduced voltages, process variations can significantly increase the failure rate of on-chip SRAMs designed with small transistors to achieve high integration density, resulting in low yields. Consequently, the size of transistors in SRAM cells and cell size needs to be increased to satisfy the target yield. However, this can result in high area overhead since on-chip memories consume a significant portion of the die area. In this paper, we present a scratchpad memory design that exploits the tradeoffs between SRAM cell sizes, their failure rates, the minimum operating voltage for target yield (Vddmin), and application characteristics to achieve an on-chip memory area reduction of up to 17%. Our approach reduces Vddmin, which allows dynamic and leakage power savings of 42% and 36% respectively with DVFS. Moreover, for error-tolerant DSP applications we allow voltage scaling below Vddminto achieve further power savings while incurring lower mean error as compared to short word-length memory. Finally, for error-sensitive applications, we propose a reconfigurable memory organization that trades memory capacity for higher precision at a lower Vddmin. Syed Zohaib Gilani, Nam Sung Kim, Michael J. Schulte |
DATE | 3 |
| 2011 | Hardware Designs for Binary Integer Decimal-Based RoundingabstractDecimal floating-point (DFP) arithmetic is becoming increasingly important and specifications for it are included in the revised IEEE 754 standard for floating-point arithmetic (IEEE 754-2008). The binary encoding of DFP numbers specified in IEEE 754-2008 is commonly referred to as Binary-Integer Decimal (BID). BID uses a binary integer to encode the significant, which allows it to leverage existing high-speed binary circuits. However, performing decimal rounding on these binary significant is challenging. In this paper, we propose and evaluate several approaches to perform decimal rounding in hardware for DFP numbers that use the BID encoding. We summarize several rounding techniques, present the theory and design of each proposed rounding unit, and use synthesis results to evaluate the critical path delay, latency, and area of rounding units for 64-bit BID numbers. Our results indicate that the bulk of each rounder design is occupied by a binary fixed-point multiplier that can be shared with other integer and floating-point operations. This is the first paper to present and compare a variety of techniques for BID-based rounding hardware. These techniques are valuable to designers of BID-based DFP solutions. Samuel Tsen, Sonia Gonzalez-Navarro, Michael J. Schulte, Katherine Compton |
IEEE Trans. Computers | 3 |
| 2010 | Galois field hardware architectures for network codingabstractThis paper presents and analyzes novel hardware designs for high-speed network coding. Our designs provide efficient methods to perform Galois field (GF) dot products and matrix inversions, which are important operations in network coding. Encoder designs that perform GF dot products and vary with respect to the number of messages combined, Galois field size, and input message size are implemented and analyzed to evaluate design tradeoffs. We investigate single cycle, multicycle, and pipelined designs with and without feedback mechanisms for encoding multiple sets of messages. The decoder is implemented as a multi-cycle design and performs GF matrix inversion followed by multiple GF dot products. Our designs are synthesized with a 65nm standard cell library and compared in terms of area, critical path delay, and throughput. Designs combining four messages achieve throughputs of more than 30 Gbps. Our designs can scale to achieve much higher throughput through the use of additional hardware. Aishwarya Nagarajan, Michael J. Schulte, Parameswaran Ramanathan |
ANCS | 2 |
| 2010 | ERCBench: An Open-Source Benchmark Suite for Embedded and Reconfigurable ComputingabstractResearchers in embedded and reconfigurable computing are often hindered by a lack of suitable benchmarks with which to accurately evaluate their work. Without a suitable benchmark suite, researchers use either outdated, unrealistic benchmarks or spend valuable time creating their own. In this paper, we present ERCBench—a freely-available, open-source benchmark suite geared towards embedded and reconfigurable computing research. ERCBench benchmarks represent a variety of application areas, including multimedia processing, wireless communications, and cryptography. They consist of synthesizable Verilog models for hardware accelerators and hybrid hardware/software applications that combine software-based control flow with hardware-based computation tasks. Daniel W. Chang, Christipher D. Jenkins, Philip C. Garcia, Syed Zohaib Gilani, Paula Aguilera, Aishwarya Nagarajan, Michael J. Anderson, Matthew A. Kenny, Sean M. Bauer, Michael J. Schulte, Katherine Compton |
FPL | 10 |
| 2009 | A Decimal Floating-Point Adder with Decoded Operands and a Decimal Leading-Zero AnticipatorabstractThe IEEE 754-2008 Standard for Floating-Point Arithmetic was officially approved this year. One of the most important revisions to IEEE 754-1985 is the introduction of decimal floating-point (DFP) formats and operations. Since IEEE 754-1985 was revised, major microprocessor vendors have been working on hardware designs and software libraries for decimal arithmetic. Because the new standard has been approved, many software vendors are planning to adapt the new decimal formats into their applications. Therefore, it is important to investigate efficient algorithms and hardware designs for common DFP arithmetic operations to improve the performance of these applications. This paper presents a novel DFP adder with decoded operands and a decimal leading-zero anticipator (LZA). The DFP adder is based on a previous DFP adder design with several new features, including a new internal format, an improved operand pre-correction stage, and a novel decimal LZA to obtain better timing for decimal addition and subtraction. Synthesis results show that the new DFP adder is roughly 14% faster than the previous design. Liang-Kai Wang, Michael J. Schulte |
IEEE Symposium on Computer Arithmetic | 2 |
| 2009 | A Combined Decimal and Binary Floating-Point MultiplierabstractIn this paper, we describe the first hardware design of a combined binary and decimal floating-point multiplier, based on specifications in the IEEE 754-2008 floating-point standard. The multiplier design operates on either (1) 64-bit binary encoded decimal floating-point (DFP) numbers or (2) 64-bit binary floating-point (BFP) numbers. It returns properly rounded results for the rounding modes specified in IEEE 754-2008. The design shares the following hardware resources between the two floating-point datatypes: a 54-bit by 54-bit binary multiplier, portions of the operand encoding/decoding, a 54-bit right shifter, exponent calculation logic, and rounding logic. Our synthesis results show that hardware sharing is feasible and has a reasonable impact on area, latency, and delay. The combined BFP and DFP multiplier occupies only 58% of the total area that would be required by separate BFP and DFP units. Furthermore, the critical path delay of a combined multiplier has a negligible increase over a standalone DFP multiplier, without increasing the number of cycles to perform either BFP or DFP multiplication. Charles Tsen, Sonia Gonzalez-Navarro, Michael J. Schulte, Brian J. Hickmann, Katherine Compton |
ASAP | 3 |
| 2009 | FPGA Design Analysis of the Clustering Algorithm for the CERN Large Hadron ColliderabstractThe Compact Muon Solenoid (CMS) Trigger of the Large Hadron Collider (LHC) particle accelerator at CERN selects potentially interesting particle collision data to process and archive for further study. The first stage of the trigger system, the hardware-based L1 Trigger, must sift through roughly 3 terabits/second of data describing the energy distribution of particles generated in the collisions, and reduce it to 100 megabits/second of event data that subsequent systems can handle. Without the CMS Trigger, the amount of experiment-generated data would quickly outstrip the archiving ability of the LHC system. Because of the sheer amount of input data and the rate at which it is generated, the hardware-based L1 Trigger is subject to stringent performance requirements. These will become even more severe as the LHC is upgraded over the next ten years, requiring a careful redesign of the L1 Trigger hardware. For example, future upgrades may introduce particle motion-tracking data into the L1 Trigger, resulting in an increased input data rate of up to 40 terabits/second. The need to modify the design as the LHC system is upgraded, the low-volume cost advantages of FPGAs, and a desire for a flexible and adaptable system all point toward the use of FPGAs as a hardware implementation solution. In this paper, we present several different FPGA implementations of the electron/photon identification module, a key part of the new Clustering Algorithm for the upgraded L1 Trigger. We analyze the resource requirements and performance tradeoffs, and present a qualitative discussion of flexibility to meet the changing needs of the CMS experiment. Finally, we narrow potential design choices to the top candidates and use one in a full Clustering Algorithm implementation. Anthony E. Gregerson, Amin Farmahini Farahani, Ben Buchli, Steve Naumov, Michail Bachtis, Katherine Compton, Michael J. Schulte, Wesley H. Smith, Sridhara Dasu |
FCCM | 7 |
| 2009 | Performance analysis of decimal floating-point libraries and its impact on decimal hardware and software solutionsabstractThe IEEE Standards Committee recently approved the IEEE 754-2008 Standard for Floating-point Arithmetic, which includes specifications for decimal floating-point (DFP) arithmetic. A growing number of DFP solutions have emerged, and developers now have many DFP design choices including arbitrary or fixed precision, binary or decimal significand encodings, 64-bit or 128-bit DFP operands, and software or hardware implementations. There is a need for accurate analysis of these solutions on representative DFP benchmarks. In this paper, we expand previous DFP benchmark and performance analysis research. We employ a DFP benchmark suite that currently supports several DFP solutions and is easily extendable. We also present performance analysis that (1) provides execution profiles for various DFP encodings and types, (2) gives the average number cycles for common DFP operations and the total number of each DFP operation in each benchmark, and (3) highlights the tradeoffs between using 64-bit and 128-bit DFP operands for both binary and decimal significand encodings. This analysis can help guide the design of future DFP hardware and software solutions. Michael J. Anderson, Chuck Tsen, Liang-Kai Wang, Katherine Compton, Michael J. Schulte |
ICCD | 5 |
| 2009 | Decimal Floating-Point MultiplicationabstractDecimal multiplication is important in many commercial applications including financial analysis, banking, tax calculation, currency conversion, insurance, and accounting. This paper presents the design of two decimal floating-point multipliers: one whose partial product accumulation strategy employs decimal carry-save addition and one that employs binary carry-save addition. The multiplier based on decimal carry-save addition favors a nonpipelined iterative implementation. The multiplier utilizing binary carry-save addition allows for an efficient pipelined implementation when latency and throughput are considered more important than area. Both designs comply with specifications for decimal multiplication given in the IEEE 754 standard for floating-point arithmetic (IEEE 754-2008). The multipliers extend previously published decimal fixed-point multipliers by adding several features, including exponent generation, sticky bit generation, shifting of the intermediate product, rounding, and exception detection and handling. Novel features of the multipliers include support for decimal floating-point numbers, on-the-fly generation of the sticky bit in the iterative design, early estimation of the shift amount, and efficient decimal rounding. Iterative and parallel decimal fixed-point and floating-point multipliers are compared in terms of their area, delay, latency, and throughput based on verified Verilog register-transfer-level models. Mark A. Erle, Brian J. Hickmann, Michael J. Schulte |
IEEE Trans. Computers | 3 |
| 2009 | Low-Power Multiple-Precision Iterative Floating-Point Multiplier with SIMD SupportabstractThe demand for improved SIMD floating-point performance on general-purpose x86-compatible microprocessors is rising. At the same time, there is a conflicting demand in the low-power computing market for a reduction in power consumption. Along with this, there is the absolute necessity of backward compatibility for x86-compatible microprocessors, which includes the support of x87 scientific floating-point instructions. The combined effect is that there is a need for low-power, low-cost floating-point units that are still capable of delivering good SIMD performance while maintaining full x86 functionality. This paper presents the design of an x86-compatible floating-point multiplier (FPM) that is compliant with the IEEE-754 Standard for Binary Floating-Point Arithmetic and is specifically tailored to provide good SIMD performance in a low-cost, low-power solution while maintaining full x87 backward compatibility. The FPM efficiently supports multiple precisions using an iterative rectangular multiplier. The FPM can perform two parallel single-precision multiplies every cycle with a latency of two cycles, one double-precision multiply every two cycles with a latency of four cycles, or one extended-double-precision multiply every three cycles with a latency of five cycles. The iterative FPM also supports division, square-root, and transcendental functions. Compared to a previous design with similar functionality, the proposed iterative FPM has 60 percent less area and 59 percent less dynamic power dissipation. Dimitri Tan, Carl Lemonds, Michael J. Schulte |
IEEE Trans. Computers | 3 |
| 2009 | Hardware Designs for Decimal Floating-Point Addition and Related OperationsabstractDecimal arithmetic is often used in commercial, financial, and Internet-based applications. Due to the growing importance of decimal floating-point (DFP) arithmetic, the IEEE 754-2008 Standard for Floating-Point Arithmetic (IEEE 754-2008) includes specifications for DFP arithmetic. IBM recently announced adding DFP instructions to their POWER6, z9, and z10 microprocessor architectures. As processor support for DFP arithmetic emerges, it is important to investigate efficient arithmetic algorithms and hardware designs for common DFP arithmetic operations. This paper gives an overview of DFP arithmetic in IEEE 754-2008 and discusses previous research on decimal fixed-point and floating-point addition. It also presents novel designs for a DFP adder and a DFP multifunction unit (DFP MFU) that comply with IEEE 754-2008. To reduce their delay, the DFP adder and MFU use decimal injection-based rounding, a new form of decimal operand alignment, and a fast flag-based method for rounding and overflow detection. Synthesis results indicate that the proposed DFP adder is roughly 21 percent faster and 1.6 percent smaller than a previous DFP adder design, when implemented in the same technology. Compared to the DFP adder, the DFP MFU provides six additional operations, yet only has 2.8 percent more delay and 9.7 percent more area. A pipelined version of the DFP MFU has a latency of six cycles, a throughput of one result per cycle, an estimated critical path delay of 12.9 fanout-of-four (F04) inverter delays, and an estimated area of 45, 681 NAND2 equivalent gates. Liang-Kai Wang, Michael J. Schulte, John D. Thompson, Nandini Jairam |
IEEE Trans. Computers | 2 |
| 2008 | Implementing communications systems on an SDR SoCabstractSoftware Defined Radios (SDRs) offer a programmable and dynamically reconfigurable method of reusing hardware to implement the physical layer processing of multiple communications systems. An SDR can dynamically change protocols and update communications systems over the air as a service provider allows. In this paper we present techniques for implementing communications systems in software. We describe briefly the SB3011 platform and programming environment. We then present a number of useful techniques that can be used to implement working systems. We further describe in which systems these techniques are implemented. C. John Glossner, Daniel Iancu, Mayan Moudgill, Sanjay Jinturkar, Gary Nacer, Stuart Stanley, Andrei Iancu, Hua Ye 0003, Michael J. Schulte, Mihai Sima, Tomas Palenik, Peter Farkas, Jarmo Takala |
ICASSP | 9 |
| 2008 | Improved combined binary/decimal fixed-point multipliersabstractDecimal multiplication is important in many commercial applications including banking, tax calculation, currency conversion, and other financial areas. This paper presents several combined binary/decimal fixed-point multipliers that use the BCD-4221 recoding for the decimal digits. This allows the use of binary carry-save hardware to perform decimal addition with a small correction. Our proposed designs contain several novel improvements over previously published designs. These include an improved reduction tree organization to reduce the area and delay of the multiplier and improved reduction tree components that leverage the redundant decimal encodings to help reduce delay. A novel split reduction tree architecture is also introduced that reduces the delay of the binary product with only a small increase in total area. Area and delay estimates are presented that show that the proposed designs have significant area improvements over separate binary and decimal multipliers while still maintaining similar latencies for both decimal and binary operations. Brian J. Hickmann, Michael J. Schulte, Mark A. Erle |
ICCD | 2 |
| 2007 | Decimal Floating-Point Multiplication Via Carry-Save AdditionabstractDecimal multiplication is important in many commercial applications including financial analysis, banking, tax calculation, currency conversion, insurance, and accounting. This paper presents the design of a decimal floating-point multiplier that complies with specifications for decimal multiplication given in the draft revision of the IEEE 754 standard for floating-point arithmetic (IEEE 754R). This multiplier extends a previously published decimal fixed- point multiplier design by adding several features including exponent generation, sticky bit generation, shifting of the intermediate product, rounding, and exception detection and handling. The core of the decimal multiplication algorithm is an iterative scheme of partial product accumulation employing decimal carry-save addition to reduce the critical path delay. Novel features of the proposed multiplier include support for decimal floating-point numbers, on-the- fly generation of the sticky bit, early estimation of the shift amount, and efficient decimal rounding. Area and delay estimates are provided for a verified Verilog register transfer level model of the multiplier. Mark A. Erle, Michael J. Schulte, Brian J. Hickmann |
IEEE Symposium on Computer Arithmetic | 2 |
| 2007 | Decimal Floating-Point Adder and Multifunction Unit with Injection-Based RoundingabstractShrinking feature sizes gives more headroom for designers to extend the functionality of microprocessors. The IEEE 754R working group has revised the IEEE 754-1985 standard for binary floating-point arithmetic to include specifications for decimal floating-point arithmetic and IBM recently announced incorporating a decimal floatingpoint unit into their POWER6 processor. As processor support for decimal floating-point arithmetic emerges, it is important to investigate efficient algorithms and hardware designs for common decimal floating-point arithmetic algorithms. This paper presents novel designs for a decimal floating-point adder and a decimal floating-point multifunction unit. To reduce their delay, both the adder and the multifunction unit use decimal injection-based rounding, a new form of decimal operand alignment, and a fast flag-based method for rounding and overflow detection. Synthesis results indicate that the proposed adder is roughly 21% faster and 1.6% smaller than a previous decimal floating-point adder design, when implemented in the same technology. Compared to the decimal floating-point adder, the decimal floating-point multifunction unit provides six additional operations, yet only has 2.8%more delay and 9.7% more area. Liang-Kai Wang, Michael J. Schulte |
IEEE Symposium on Computer Arithmetic | 2 |
| 2007 | Architecture Support for Reconfigurable Multithreaded Processors in Programmable Communication SystemsabstractExplicitly multithreaded processors and reconfigurable hardware have individually proven to be useful in the design of wireless communication systems. However, new techniques are needed to satisfy the processing requirements of emerging wireless communication standards, which have high throughput requirements for a wide variety of algorithms. This paper presents an efficient technique for adding reconfigurable functional units, called Polymorphic Hardware Accelerators (PHAs), to multicore, multithreaded Digital Signal Processors (DSPs). This paper discusses architectural support to facilitate management and sharing of PHAs on a multithreaded system. The proposed technique shows an average speedup of 6.8 on important wireless communication algorithms in the EEMBC Telecom Benchmark Suite and Department of Defense's Joint Tactical Radio System Software Communication Architecture (JTRS SCA) Hardware Supplement. Suman Mamidi, Michael J. Schulte, Daniel Iancu, C. John Glossner |
ASAP | 2 |
| 2007 | Hardware Design of a Binary Integer Decimal-based IEEE P754 Rounding UnitabstractBecause of the growing importance of decimal floating-point (DFP) arithmetic, specifications for it were recently added to the draft revision of the IEEE 754 Standard (IEEE P754). In this paper, we present a hardware design for a rounding unit for 64-bit DFP numbers (decimal 64) that use the IEEE P754 binary encoding of DFP numbers, which is widely known as the Binary Integer Decimal (BID) encoding. We summarize the technique used for rounding, present the theory and design of the BID rounding unit, and evaluate its critical path delay, latency, and area for combinational and pipelined designs. Over 86% of the rounding unit's area is due to a 55-bit by 54-bit binary multiplier, which can be shared with a double-precision binary floating-point multiplier. To our knowledge, this is the first hardware design for rounding IEEE P754 BID-encoded DFP numbers. Charles Tsen, Michael J. Schulte, Sonia Gonzalez-Navarro |
ASAP | 2 |
| 2007 | A parallel IEEE P754 decimal floating-point multiplierabstractDecimal floating-point multiplication is important in many commercial applications including banking, tax calculation, currency conversion, and other financial areas. This paper presents a fully parallel decimal floating-point multiplier compliant with the recent draft of the IEEE P754 Standard for Floating-point Arithmetic (IEEE P754). The novelty of the design is that it is the first parallel decimal floating-point multiplier offering low latency and high throughput. This design is based on a previously published parallel fixed-point decimal multiplier which uses alternate decimal digit encodings to reduce area and delay. The fixed-point design is extended to support floating-point multiplication by adding several components including exponent generation, rounding, shifting, and exception handling. Area and delay estimates are presented that show a significant latency and throughput improvement with a substantial increase in area as compared to the only published IEEE P754 compliant sequential floating-point multiplier. To the best of our knowledge, this is the first publication to present a fully parallel decimal floating-point multiplier that complies with IEEE P754. Brian J. Hickmann, Andrew Krioukov, Michael J. Schulte, Mark A. Erle |
ICCD | 3 |
| 2007 | Floating-point division algorithms for an x86 microprocessor with a rectangular multiplierabstractFloating-point division is an important operation in scientific computing and multimedia applications. This paper presents and compares two division algorithms for an times86 microprocessor, which utilizes a rectangular multiplier that is optimized for multimedia applications. The proposed division algorithms are based on Goldschmidt's division algorithm and provide correctly rounded results for IEEE 754 single, double, and extended precision floating-point numbers. Compared to a previous Goldschmidt division algorithm, the fastest proposed algorithm requires 25% to 37% fewer cycles, while utilizing a multiplier that is roughly 2.5 times smaller. Michael J. Schulte, Dimitri Tan, Carl Lemonds |
ICCD | 1 |
| 2007 | Hardware design of a Binary Integer Decimal-based floating-point adderabstractBecause of the growing importance of decimal floating-point (DFP) arithmetic, specifications for it are included in the IEEE Draft Standard for Floating-point Arithmetic (IEEE P754). In this paper, we present a novel algorithm and hardware design for a DFP adder. The adder performs addition and subtraction on 64-bit operands that use the IEEE P754 binary encoding of DFP numbers, widely known as the binary integer decimal (BID) encoding. The BID adder uses a novel hardware component for decimal digit counting and an enhanced version of a previously published BID rounding unit. By adding more sophisticated control, operations are performed with variable latency to optimize for common cases. We show that a BID-based DFP adder design can be achieved with a modest area increase compared to a single 2-stage pipelined 64-bit fixed-point multiplier. Over 70% of the BID adderpsilas area is due the 64-bit fixed-point multiplier, which can be shared with a binary floating-point multiplier and hardware for other DFP operations. To our knowledge, this is the first hardware design for adding and subtracting IEEE P754 BID-encoded DFP numbers. Charles Tsen, Sonia Gonzalez-Navarro, Michael J. Schulte |
ICCD | 3 |
| 2007 | Benchmarks and performance analysis of decimal floating-point applicationsabstractThe IEEE P754 draft standard for floating-point arithmetic provides specifications for decimal floating-point (DFP) formats and operations. Based on this standard, many developers will provide support for DFP calculations. We present a benchmark suite for DFP applications and use this suite to evaluate the performance of hardware and software DFP solutions. Our benchmarks include banking, commerce, risk-management, tax, and telephone billing applications organized into a suite of five macro benchmarks. In addition to developing our own applications, we leverage open-source projects and academic financial analysis applications. The benchmarks are modular, making them easy to adapt for different DFP solutions. We use the benchmarks to evaluate the performance of the decNumber DFP library and an extended version of the SimpleScalar PISA architecture with hardware and instruction set support for DFP operations. Our analysis shows that providing processor support for high-speed DFP operations significantly improves the performance of DFP applications. Liang-Kai Wang, Charles Tsen, Michael J. Schulte, Divya Jhalani |
ICCD | 3 |
| 2007 | Software Solutions for Converting a MIMO-OFDM Channel into Multiple SISO-OFDM Channels
Mihai Sima, Murugappan Senthilvelan, Daniel Iancu, C. John Glossner, Mayan Moudgill, Michael J. Schulte |
WiMob | 6 |
| 2006 | Dual-mode floating-point multiplier architectures with parallel operations
Ahmet Akkas, Michael J. Schulte |
J. Syst. Archit. | 2 |
| 2006 | Integer Multipliers with Overflow DetectionabstractThis paper presents a general approach for designing array and tree integer multipliers with overflow detection. The overflow detection techniques are based on an analysis of the magnitudes of the input operands. The overflow detection circuits operate in parallel with a simplified multiplier to reduce the overall area and delay Mustafa Gök, Michael J. Schulte, Mark G. Arnold |
IEEE Trans. Computers | 2 |
| 2006 | Generation and visualization of four-dimensional MR angiography data using an undersampled 3-D projection trajectoryabstractTime-resolved contrast-enhanced magnetic resonance (MR) angiography (CE-MRA) has gained in popularity relative to X-ray Digital Subtraction Angiography because it provides three-dimensional (3-D) spatial resolution and it is less invasive. We have previously presented methods that improve temporal resolution in CE-MRA while providing high spatial resolution by employing an undersampled 3-D projection (3D PR) trajectory. The increased coverage and isotropic resolution of the 3D PR acquisition simplify visualization of the vasculature from any perspective. We present a new algorithm to develop a set of time-resolved 3-D image volumes by preferentially weighting the 3D PR data according to its acquisition time. An iterative algorithm computes a series of density compensation functions for a regridding reconstruction, one for each time frame, that exploit the variable sampling density in 3D PR. The iterative weighting procedure simplifies the calculation of appropriate density compensation for arbitrary sampling patterns, which improve sampling efficiency and, thus, signal-to-noise ratio and contrast-to-noise ratio, since it is does not require a closed-form calculation based on geometry. Current medical workstations can display these large four-dimensional studies, however, interactive cine animation of the data is only possible at significantly degraded resolution. Therefore, we also present a method for interactive visualization using powerful graphics cards and distributed processing. Results from volunteer and patient studies demonstrate the advantages of dynamic imaging with high spatial resolution. Michael J. Redmond, Ethan K. Brodsky, Andrew L. Alexander, Aiming Lu, F. J. Thornton, Michael J. Schulte, Tom M. Grist, James G. Pipe, Walter F. Block |
IEEE Trans. Medical Imaging | 7 |
| 2005 | Decimal Multiplication with Efficient Partial Product GenerationabstractDecimal multiplication is important in many commercial applications including financial analysis, banking, tax calculation, currency conversion, insurance, and accounting. This paper presents a novel design for fixed-point decimal multiplication that utilizes a simple recoding scheme to produce signed-magnitude representations of the operands thereby greatly simplifying the process of generating partial products for each multiplier digit. The partial products are generated using a digit-by-digit multiplier on a word-by-digit basis, first in a signed-digit form with two digits per position, and then combined via a combinational circuit. As the signed-digit partial products are developed one at a time while traversing the recoded multiplier operand from the least significant digit to the most significant digit, each partial product is added along with the accumulated sum of previous partial products via a signed-digit adder. This work is significantly different from other work employing digit-by-digit multipliers due to the efficiency gained by restricting the range of digits throughout the multiplication process. Mark A. Erle, Eric M. Schwarz, Michael J. Schulte |
IEEE Symposium on Computer Arithmetic | 3 |
| 2005 | Efficient Function Approximation Using Truncated Multipliers and SquarersabstractThis paper presents a technique for designing linear and quadratic interpolators for function approximation using truncated multipliers and squarers. Initial coefficient values are found using a Chebyshev series approximation, and then adjusted through exhaustive simulation to minimize the maximum absolute error of the interpolator output. This technique is suitable for any function and any precision up to 24-bits (IEEE single precision). Designs for linear and quadratic interpolators that implement the reciprocal function, f(x)=1/x, are presented and analyzed as an example. We show that a 24-bit truncated reciprocal quadratic interpolator with a design specification /spl plusmn/1 ulp error requires 24.1% fewer partial products to implement than a comparable standard interpolator with the same error specification. E. George Walters III, Michael J. Schulte |
IEEE Symposium on Computer Arithmetic | 2 |
| 2005 | Instruction Set Extensions for Reed-Solomon Encoding and DecodingabstractReed-Solomon codes are an important class of error correcting codes used in many applications related to communications and digital storage. The fundamental operations in Reed-Solomon encoding and decoding involve Galois field arithmetic which is not directly supported in general purpose processors. On the other hand, pure hardware implementations of Reed-Solomon coders are not programmable. In this paper, we present a novel algorithm to perform Reed-Solomon encoding. We also propose four new instructions for Galois field arithmetic. We show that by using the instructions, we can speedup Reed-Solomon decoding by a factor of 12 compared to a pure software approach, while still maintaining programmability. Suman Mamidi, Daniel Iancu, Andrei Iancu, Michael J. Schulte, C. John Glossner |
ASAP | 4 |
| 2005 | Decimal Floating-Point Square Root Using Newton-Raphson IterationabstractWith continued reductions in feature size, additional functionality may be added to future microprocessors to boost the performance of important application domains. Due to growth in commercial, financial, and Internet-based applications, decimal floating point arithmetic is now attracting more attention and hardware support for decimal operations is being considered by various computer manufacturers. In order to standardize decimal number formats and operations, specifications for decimal floating point arithmetic have been added to the draft revision of the IEEE-754 standard for floating point arithmetic (IEEE-754R). This paper presents an efficient arithmetic algorithm and hardware design for decimal floating point square root. This design uses an optimized piecewise linear approximation, a modified Newton-Raphson iteration, a specialized rounding technique, and a modified decimal multiplier. Synthesis results show that a 64-bit (16 digit) implementation of decimal square root, which is compliant with IEEE-754R, has an estimated critical path delay of 0.95 ns and a maximum latency of 210 clock cycles when implemented using a sequential multiplier and LSI Logic's 0.11 micron Gflx-P standard cell library. Liang-Kai Wang, Michael J. Schulte |
ASAP | 2 |
| 2005 | Instruction set extensions for software defined radio on a multithreaded processorabstractSoftware defined radios, which provide a programmable solution for implementing the physical layer processing of multiple communication standards, are widely recognized as one of the most important new technologies for wireless communication systems. Emerging communication standards, however, require tremendous processing capabilities to perform high-bandwidth physical-layer processing in real time. In this paper, we present instruction set extensions for several important communication algorithms including convolutional encoding, Viterbi decoding, turbo decoding, and Reed-Solomon encoding and decoding. The performance benefits of these extensions are evaluated using a supercomputer class vectorizing compiler and the Sandblaster low-power multithreaded processor for software defined radio. The proposed instruction set extensions provide significant performance improvements, while maintaining a high degree of programmability. Suman Mamidi, Emily R. Blem, Michael J. Schulte, C. John Glossner, Daniel Iancu, Andrei Iancu, Mayan Moudgill, Sanjay Jinturkar |
CASES | 3 |
| 2005 | High-Speed Multioperand Decimal AddersabstractThere is increasing interest in hardware support for decimal arithmetic as a result of recent growth in commercial, financial, and Internet-based applications. Consequently, new specifications for decimal floating-point arithmetic have been added to the draft revision of the IEEE-754 Standard for floating-point arithmetic. This paper introduces and analyzes three techniques for performing fast decimal addition on multiple binary coded decimal (BCD) operands. Two of the techniques speculate BCD correction values and correct intermediate results while adding the input operands. The first speculates over one addition. The second speculates over two additions. The third technique uses a binary carry-save adder tree and produces a binary sum. Combinational logic is then used to correct the sum and determine the carry into the next more significant digit. Multioperand adder designs are constructed and synthesized for four to 16 input operands. Analyses are performed on the synthesis results and the merits of each technique are discussed. Finally, these techniques are compared to several previous techniques for high-speed decimal addition. Robert D. Kenney, Michael J. Schulte |
IEEE Trans. Computers | 2 |
| 2004 | A Low-Power Carry Skip Adder with Fast Saturation
Michael J. Schulte, Kai Chirca, C. John Glossner, Haoran Wang 0003, Suman Mamidi, Pablo I. Balzola, Stamatis Vassiliadis |
ASAP | 1 |
| 2004 | Decimal Floating-Point Division Using Newton-Raphson Iteration
Liang-Kai Wang, Michael J. Schulte |
ASAP | 2 |
| 2004 | A Static Low-Power, High-Performance 32-bit Carry Skip AdderabstractIn this paper, we present a full-static carry-skip adder designed to achieve low power dissipation and high-performance operation. To reduce the adder's delay and power consumption, the adder is divided into variable-sized blocks that balance the inputs to the carry chain. The optimum block sizes for minimizing the critical path delay with complementary carry generation are achieved. Within blocks, highly optimized carry look-ahead logic, which computes block generate and block propagate signals, is used to further decrease delay. The adder architecture decreases power consumption by reducing the number of logic levels, glitches, and transistors. To achieve balanced delay, input bits are grouped unevenly in the carry chain. This grouping reduces active power by minimizing extraneous glitches and transitions. The adder has been implemented in 130nm CMOS technology. At 1.2V and 25C, typical performance is 1.086GHz and power dissipation normalized to 600 MHz operation is 0.786 mW. Kai Chirca, Michael J. Schulte, C. John Glossner, Haoran Wang 0003, Suman Mamidi, Pablo I. Balzola, Stamatis Vassiliadis |
DSD | 2 |
| 2004 | A High-Frequency Decimal MultiplierabstractDecimal arithmetic is regaining popularity in the computing community due to the growing importance of commercial, financial, and Internet-based applications, which process decimal data. This paper presents an iterative decimal multiplier, which operates at high clock frequencies and scales well to large operand sizes. The multiplier uses a new decimal representation for intermediate products, which allows for a very fast two-stage iterative multiplier design. Decimal multipliers, which are synthesized using a 0.11 micron CMOS standard cell library, operate at clock frequencies close to 2 GHz. The latency of the proposed design to multiply two n-digit BCD operands is (n+8) cycles with a new multiplication able to begin every (n+1) cycles. Robert D. Kenney, Michael J. Schulte, Mark A. Erle |
ICCD | 2 |
| 2003 | The Interval Logarithmic Number SystemabstractWe introduce the interval logarithmic number system (ILNS), in which the logarithmic number system (LNS) is used as the underlying number system for interval arithmetic. The basic operations in ILNS are introduced and an efficient method for performing ILNS addition and subtraction is presented. We compare ILNS to interval floating point (IFP) for a few sample applications. For applications like the N-body problem, which have a large percentage of multiplies, divides and square roots, ILNS provides much narrower intervals than IFP. In other applications, like the fast Fourier transform, where addition and subtraction dominate, ILNS and IFP produce intervals having similar widths. Based on our analysis, ILNS is an attractive alternative to IFP for application that can tolerate low to moderate precisions. Mark G. Arnold, Michael J. Schulte |
IEEE Symposium on Computer Arithmetic | 3 |
| 2003 | Decimal Multiplication Via Carry-Save AdditionabstractDecimal multiplication is important in many commercial applications including financial analysis, banking, tax calculation, currency conversion, insurance, and accounting. We present two novel designs for fixed-point decimal multiplication that utilize decimal carry-save addition to reduce the critical path delay. First, a multiplier that stores a reduced number of multiplicand multiples and uses decimal carry-save addition in the iterative portion of the design is presented. Then, a second multiplier design is proposed with several notable improvements including fast generation of multiplicand multiples that do not need to be stored, the use of decimal (4:2) compressors, and a simplified decimal carry-propagate addition to produce the final product. When multiplying two n-digit operands to produce a 2n-digit product, the improved multiplier design has a worst-case latency of n+4 cycles and an initiation interval of n+1 cycles. Three data-dependent optimizations, which help reduce the multipliers' average latency, are also described. The multipliers presented can be extended to support decimal floating-point multiplication. Mark A. Erle, Michael J. Schulte |
ASAP | 2 |
| 2003 | Combined Multiplication and Sum-of-Squares UnitsabstractMultiplication and squaring are important operations in digital signal processing and multimedia applications. We present designs for units that implement either multiplication, A/spl times/B, or sum-of-squares computations, A/sup 2/+B/sup 2/, based on an input control signal. Compared to conventional parallel multipliers, these units have a modest increase in area and delay, but allow either multiplication or sum-of-squares computations to be performed. Combined multiplication and sum-of-squares units for unsigned and two's complement operands are presented, along with integrated designs that can operate on either unsigned or two's complement operands. The designs can also be extended to work with a third accumulator operand to compute either Z+A/spl times/B or Z+A/sup 2/+B/sup 2/. Synthesis results indicate that a combined multiplication and sum-of-squares unit for 32-bit two's complement operands can be implemented with roughly 15% more area and nearly the same worst case delay as a conventional 32-bit two's complement multiplier. Michael J. Schulte, Louis Marquette, Shankar Krithivasan, E. George Walters III, C. John Glossner |
ASAP | 1 |
| 2003 | A Quadruple Precision and Dual Double Precision Floating-Point MultiplierabstractDouble precision floating-point arithmetic is inadequate for many scientific computations. This paper presents the design of a quadruple precision floating-point multiplier that also supports two parallel double precision multiplications. Since hardware support for quadruple precision arithmetic is expensive, a new technique is presented that requires much less hardware than a fully parallel quadruple precision multiplier. With this implementation, quadruple precision multiplication has a latency of three cycles and two parallel double precision multiplications have a latency of only two cycles. The design is pipelined so that two double precision multiplications can be started every cycle or a quadruple precision multiplication can be started every other cycle. Ahmet Akkas, Michael J. Schulte |
DSD | 2 |
| 2001 | Analysis of Column Compression MultipliersabstractColumn compression multipliers are frequently used in high-performance computer systems due to their short worst case delay. This paper examines the area, delay, and power characteristics of Dadda (1965) and Wallace (1964) column compression multipliers in deep submicron technology. Our analysis shows that Wallace multipliers have slightly more area and approximately the same worst case delay as Dadda multipliers. It also shows the importance of considering parasitic capacitances when determining the delay of column compression multipliers, since parasitics can increase the delay of the multiplier by over 60%. As multiplier size increases, the ratio of power to area also increases, due to longer interconnect lines and increased glitching. K'Andrea C. Bickerstaff, Earl E. Swartzlander Jr., Michael J. Schulte |
IEEE Symposium on Computer Arithmetic | 3 |
| 2001 | FPGA Resource Reduction Through Truncated Multiplication
Kent E. Wires, Michael J. Schulte, Don McCarley |
FPL | 2 |
| 2001 | Design Alternatives for Parallel Saturating Multioperand AddersabstractParallel saturating multioperand adders significantly improve the performance of global system for mobile (GSM) speech coders by giving compilers and assembly language programmers the ability to parallelize loops containing saturating dot products, while maintaining GSM compliant results. This paper presents four designs for parallel saturating multioperand adders. These designs have at most one carry-propagate adder on their critical delay path, yet produce the same results that would be obtained if the additions were performed serially with saturation after each addition. The four parallel designs offer tradeoffs in terms of area, worst case delay, and dot product latency. Compared to a 5-input serial design, the 5-input parallel designs have delays up to 3.51 times shorter. Pablo I. Balzola, Michael J. Schulte, Jie Ruan, C. John Glossner, Erdem Hokenek |
ICCD | 2 |
| 2001 | Combined IEEE Compliant and Truncated Floating Point Multipliers for Reduced Power DissipationabstractTruncated multiplication can be used to significantly reduce power dissipation for applications that do not require correctly rounded results. This paper presents a power efficient method for designing floating point multipliers that can perform either correctly rounded IEEE compliant multiplication or truncated multiplication, based on an input control signal. Compared to conventional IEEE floating point multipliers, these multipliers require only a small amount of additional area and delay, yet provide a significant reduction in power dissipation for applications that do not require IEEE compliant results. Kent E. Wires, Michael J. Schulte, James E. Stine |
ICCD | 2 |
| 2000 | A Hardware Algorithm for Variable-Precision LogarithmabstractThis paper presents an efficient hardware algorithm for variable-precision logarithm. The algorithm uses an iterative technique that employs table lookups and polynomial approximations. Compared to similar algorithms, it reduces the number of fixed-precision operations by avoiding full precision computations and dynamically varying the precision of intermediate results. It also uses significantly smaller tables than related algorithms. For a specified hardware implementation, the algorithm requires fewer than 2L/sup 2/ fixed-precision multiplications to evaluate the logarithm to L words of precision. An error analysis for the algorithm is also presented. Javier Hormigo, Julio Villalba, Michael J. Schulte |
ASAP | 3 |
| 2000 | Parallel saturating multioperand addersabstractThis paper presents designs for parallel saturating multioperand adders. These adders have only a single carrypropagate adder on the critical delay path, yet produce the same results that would be obtained if the additions were performed serially with saturation after each operation. When used with parallel saturating multipliers or multiplyaccumulate units, these adders signi cantly improve the performance of GSM speech coders. They can also easily be modi ed to perform either saturating or wraparound multioperand addition, based on an input control signal. Since parallel saturating multioperand adders have more area and less delay than serial saturating multioperand adders, they are suitable for high-performance digital signal processing systems. 1. Michael J. Schulte, Pablo I. Balzola, Jie Ruan, C. John Glossner |
CASES | 1 |
| 2000 | Integer Multiplication with Overflow Detection or SaturationabstractHigh-speed multiplication is frequently used in general-purpose and application-specific computer systems. These systems often support integer multiplication, where two n-bit integers are multiplied to produce a 2n-bit product. To prevent growth in word length, processors typically return the n least significant bits of the product and a flag that indicates whether or not overflow has occurred. Alternatively, some processors saturate results that overflow to the most positive or most negative representable number. This paper presents efficient methods for performing unsigned or two's complement integer multiplication with overflow detection or saturation. These methods have significantly less area and delay than conventional methods for integer multiplication with overflow detection or saturation. Michael J. Schulte, Pablo I. Balzola, Ahmet Akkas, Robert W. Brocato |
IEEE Trans. Computers | 1 |
| 2000 | A Family of Variable-Precision Interval Arithmetic ProcessorsabstractTraditional computer systems often suffer from roundoff error and catastrophic cancellation in floating point computations. These systems produce apparently high precision results with little or no indication of the accuracy. This paper presents hardware designs, arithmetic algorithms, and software support for a family of variable-precision, interval arithmetic processors. These processors give the programmer the ability to detect and, if desired, to correct implicit errors in finite precision numerical computations. They also provide the ability to solve problems that cannot be solved efficiently using traditional floating point computations. Execution time estimates indicate that these processors are two to three orders of magnitude faster than software packages that provide similar functionality. Michael J. Schulte, Earl E. Swartzlander Jr. |
IEEE Trans. Computers | 1 |
| 1999 | High-Speed Inverse Square RootsabstractInverse square roots are used in several digital signal processing, multimedia, and scientific computing applications. This paper presents a high-speed method for computing inverse square roots. This method uses a table lookup, operand modification, and multiplication to obtain an initial approximation to the inverse square root. This is followed by a modified Newton-Raphson iteration, consisting of one square, one multiply-complement, and one multiply-add operation. The initial approximation and Newton-Raphson iteration employ specialized hardware to reduce the delay, area, and power dissipation. Application of this method is illustrated through the design of an inverse square root unit operands in the IEEE single precision format. An implementation of this unit with a 4-layer metal, 2.5 Volt, 0.25 micron CMOS standard cell library has a cycle rime of 6.7 ns, an area of 0.41 mm/sup 2/, a latency of five cycles, and a throughput of one result per cycle. Michael J. Schulte, Kent E. Wires |
IEEE Symposium on Computer Arithmetic | 1 |
| 1999 | Parallel Saturating Fractional Arithmetic UnitsabstractThis paper describes the designs of a saturating adder multiplier single MAC unit, and dual MAC unit with single cycle latencies. The dual MAC unit can perform two saturating MAC operations in parallel and accumulate the results with saturation. Specialized saturation logic ensures that the output of the dual MAC unit is identical to the result of the operations performed serially with saturation after each multiplication and each addition. Navindra Yadav, Michael J. Schulte, C. John Glossner |
Great Lakes Symposium on VLSI | 2 |
| 1999 | Approximating Elementary Functions with Symmetric Bipartite TablesabstractThis paper presents a high-speed method for function approximation that employs symmetric bipartite tables. This method performs two parallel table lookups to obtain a carry-save (borrow-save) function approximation, which is either converted to a two's complement number or is Booth encoded. Compared to previous methods for bipartite table approximations, this method uses less memory by taking advantage of symmetry and leading zeros in one of the two tables. It also has a closed-form solution for the table entries, provides tight bounds on the maximum absolute error, and can be applied to a wide range of functions. A variation of this method provides accurate initial approximations that are useful in multiplicative divide and square root algorithms. Michael J. Schulte, James E. Stine |
IEEE Trans. Computers | 1 |
| 1998 | A Combined Interval and Floating Point MultiplierabstractInterval arithmetic provides an efficient method for monitoring and controlling errors in numerical calculations. However, existing software packages for interval arithmetic are often too slow for numerically intensive computations. This paper presents the design of a multiplier that performs either interval or floating point multiplication. This multiplier requires only slightly more area and delay than a conventional floating point multiplier, and is one to two orders of magnitude faster than software implementations of interval multiplication. James E. Stine, Michael J. Schulte |
Great Lakes Symposium on VLSI | 2 |
| 1997 | Symmetric Bipartite Tables for Accurate Function ApproximationabstractThe paper presents a methodology for designing bipartite tables for accurate function approximation. Bipartite tables use two parallel table lookups to obtain a carry-save (borrow-save) function approximation. A carry propagate adder can then convert this approximation to a two's complement number or the approximation can be directly Booth encoded. Our method for designing bipartite tables, called the Symmetric Bipartite Table Method, utilizes symmetry in the table entries to reduce the overall memory requirements. It has several advantages over previous bipartite table methods in that it: (1) provides a closed form solution for the table entries; (2) has right bounds on the maximum absolute error; (3) requires smaller table lookups to achieve a given accuracy; and (4) can be applied to a wide range of functions. Compared to conventional table lookups, the symmetric bipartite tables presented are 15.0 to 41.7 times smaller when the operand size is 16 bits and 99.1 to 273.9 times smaller when the operand size is 24 bits. Michael J. Schulte, James E. Stine |
IEEE Symposium on Computer Arithmetic | 1 |
| 1997 | Accurate Function Approximations by Symmetric Table Lookup and AdditionabstractThis paper presents a high-speed method for accurate function approximations. This method employs parallel table lookups followed by multi-operand addition. It takes advantage of leading zeros and symmetry in the table entries to reduce the table sizes. By increasing the number of tables and the number of operands in the multi-operand addition, the amount of memory is significantly reduced. This method provides a closed form solution for the table entries and can be applied to a variety of elementary functions. Compared to conventional table lookups, it requires two to three orders of magnitude less memory. The design of elementary function generators that use this method are presented and compared to similar methods for elementary function generation. Michael J. Schulte, James E. Stine |
ASAP | 1 |
| 1995 | The K5 transcendental functionsabstractThis paper describes the development of the transcendental instructions for the K5, AMD's recently completed x86 compatible superscalar microprocessor. A multi-level development cycle, with testing between levels, facilitated the early detection of errors and limited their effect on the design schedule. The algorithms for the transcendental functions use table-driven reductions followed by polynomial approximations. Multiprecision arithmetic operations are used when necessary to maintain sufficient accuracy and to ensure that the transcendental functions have a maximum error of one unit in the last place.> Thomas W. Lynch, Ashraf Ahmed, Michael J. Schulte, Thomas K. Callaway, Robert Tisdale |
IEEE Symposium on Computer Arithmetic | 3 |
| 1995 | Hardware Design and Arithmetic Algorithms for a Variable-Precision, Interval Arithmetic CoprocessorabstractThis paper presents the hardware design and arithmetic algorithms for a coprocessor that performs variable-precision, interval arithmetic. The coprocessor gives the programmer the ability to specify the precision of the computation, determine the accuracy of the result, and recompute inaccurate results with higher precision. Direct hardware support and efficient algorithms for variable-precision, interval arithmetic greatly improve the speed, accuracy, and reliability of numerical computations. Performance estimates indicate that the coprocessor is 200 to 1,000 times faster than a software package for variable-precision, interval arithmetic. The coprocessor can be implemented on a single chip with a cycle time that is comparable to IEEE double-precision floating point coprocessors.> Michael J. Schulte, Earl E. Swartzlander Jr. |
IEEE Symposium on Computer Arithmetic | 1 |
| 1995 | A Processor for Staggered Interval ArithmeticabstractThe paper presents the design of a high-speed processor which performs staggered interval arithmetic. Each staggered interval is represented as the sum of a set of floating point numbers plus an interval, which consists of two floating point endpoints. Staggered interval arithmetic allows the precision of the computation to be specified and the accuracy of the result to be determined. Efficient arithmetic algorithms, which reduce the number of floating point operations needed to perform staggered interval arithmetic, are introduced. To achieve high performance, the processor employs an array of pipelined floating point arithmetic units and two long accumulators. The processor provides direct hardware support for accurate and numerically reliable vector and matrix computations. Michael J. Schulte, Earl E. Swartzlander Jr. |
ASAP | 1 |
| 1995 | A coprocessor for accurate and reliable numerical computationsabstractThis paper presents the architecture and hardware design of a special-purpose coprocessor that performs variable-precision, interval arithmetic. Variable-precision arithmetic allows the precision of the computation to be specified, based on the problem to be solved and the required accuracy of the results. Interval arithmetic produces two values for each result, such that the true result is guaranteed to be between the two values. The coprocessor gives the programmer the ability to specify the precision of the computation, determine the accuracy of the results, and recompute inaccurate results with higher precision. Direct hardware support for variable-precision, interval arithmetic greatly improves the accuracy and reliability of numerical computations. Execution time estimates indicate that the coprocessor is two to three orders of magnitude faster than an existing software package for variable-precision, interval arithmetic. Michael J. Schulte, Earl E. Swartzlander Jr. |
ICCD | 1 |
| 1994 | A variable-precision interval arithmetic processorabstractThis paper presents a special-purpose processor which implements variable-precision, interval arithmetic. Variable-precision arithmetic allows the precision of the computation to be specified, based on the problem to be solved and the required accuracy of the computation. Interval arithmetic produces two values for each result, such that the true result is guaranteed to be between the two values. The distance between the two values gives an upper bound on the error. Direct hardware support for variable-precision, interval arithmetic greatly improves the accuracy of the computation, and is much faster than existing software methods for controlling numerical error. Area and delay estimates indicate that the processor can be implemented on a single chip with a cycle time which is comparable to existing IEEE double-precision floating point processors. For computationally intensive problems, an application-specific array of variable-precision, interval arithmetic processors can execute in parallel to provide high-performance and numerically reliable results.> Michael J. Schulte, Earl E. Swartzlander Jr. |
ASAP | 1 |
| 1994 | Hardware Designs for Exactly Rounded Elemantary FunctionsabstractThis paper presents hardware designs that produce exactly rounded results for the functions of reciprocal, square-root, 2/sup x/, and log/sub 2/(x). These designs use polynomial approximation in which the terms in the approximation are generated in parallel, and then summed by using a multi-operand adder. To reduce the number of terms in the approximation, the input interval is partitioned into subintervals of equal size, and different coefficients are used for each subinterval. The coefficients used in the approximation are initially determined based on the Chebyshev series approximation. They are then adjusted to obtain exactly rounded results for all inputs. Hardware designs are presented, and delay and area comparisons are made based on the degree of the approximating polynomial and the accuracy of the final result. For single-precision floating point numbers, a design that produces exactly rounded results for all four functions has an estimated delay of 80 ns and a total chip area of 98 mm/sup 2/ in a 1.0-micron CMOS technology. Allowing the results to have a maximum error of one unit in the last place reduces the computational delay by 5% to 30% and the area requirements by 33% to 77%.> Michael J. Schulte, Earl E. Swartzlander Jr. |
IEEE Trans. Computers | 1 |
| 1993 | Exact rounding of certain elementary functionsabstractAn algorithm is described which produces exactly rounded results for the functions of reciprocal, square root, 2/sup x/, and log 2/sup x/. Hardware designs based on this algorithm are presented for floating point numbers with 16- and 24-b significands. These designs use a polynomial approximation in which coefficients are originally selected based on the Chebyshev series approximation and are then adjusted to ensure exactly rounded results for all inputs. To reduce the number of terms in the approximation, the input interval is divided into subintervals of equal size and different coefficients are used for each subinterval. For floating point numbers with 16-b significands, the exactly rounded value of the function can be computed in 51 ns on a 20-mm/sup 2/ chip. For floating point numbers with 24-b significands, the functions can be computed in 80 ns on a 98-mm/sup 2/ chip.> Michael J. Schulte, Earl E. Swartzlander Jr. |
IEEE Symposium on Computer Arithmetic | 1 |
| 1993 | Reduced area multipliersabstractAs developed by Wallace (1964) and Dadda (1965), a high-speed method for the parallel multiplication of two binary numbers is to reduce their partial products to two numbers whose sum is equal to the product. The resulting two numbers are then summed using a fast carry-propagate adder. The authors present a multiplier, the reduced area multiplier, with a novel reduction scheme which results in fewer components and less interconnect overhead than either Wallace or Dadda multipliers. This reduction scheme is especially useful for pipelined multipliers, because it minimizes the number of latches required in the reduction of the partial products. Equations are given for determining the number of components and a method is presented for estimating the interconnect overhead for Wallace, Dadda and reduced area multipliers. Area estimates indicate that pipelined reduced area multipliers require 3 to 8% less area than equivalent Wallace multipliers and 15 to 25% less area than equivalent Dadda multipliers.> K'Andrea C. Bickerstaff, Michael J. Schulte, Earl E. Swartzlander Jr. |
ASAP | 2 |