John L. Gustafson

dblp:44/4606 · DBLP profile ↗
← Back
29ranked-venue papers
10as first author
5since 2021 · last 2026
0000-0002-2957-1304ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 8 first-author · 2 since 2021Software engineering, systems software and programming languages · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 The Early Electronic Digital Computers - Panel Summary
B. Jack Copeland, Vladimir Getov, John L. Gustafson, Brian Stuart
COMPSAC3
2026 Energy-efficient posit systolic array for real-time CNN inference using approximate computing on FPGA
abstract
With the increasing adoption of Convolutional Neural Networks (CNNs) in edge and embedded applications, there is a growing demand for hardware accelerators that provide high computational accuracy while maintaining low energy consumption. In this work, we present a reconfigurable CNN accelerator specifically designed for real-time object detection. The accelerator is based on a systolic array architecture and integrates the posit number system to enhance both accuracy and dynamic range in numerical representation. To further improve energy efficiency, the design incorporates a set of lightweight approximation techniques applied to arithmetic operations. In particular, the Multiply and Accumulate (MAC) units employ approximate strategies for multiplication, normalization, and rounding of the fraction part, selectively simplifying low-significance computations to reduce hardware cost, while ensuring that the final classification accuracy remains within acceptable limits. The proposed architecture is implemented at the Register-Transfer Level (RTL) level and evaluated on a Xilinx ZCU102 FPGA using ResNet-50 and the ImageNet dataset. Experimental results demonstrate a 21% reduction in power consumption and an 18% decrease in resource utilization compared to conventional posit-based implementations. The proposed 32 32 systolic array achieves a 3% improvement in inference accuracy compared to the floating-point implementation. These findings indicate that the combination of posit arithmetic with targeted approximation techniques provides a promising solution for achieving accurate and energy-efficient deep learning inference at the edge.
Emadodin Sakhaee, Mahdi Kalbasi, Hooman Nikmehr, Naser Movahhedinia, John L. Gustafson
Neurocomputing5
2024 Algorithm-Hardware Co-Design of Distribution-Aware Logarithmic-Posit Encodings for Efficient DNN Inference
abstract
Traditional Deep Neural Network (DNN) quantization methods using integer or floating-point data types struggle to capture diverse DNN parameter distributions and often require large silicon overhead and intensive quantization-aware training. In this study, we introduce Logarithmic Posits (LP), an adaptive, hardware-friendly data type inspired by posits that dynamically adapts to DNN weight/activation distributions by parameterizing LP bit fields. We also develop a novel genetic-algorithm based framework, LP Quantization (LPQ), to find optimal layer-wise LP parameters while reducing representational divergence between quantized and full-precision models through a novel global-local contrastive objective. Additionally, we design a LP accelerator (LPA) architecture comprising of mixed-precision LP processing elements (PEs). Our algorithmhardware co-design demonstrates on average <1% drop in top-1 accuracy across various CNN and ViT models. It also achieves ~ 2× improvements in performance per unit area and 2.2× gains in energy efficiency compared to state-of-the-art quantization accelerators using different data types.
Akshat Ramachandran, Zishen Wan, Geonhwa Jeong, John L. Gustafson, Tushar Krishna
DAC4
2021 Posit Arithmetic for the Training and Deployment of Generative Adversarial Networks
abstract
This paper proposes a set of methods that enables low precision posit ™ arithmetic to be successfully used for the training of generative adversarial networks (GANs) with minimal quality loss. We show that ultra low precision posits, as small as 6 bits, can achieve high quality output for the generation phase after training. We also evaluate floating-point (float) formats and compare them to 8-bit posits in the context of GAN training. Our scaling and adaptive calibration techniques are capable of producing superior training quality for 8-bit posits that surpasses 8-bit floats and matches the results of 16-bit floats. Hardware simulation results indicate that our methods have higher energy efficiency compared to both 16- and 8-bit float training systems.
Nhut-Minh Ho, Duy Thanh Nguyen, Himeshi De Silva, John L. Gustafson, Weng-Fai Wong, Ik Joon Chang
DATE4
2021 An approach to generate correctly rounded math libraries for new floating point variants
abstract
Given the importance of floating point (FP) performance in numerous domains, several new variants of FP and its alternatives have been proposed (e.g., Bfloat16, TensorFloat32, and posits). These representations do not have correctly rounded math libraries. Further, the use of existing FP libraries for these new representations can produce incorrect results. This paper proposes a novel approach for generating polynomial approximations that can be used to implement correctly rounded math libraries. Existing methods generate polynomials that approximate the real value of an elementary function 𝑓 (𝑥) and produce wrong results due to approximation errors and rounding errors in the implementation. In contrast, our approach generates polynomials that approximate the correctly rounded value of 𝑓 (𝑥) (i.e., the value of 𝑓 (𝑥) rounded to the target representation). It provides more margin to identify efficient polynomials that produce correctly rounded results for all inputs. We frame the problem of generating efficient polynomials that produce correctly rounded results as a linear programming problem. Using our approach, we have developed correctly rounded, yet faster, implementations of elementary functions for multiple target representations.
Jay P. Lim, Mridul Aanjaneya, John L. Gustafson, Santosh Nagarakatte
Proc. ACM Program. Lang.3
2020 Next Generation Arithmetic for Edge Computing
abstract
Arithmetic is a key component and is ubiquitous in today’s digital world, ranging from embedded to high-performance computing systems. With machine learning at the fore in a wide range of application domains from wearables to automotive to avionics to weather prediction, sufficiently accurate yet low-cost arithmetic is the need for the day. Recently, there have been several advances in the domain of computer arithmetic, which includes high-precision anchored numbers from ARM, posit arithmetic, bfloat16, etc. as an alternative to IEEE 754-2008 compliant arithmetic. Optimizations on fixed-point and integer arithmetic are also being pursued actively for low-power computing architectures. Furthermore, approximate computing and transprecision/mixed-precision computing have been exciting areas of research forever. While academic research in the domain of computer arithmetic has a long history, industrial adoption of some of these new data types and techniques is in its early stages and expected to increase in the future. bfloat16 is an excellent example for this. In this paper, we bring academia and industry together to discuss the latest results and future directions for research in the domain of next-generation computer arithmetic, especially for edge computing.
Andre Guntoro, Cecilia De la Parra, Farhad Merchant, Florent de Dinechin, John L. Gustafson, Martin Langhammer, Rainer Leupers, Sangeeth Nambiar
DATE5
2019 Next-generation arithmetic: major performance gains with minimal disruption
abstract
Moore's law made application developers lazy, since they could rely on increases in clock speeds and transistor density to improve the performance of their codes with little or no rewriting required. The frontiers of supercomputing, such as quantum computing, are certainly exciting and promising, but also highly disruptive... even more so than the shift from serial to parallel computing. They require a complete rewrite of millions of lines of software, and the invention of completely different algorithms.
John L. Gustafson
CF1
2019 Deep Positron: A Deep Neural Network Using the Posit Number System
abstract
The recent surge of interest in Deep Neural Networks (DNNs) has led to increasingly complex networks that tax computational and memory resources. Many DNNs presently use 16-bit or 32-bit floating point operations. Significant performance and power gains can be obtained when DNN accelerators support low-precision numerical formats. Despite considerable research, there is still a knowledge gap on how low-precision operations can be realized for both DNN training and inference. In this work, we propose a DNN architecture, Deep Positron, with posit numerical format operating successfully at ≤8 bits for inference. We propose a precision-adaptable FPGA soft core for exact multiply-and-accumulate for uniform comparison across three numerical formats, fixed, floating-point and posit. Preliminary results demonstrate that 8-bit posit has better accuracy than 8-bit fixed or floating-point for three different low-dimensional datasets. Moreover, the accuracy is comparable to 32-bit floating-point on a Xilinx Virtex-7 FPGA device. The trade-offs between DNN performance and hardware resources, i.e. latency, power, and resource utilization, show that posit outperforms in accuracy and latency at 8-bit and below.
Zachariah Carmichael, Hamed Fatemi Langroudi, Char Khazanov, Jeffrey Lillie, John L. Gustafson, Dhireesha Kudithipudi
DATE5
2018 Making Strassen Matrix Multiplication Safe
abstract
Strassen's recursive algorithm for matrix-matrix multiplication has seen slow adoption in practical applications despite being asymptotically faster than the traditional algorithm. A primary cause for this is the comparatively weaker numerical stability of its results. Techniques that aim to improve the errors of Strassen stand the risk of losing any potential performance gain. Moreover, current methods of evaluating such techniques for safety are overly pessimistic or error prone and generally do not allow for quick and accurate comparisons. In this paper we present an efficient technique to obtain rigorous error bounds for floating point computations based on an implementation of unum arithmetic. Using it, we evaluate three techniques - exact dot product, fused multiply-add, and matrix quadrant rotation - that can potentially improve the numerical stability of Strassen's algorithm for practical use. We also propose a novel error-based heuristic rotation scheme for matrix quadrant rotation. Finally we apply techniques that improve numerical safety with low overhead to a LINPACK linear solver to demonstrate the usefulness of the Strassen algorithm in practice.
Himeshi De Silva, John L. Gustafson, Weng-Fai Wong
HiPC2
2018 Parameterized Posit Arithmetic Hardware Generator
abstract
Hardware implementation of Floating Point Units (FPUs) has been a key area of research due to their massive area and energy footprints. Recently, a proposal was made to replace IEEE 754-2008 technical standard compliant FPUs with Posit Arithmetic Units (PAUs) due to the greater accuracy, speed, and simpler hardware design. In this paper, we present the architecture of a parameterized PAU generator that can generate PAU adders and PAU multipliers of any bit-width pre-synthesis. We synthesize generated arithmetic units using the parameterized PAU generator for 8-bit, 16-bit, and 32-bit adders and multipliers and compare them with IEEE 754-2008 compliant adders and multipliers. Both, synthesis for Field Programmable Gate Array (FPGA) and Application Specific Integrated Circuit (ASIC) are performed. In our comparison of m-bit PAU units with n-bit IEEE 754-2008 compliant units, it is observed that the area and energy of a PAU adder and multiplier are comparable to their IEEE 754-2008 compliant counterparts where m=n. We argue that an n-bit IEEE 754-2008 adder and multiplier can be safely replaced with an m-bit PAU adder and multiplier where m
Rohit Chaurasiya, John L. Gustafson, Rahul Shrestha, Jonathan Neudorfer, Sangeeth Nambiar, Kaustav Niyogi, Farhad Merchant, Rainer Leupers
ICCD2
2010 Challenges and Future Directions of Software Technology: The Need for Explicit Programming Environments
abstract
Discussion of the future software increasingly requires a careful distinction between application-facing software and hardware-facing software. Programmers of application-facing software will increasingly have to balance speed, reliability, and accuracy as competing goals. Programmers of hardware-facing software will increasingly have to manage data placement, power consumption, and the choices presented by heterogeneous processors. By making these tradeoffs explicit for both programming environments, we will be able to overcome these challenges and potentially will discover new approaches that are not possible with presently available tools.
John L. Gustafson
COMPSAC1
2006 Innovative technologies II - Acceleration technologies: understanding the differences and assessing what's right for you
abstract
Acceleration technologies have been gaining much attention in the past year and increasingly being integrated in cluster solutions ranging from top500 class installations to production oriented systems such as IBM's System Cluster 1350. Options include games processors such as the Cell processor, Graphics Processing Units (GPUs,) Field Programmable Gate Array's (FPGAs,) and purpose designed coprocessors such as ClearSpeed's technology.So what are the differences between the acceleration technology categories, and how do you know if one of the options is right for you?From this presentation, you will acquire an understanding of how to assess the range of characteristics of accelerators which include single and double precision performance, power consumption, programming models, and how they relate to your application needs.
John L. Gustafson
SC1
2004 The Speed of Light Isn't What it Used to Be
abstract
As computer designs run up against the limits of physics, we increasingly see communication latencies limited by the speed of light. However, the "speed of light" is far less than the usual rule of thumb, "one nanosecond per foot of distance"; on-chip traces will soon be able to move signals at only 4% of the speed of light. This accelerates the need for ways to cope with an ever-increasing ratio of communication to computation time, cost, and power consumption in modern computer designs. We examine the impact this has on designers and users, and present some novel approaches for dealing with this effect both in hardware and software.
John L. Gustafson
ICPADS1
1999 Conventional Benchmarks as a Sample of the Performance Spectrum
John L. Gustafson, Rajat Todi
J. Supercomput.1
1998 Distribution-Independent Hierarchical Algorithms for the N-body Problem
Srinivas Aluru, John L. Gustafson, Gurpur M. Prabhu, Fatih Erdogan Sevilgen
J. Supercomput.2
1997 Parallel Hierarchical Global Illumination
abstract
This paper presents an algorithm that solves the Rendering Equation to any desired accuracy, and can be run in parallel on distributed memory or shared memory computer systems with excellent scaling properties. It appears superior in both speed and physical correctness to recent published methods involving bidirectional ray tracing or hybrid treatments of diffuse and specular surfaces. Like "progressive radiosity" methods, it dynamically refines the geometry decomposition where required, but does so without the excessive storage requirements for "ray histories.".
Quinn Snell, John L. Gustafson
HPDC2
1996 An Analytical Model of the HINT Performance Metric
abstract
The HINT Benchmark was developed to provide a broad-spectrum metric for computers, and to measure performance over the full range of memory sizes and time scales. We have extended our understanding of why HINT performance curves look the way they do, and can now predict the curves using an analytical model based on simple hardware specifications as input parameters.
Quinn Snell, John L. Gustafson
SC2
1994 Truly distribution-independent algorithms for the N-body problem
abstract
The N-body problem is to simulate the motion of N particles under the influence of mutual force fields based on an inverse square law. Greengard's algorithm claims to compute the cumulative force on each particle in O(N) time for a fixed precision irrespective of the distribution of the particles. In this paper, we show that Greengard's algorithm is distribution dependent and has a lower bound of /spl Omega/(N log/sup 2/ N) in two dimensions and /spl Omega/(N log/sup 4/ N) in three dimensions. We analyze the Greengard and Barnes-Hut algorithms and show that they are unbounded for arbitrary distributions. We also present a truly distribution independent algorithm for solving the N-body problem in O(N log N) time in two dimensions and in O(N log/sup 2/ N) time in three dimensions.>
Srinivas Aluru, Gurpur M. Prabhu, John L. Gustafson
SC3
1993 A Massively Parallel Optimizer for Expression Evaluation
abstract
A number of “tricks” are known that trade multiplications for additions. The term “tricks” reflects the way these methods seem not to proceed from any general theory, but instead jump into existence as recipes that work. The Strassen method for 2 by 2 matrix product with 7 multiplications is a well-known example, as is the method for finding a complex number product in 3 multiplications. We have created a practical computer program for finding such tricks automatically, where massive parallelism makes the combinatorially explosive search tolerable for small problems. One result of this program is a method for computing cross products of 3-vectors using only 5 multiplications.
Srinivas Aluru, John L. Gustafson
International Conference on Supercomputing2
1993 Commentary - The "Tar Baby" of Computing: Performance Analysis
abstract
It looks so easy. To find the speed of a computer, just take a job and time it. To compare two speeds, or try to establish speed in some type of universal unit, well…. Reporting performance is the “tar baby” of computing, and many authors and researchers have attacked this issue only to find themselves stuck in a mess of shaky logic and shakier assumptions. The more scientific one tries to be, the more difficult the problem is. Barr and Hickman have aimed their efforts at this issue, and have generated some thought-provoking ideas. INFORMS Journal on Computing, ISSN 1091-9856, was published as ORSA Journal on Computing from 1989 to 1995 under ISSN 0899-1499.
John L. Gustafson
INFORMS J. Comput.1
1992 A random number generator for parallel computers
Srinivas Aluru, Gurpur M. Prabhu, John L. Gustafson
Parallel Comput.3
1991 SIZEUP: A New Parallel Performance Metric
Xian-He Sun, John L. Gustafson
ICPP (2)2
1991 A Threshold Test for Dynamic Load Balancers
Milton C. Wikstrom, John L. Gustafson, Gurpur M. Prabhu
ICPP (2)2
1991 The Design of a Scalable, Fixed-Time Computer Benchmark
John L. Gustafson, Diane T. Rover, Stephen T. Elbert, Michael Carter
J. Parallel Distributed Comput.1
1991 Signal-Processing Algorithms on Parallel Architectures: A Performance Update
Diane T. Rover, Vicki Tsai, Yin-Shan Chow, John L. Gustafson
J. Parallel Distributed Comput.4
1991 Toward a better parallel performance metric
Xian-He Sun, John L. Gustafson
Parallel Comput.2
1989 A radar simulation program for a 1024-processor hypercube
abstract
We have developed a fast parallel version of an existing synthetic aperture radar (SAR) simulation program, SRIM. On a 1024-processor NCUBE hypercube it runs an order of magnitude faster than on a CRAY X-MP or CRAY Y-MP processor. This speed advantage is coupled with an order of magnitude advantage in machine acquisition cost. SRIM is a somewhat large (30,000 lines of Fortran 77) program designed for uniprocessors; its restructuring for a hypercube provides new lessons in the task of altering older serial programs to run well on modern parallel architectures. We describe the techniques used for parallelization, and the performance obtained. Several novel parallel approaches to problems of task distribution, data distribution, and direct output were required. These techniques increase performance and appear to have general applicability for massive parallelism. We describe the hierarchy necessary to dynamically manage (i.e., load balance) a large ensemble. The ensemble is used in a heterogeneous manner, with different programs on different parts of the hypercube. The heterogeneous approach takes advantage of the independent instruction streams possible on MIMD machines.
John L. Gustafson, Robert E. Benner, Mark P. Sears, Thomas D. Sullivan
SC1
1986 The Architecture of a Homogeneous Vector Supercomputer
John L. Gustafson, Stuart Hawkinson, Ken Scott
ICPP1
1986 The Architecture of a Homogeneous Vector Supercomputer
John L. Gustafson, Stuart Hawkinson, Ken Scott
J. Parallel Distributed Comput.1