VLDB 2026 Research / reviewers in the wild / expert
John L. Gustafson
dblp:44/4606
· DBLP profile ↗
29ranked-venue papers
10as first author
5since 2021 · last 2026
0000-0002-2957-1304ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 8 first-author · 2 since 2021Software engineering, systems software and programming languages · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Early Electronic Digital Computers - Panel Summary
B. Jack Copeland, Vladimir Getov, John L. Gustafson, Brian Stuart |
COMPSAC | 3 |
| 2026 | Energy-efficient posit systolic array for real-time CNN inference using approximate computing on FPGAabstractWith the increasing adoption of Convolutional Neural Networks (CNNs) in edge and embedded applications, there is a growing demand for hardware accelerators that provide high computational accuracy while maintaining low energy consumption. In this work, we present a reconfigurable CNN accelerator specifically designed for real-time object detection. The accelerator is based on a systolic array architecture and integrates the posit number system to enhance both accuracy and dynamic range in numerical representation. To further improve energy efficiency, the design incorporates a set of lightweight approximation techniques applied to arithmetic operations. In particular, the Multiply and Accumulate (MAC) units employ approximate strategies for multiplication, normalization, and rounding of the fraction part, selectively simplifying low-significance computations to reduce hardware cost, while ensuring that the final classification accuracy remains within acceptable limits. The proposed architecture is implemented at the Register-Transfer Level (RTL) level and evaluated on a Xilinx ZCU102 FPGA using ResNet-50 and the ImageNet dataset. Experimental results demonstrate a 21% reduction in power consumption and an 18% decrease in resource utilization compared to conventional posit-based implementations. The proposed 32 32 systolic array achieves a 3% improvement in inference accuracy compared to the floating-point implementation. These findings indicate that the combination of posit arithmetic with targeted approximation techniques provides a promising solution for achieving accurate and energy-efficient deep learning inference at the edge. Emadodin Sakhaee, Mahdi Kalbasi, Hooman Nikmehr, Naser Movahhedinia, John L. Gustafson |
Neurocomputing | 5 |
| 2024 | Algorithm-Hardware Co-Design of Distribution-Aware Logarithmic-Posit Encodings for Efficient DNN InferenceabstractTraditional Deep Neural Network (DNN) quantization methods using integer or floating-point data types struggle to capture diverse DNN parameter distributions and often require large silicon overhead and intensive quantization-aware training. In this study, we introduce Logarithmic Posits (LP), an adaptive, hardware-friendly data type inspired by posits that dynamically adapts to DNN weight/activation distributions by parameterizing LP bit fields. We also develop a novel genetic-algorithm based framework, LP Quantization (LPQ), to find optimal layer-wise LP parameters while reducing representational divergence between quantized and full-precision models through a novel global-local contrastive objective. Additionally, we design a LP accelerator (LPA) architecture comprising of mixed-precision LP processing elements (PEs). Our algorithmhardware co-design demonstrates on average <1% drop in top-1 accuracy across various CNN and ViT models. It also achieves ~ 2× improvements in performance per unit area and 2.2× gains in energy efficiency compared to state-of-the-art quantization accelerators using different data types. Akshat Ramachandran, Zishen Wan, Geonhwa Jeong, John L. Gustafson, Tushar Krishna |
DAC | 4 |
| 2021 | Posit Arithmetic for the Training and Deployment of Generative Adversarial NetworksabstractThis paper proposes a set of methods that enables low precision posit ™ arithmetic to be successfully used for the training of generative adversarial networks (GANs) with minimal quality loss. We show that ultra low precision posits, as small as 6 bits, can achieve high quality output for the generation phase after training. We also evaluate floating-point (float) formats and compare them to 8-bit posits in the context of GAN training. Our scaling and adaptive calibration techniques are capable of producing superior training quality for 8-bit posits that surpasses 8-bit floats and matches the results of 16-bit floats. Hardware simulation results indicate that our methods have higher energy efficiency compared to both 16- and 8-bit float training systems. Nhut-Minh Ho, Duy Thanh Nguyen, Himeshi De Silva, John L. Gustafson, Weng-Fai Wong, Ik Joon Chang |
DATE | 4 |
| 2021 | An approach to generate correctly rounded math libraries for new floating point variantsabstractGiven the importance of floating point (FP) performance in numerous domains, several new variants of FP and its alternatives have been proposed (e.g., Bfloat16, TensorFloat32, and posits). These representations do not have correctly rounded math libraries. Further, the use of existing FP libraries for these new representations can produce incorrect results. This paper proposes a novel approach for generating polynomial approximations that can be used to implement correctly rounded math libraries. Existing methods generate polynomials that approximate the real value of an elementary function 𝑓 (𝑥) and produce wrong results due to approximation errors and rounding errors in the implementation. In contrast, our approach generates polynomials that approximate the correctly rounded value of 𝑓 (𝑥) (i.e., the value of 𝑓 (𝑥) rounded to the target representation). It provides more margin to identify efficient polynomials that produce correctly rounded results for all inputs. We frame the problem of generating efficient polynomials that produce correctly rounded results as a linear programming problem. Using our approach, we have developed correctly rounded, yet faster, implementations of elementary functions for multiple target representations. Jay P. Lim, Mridul Aanjaneya, John L. Gustafson, Santosh Nagarakatte |
Proc. ACM Program. Lang. | 3 |
| 2020 | Next Generation Arithmetic for Edge ComputingabstractArithmetic is a key component and is ubiquitous in today’s digital world, ranging from embedded to high-performance computing systems. With machine learning at the fore in a wide range of application domains from wearables to automotive to avionics to weather prediction, sufficiently accurate yet low-cost arithmetic is the need for the day. Recently, there have been several advances in the domain of computer arithmetic, which includes high-precision anchored numbers from ARM, posit arithmetic, bfloat16, etc. as an alternative to IEEE 754-2008 compliant arithmetic. Optimizations on fixed-point and integer arithmetic are also being pursued actively for low-power computing architectures. Furthermore, approximate computing and transprecision/mixed-precision computing have been exciting areas of research forever. While academic research in the domain of computer arithmetic has a long history, industrial adoption of some of these new data types and techniques is in its early stages and expected to increase in the future. bfloat16 is an excellent example for this. In this paper, we bring academia and industry together to discuss the latest results and future directions for research in the domain of next-generation computer arithmetic, especially for edge computing. Andre Guntoro, Cecilia De la Parra, Farhad Merchant, Florent de Dinechin, John L. Gustafson, Martin Langhammer, Rainer Leupers, Sangeeth Nambiar |
DATE | 5 |
| 2019 | Next-generation arithmetic: major performance gains with minimal disruptionabstractMoore's law made application developers lazy, since they could rely on increases in clock speeds and transistor density to improve the performance of their codes with little or no rewriting required. The frontiers of supercomputing, such as quantum computing, are certainly exciting and promising, but also highly disruptive... even more so than the shift from serial to parallel computing. They require a complete rewrite of millions of lines of software, and the invention of completely different algorithms. John L. Gustafson |
CF | 1 |
| 2019 | Deep Positron: A Deep Neural Network Using the Posit Number SystemabstractThe recent surge of interest in Deep Neural Networks (DNNs) has led to increasingly complex networks that tax computational and memory resources. Many DNNs presently use 16-bit or 32-bit floating point operations. Significant performance and power gains can be obtained when DNN accelerators support low-precision numerical formats. Despite considerable research, there is still a knowledge gap on how low-precision operations can be realized for both DNN training and inference. In this work, we propose a DNN architecture, Deep Positron, with posit numerical format operating successfully at ≤8 bits for inference. We propose a precision-adaptable FPGA soft core for exact multiply-and-accumulate for uniform comparison across three numerical formats, fixed, floating-point and posit. Preliminary results demonstrate that 8-bit posit has better accuracy than 8-bit fixed or floating-point for three different low-dimensional datasets. Moreover, the accuracy is comparable to 32-bit floating-point on a Xilinx Virtex-7 FPGA device. The trade-offs between DNN performance and hardware resources, i.e. latency, power, and resource utilization, show that posit outperforms in accuracy and latency at 8-bit and below. Zachariah Carmichael, Hamed Fatemi Langroudi, Char Khazanov, Jeffrey Lillie, John L. Gustafson, Dhireesha Kudithipudi |
DATE | 5 |
| 2018 | Making Strassen Matrix Multiplication SafeabstractStrassen's recursive algorithm for matrix-matrix multiplication has seen slow adoption in practical applications despite being asymptotically faster than the traditional algorithm. A primary cause for this is the comparatively weaker numerical stability of its results. Techniques that aim to improve the errors of Strassen stand the risk of losing any potential performance gain. Moreover, current methods of evaluating such techniques for safety are overly pessimistic or error prone and generally do not allow for quick and accurate comparisons. In this paper we present an efficient technique to obtain rigorous error bounds for floating point computations based on an implementation of unum arithmetic. Using it, we evaluate three techniques - exact dot product, fused multiply-add, and matrix quadrant rotation - that can potentially improve the numerical stability of Strassen's algorithm for practical use. We also propose a novel error-based heuristic rotation scheme for matrix quadrant rotation. Finally we apply techniques that improve numerical safety with low overhead to a LINPACK linear solver to demonstrate the usefulness of the Strassen algorithm in practice. Himeshi De Silva, John L. Gustafson, Weng-Fai Wong |
HiPC | 2 |
| 2018 | Parameterized Posit Arithmetic Hardware GeneratorabstractHardware implementation of Floating Point Units (FPUs) has been a key area of research due to their massive area and energy footprints. Recently, a proposal was made to replace IEEE 754-2008 technical standard compliant FPUs with Posit Arithmetic Units (PAUs) due to the greater accuracy, speed, and simpler hardware design. In this paper, we present the architecture of a parameterized PAU generator that can generate PAU adders and PAU multipliers of any bit-width pre-synthesis. We synthesize generated arithmetic units using the parameterized PAU generator for 8-bit, 16-bit, and 32-bit adders and multipliers and compare them with IEEE 754-2008 compliant adders and multipliers. Both, synthesis for Field Programmable Gate Array (FPGA) and Application Specific Integrated Circuit (ASIC) are performed. In our comparison of m-bit PAU units with n-bit IEEE 754-2008 compliant units, it is observed that the area and energy of a PAU adder and multiplier are comparable to their IEEE 754-2008 compliant counterparts where m=n. We argue that an n-bit IEEE 754-2008 adder and multiplier can be safely replaced with an m-bit PAU adder and multiplier where m Rohit Chaurasiya, John L. Gustafson, Rahul Shrestha, Jonathan Neudorfer, Sangeeth Nambiar, Kaustav Niyogi, Farhad Merchant, Rainer Leupers |
ICCD | 2 |
| 2010 | Challenges and Future Directions of Software Technology: The Need for Explicit Programming EnvironmentsabstractDiscussion of the future software increasingly requires a careful distinction between application-facing software and hardware-facing software. Programmers of application-facing software will increasingly have to balance speed, reliability, and accuracy as competing goals. Programmers of hardware-facing software will increasingly have to manage data placement, power consumption, and the choices presented by heterogeneous processors. By making these tradeoffs explicit for both programming environments, we will be able to overcome these challenges and potentially will discover new approaches that are not possible with presently available tools. John L. Gustafson |
COMPSAC | 1 |
| 2006 | Innovative technologies II - Acceleration technologies: understanding the differences and assessing what's right for youabstractAcceleration technologies have been gaining much attention in the past year and increasingly being integrated in cluster solutions ranging from top500 class installations to production oriented systems such as IBM's System Cluster 1350. Options include games processors such as the Cell processor, Graphics Processing Units (GPUs,) Field Programmable Gate Array's (FPGAs,) and purpose designed coprocessors such as ClearSpeed's technology.So what are the differences between the acceleration technology categories, and how do you know if one of the options is right for you?From this presentation, you will acquire an understanding of how to assess the range of characteristics of accelerators which include single and double precision performance, power consumption, programming models, and how they relate to your application needs. John L. Gustafson |
SC | 1 |
| 2004 | The Speed of Light Isn't What it Used to BeabstractAs computer designs run up against the limits of physics, we increasingly see communication latencies limited by the speed of light. However, the "speed of light" is far less than the usual rule of thumb, "one nanosecond per foot of distance"; on-chip traces will soon be able to move signals at only 4% of the speed of light. This accelerates the need for ways to cope with an ever-increasing ratio of communication to computation time, cost, and power consumption in modern computer designs. We examine the impact this has on designers and users, and present some novel approaches for dealing with this effect both in hardware and software. John L. Gustafson |
ICPADS | 1 |
| 1999 | Conventional Benchmarks as a Sample of the Performance Spectrum
John L. Gustafson, Rajat Todi |
J. Supercomput. | 1 |
| 1998 | Distribution-Independent Hierarchical Algorithms for the N-body Problem
Srinivas Aluru, John L. Gustafson, Gurpur M. Prabhu, Fatih Erdogan Sevilgen |
J. Supercomput. | 2 |
| 1997 | Parallel Hierarchical Global IlluminationabstractThis paper presents an algorithm that solves the Rendering Equation to any desired accuracy, and can be run in parallel on distributed memory or shared memory computer systems with excellent scaling properties. It appears superior in both speed and physical correctness to recent published methods involving bidirectional ray tracing or hybrid treatments of diffuse and specular surfaces. Like "progressive radiosity" methods, it dynamically refines the geometry decomposition where required, but does so without the excessive storage requirements for "ray histories.". Quinn Snell, John L. Gustafson |
HPDC | 2 |
| 1996 | An Analytical Model of the HINT Performance MetricabstractThe HINT Benchmark was developed to provide a broad-spectrum metric for computers, and to measure performance over the full range of memory sizes and time scales. We have extended our understanding of why HINT performance curves look the way they do, and can now predict the curves using an analytical model based on simple hardware specifications as input parameters. Quinn Snell, John L. Gustafson |
SC | 2 |
| 1994 | Truly distribution-independent algorithms for the N-body problemabstractThe N-body problem is to simulate the motion of N particles under the influence of mutual force fields based on an inverse square law. Greengard's algorithm claims to compute the cumulative force on each particle in O(N) time for a fixed precision irrespective of the distribution of the particles. In this paper, we show that Greengard's algorithm is distribution dependent and has a lower bound of /spl Omega/(N log/sup 2/ N) in two dimensions and /spl Omega/(N log/sup 4/ N) in three dimensions. We analyze the Greengard and Barnes-Hut algorithms and show that they are unbounded for arbitrary distributions. We also present a truly distribution independent algorithm for solving the N-body problem in O(N log N) time in two dimensions and in O(N log/sup 2/ N) time in three dimensions.> Srinivas Aluru, Gurpur M. Prabhu, John L. Gustafson |
SC | 3 |
| 1993 | A Massively Parallel Optimizer for Expression EvaluationabstractA number of “tricks” are known that trade multiplications for additions. The term “tricks” reflects the way these methods seem not to proceed from any general theory, but instead jump into existence as recipes that work. The Strassen method for 2 by 2 matrix product with 7 multiplications is a well-known example, as is the method for finding a complex number product in 3 multiplications. We have created a practical computer program for finding such tricks automatically, where massive parallelism makes the combinatorially explosive search tolerable for small problems. One result of this program is a method for computing cross products of 3-vectors using only 5 multiplications. Srinivas Aluru, John L. Gustafson |
International Conference on Supercomputing | 2 |
| 1993 | Commentary - The "Tar Baby" of Computing: Performance AnalysisabstractIt looks so easy. To find the speed of a computer, just take a job and time it. To compare two speeds, or try to establish speed in some type of universal unit, well…. Reporting performance is the “tar baby” of computing, and many authors and researchers have attacked this issue only to find themselves stuck in a mess of shaky logic and shakier assumptions. The more scientific one tries to be, the more difficult the problem is. Barr and Hickman have aimed their efforts at this issue, and have generated some thought-provoking ideas. INFORMS Journal on Computing, ISSN 1091-9856, was published as ORSA Journal on Computing from 1989 to 1995 under ISSN 0899-1499. John L. Gustafson |
INFORMS J. Comput. | 1 |
| 1992 | A random number generator for parallel computers
Srinivas Aluru, Gurpur M. Prabhu, John L. Gustafson |
Parallel Comput. | 3 |
| 1991 | SIZEUP: A New Parallel Performance Metric
Xian-He Sun, John L. Gustafson |
ICPP (2) | 2 |
| 1991 | A Threshold Test for Dynamic Load Balancers
Milton C. Wikstrom, John L. Gustafson, Gurpur M. Prabhu |
ICPP (2) | 2 |
| 1991 | The Design of a Scalable, Fixed-Time Computer Benchmark
John L. Gustafson, Diane T. Rover, Stephen T. Elbert, Michael Carter |
J. Parallel Distributed Comput. | 1 |
| 1991 | Signal-Processing Algorithms on Parallel Architectures: A Performance Update
Diane T. Rover, Vicki Tsai, Yin-Shan Chow, John L. Gustafson |
J. Parallel Distributed Comput. | 4 |
| 1991 | Toward a better parallel performance metric
Xian-He Sun, John L. Gustafson |
Parallel Comput. | 2 |
| 1989 | A radar simulation program for a 1024-processor hypercubeabstractWe have developed a fast parallel version of an existing synthetic aperture radar (SAR) simulation program, SRIM. On a 1024-processor NCUBE hypercube it runs an order of magnitude faster than on a CRAY X-MP or CRAY Y-MP processor. This speed advantage is coupled with an order of magnitude advantage in machine acquisition cost. SRIM is a somewhat large (30,000 lines of Fortran 77) program designed for uniprocessors; its restructuring for a hypercube provides new lessons in the task of altering older serial programs to run well on modern parallel architectures. We describe the techniques used for parallelization, and the performance obtained. Several novel parallel approaches to problems of task distribution, data distribution, and direct output were required. These techniques increase performance and appear to have general applicability for massive parallelism. We describe the hierarchy necessary to dynamically manage (i.e., load balance) a large ensemble. The ensemble is used in a heterogeneous manner, with different programs on different parts of the hypercube. The heterogeneous approach takes advantage of the independent instruction streams possible on MIMD machines. John L. Gustafson, Robert E. Benner, Mark P. Sears, Thomas D. Sullivan |
SC | 1 |
| 1986 | The Architecture of a Homogeneous Vector Supercomputer
John L. Gustafson, Stuart Hawkinson, Ken Scott |
ICPP | 1 |
| 1986 | The Architecture of a Homogeneous Vector Supercomputer
John L. Gustafson, Stuart Hawkinson, Ken Scott |
J. Parallel Distributed Comput. | 1 |