VLDB 2026 Research / reviewers in the wild / expert
David Raymond Lutz
dblp:63/5918
· DBLP profile ↗
9ranked-venue papers
8as first author
2since 2021 · last 2026
0000-0002-7921-5455ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Theory of computation · 5 · 5 first-author · 1 since 2021Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fused FP8 Many-Terms Dot Product With Scaling and FP32 Accumulation
David Raymond Lutz, Anisha Saini, Mairin Kroes, Thomas Elmer, Harsha Valsaraju, Javier D. Bruguera |
IEEE Trans. Computers | 1 |
| 2024 | Fused FP8 4-Way Dot Product With Scaling and FP32 AccumulationabstractFor a variety of ML applications, generalized matrix multiply (GEMM) with DOT product is the most computationally intensive operation. This paper presents a microarchitecture exploration of fused multi-way reduced precision floating point multiply-accumulate with single rounding, resulting in power and area efficient characteristics. We propose two different microarchitectures, implementing novel design techniques for computing fused FP8 DOT4 accumulating to higher precision FP32 with scaling to adjust the dynamic range. Our first design, dot product with late accumulation, computes fused FP8 DOT4 by calculating the dot product in the first two cycles, expanding the products to fixed point format and another two cycles for the accumulation operation. This design allows the reuse of a slightly modified, FMA capable, FP32 adder. Our second design, dot product with early accumulation is implemented as a standalone FP8 datapath computing products and accumulation in first two cycles, and another two cycles for normalization and single rounding operation. This design aligns addends (products and accumulator) from an “Anchor” for efficient, arithmetically fused, N-way FP DOT product computation. Furthermore, we synthesized the two designs proposed in a 5nm technology node and compared the cost of implementation. David Raymond Lutz, Anisha Saini, Mairin Kroes, Thomas Elmer, Harsha Valsaraju |
ARITH | 1 |
| 2019 | ARM Floating Point 2019: Latency, Area, PowerabstractWe have had little or no speed increase from process in the past few years, but latency continues to decrease due to algorithmic improvements [1] and a decision to spend more area on CPU datapaths [2]. A binary64 floating-point (FP) add now takes two cycles when done as part of a 2+2-cycle FMA, and even one cycle when done as part of an in-order vector reduction. Smaller and more specialized FP operations (bfloat16) are even faster. Finally, the decision to spend more area on datapath logic took a new twist this year when we applied it to GPUs, cutting dynamic power there by a third. David Raymond Lutz |
ARITH | 1 |
| 2019 | High-Precision Anchored Accumulators for Reproducible Floating-Point SummationabstractThis paper introduces a new datatype, the High-Precision Anchored (HPA) number, that allows reproducible accumulation of floating-point (FP) numbers in a programmer-selectable range. The new datatype has a larger significand and a smaller range than existing FP formats and has much better arithmetic and computational properties. In particular, it is associative, parallelizable, reproducible and correct. The paper also describes how HPA processing can be implemented as part of Arm's new Scalable Vector Extension (SVE) together with proposals for new instructions aimed specifically at the new datatype. For the modest ranges that will accommodate most problems, HPA processing is much faster than FP arithmetic: performance modelling shows 2-lane HPA accumulation of FP64 operands is 9.5 times faster on Arm's new vector architecture than double double accumulation and accelerates a recently published software algorithm for 3-lane reproducible FP summation by a factor of 5.6. This paper also discusses instruction-level optimizations for FP32 and FP16 summations that further increase HPA performance relative to FP64 accumulations. Neil Burgess, Christopher E. Goodyer, Christopher Neal Hinds, David Raymond Lutz |
IEEE Trans. Computers | 4 |
| 2017 | High-Precision Anchored Accumulators for Reproducible Floating-Point SummationabstractThis paper introduces a new datatype that allows reproducible accumulation of floating-point (FP) numbers in a programmer-selectable range. The new datatype has a larger significand and a smaller range than existing FP formats and has much better arithmetic and computational properties. In particular, it is associative, parallelizable, reproducible and correct. For the modest ranges that will accommodate most problems, it is also much faster: 3 to 12 times faster on a single 256-bit SIMD implementation. The paper also describes a new instruction and associated datapath that support the proposed datatype, and discusses how a recently published software algorithm for reproducible FP summation could be implemented using the proposed approach. David Raymond Lutz, Christopher Neal Hinds |
ARITH | 1 |
| 2011 | Fused Multiply-Add Microarchitecture Comprising Separate Early-Normalizing Multiply and Add PipelinesabstractWe present an IEEE 754-2008 and ARM compliant floating-point micro architecture that preserves the higher performance of separate multiply and add units while decreasing the effective latency of fused multiply-adds (FMAs). The multiplier supports subnormals in a novel and faster manner, shifting the partial products so that injection rounding can be used. The early-normalizing adder retains the low latency of a split path near/far adder, but does so in a unified path with less area. The adder also allows rounding on effective subtractions involving one input that is twice the normal width, a necessary feature for handling FMAs. The resulting floating-point unit has about twice the (IPC) performance of the best previous ARM design, and can be clocked at a higher speed despite the wider paths required by FMAs. David Raymond Lutz |
IEEE Symposium on Computer Arithmetic | 1 |
| 1999 | Performance analysis for chipsets and systemsabstractPlasma is a new tool for modeling the timing of chipsets and other system components. Modeling chipsets is in some ways more difficult than modeling processors: the interfaces are more complex and more numerous, the internal queues and buffers are larger, and the traces are much more complicated. We have used Plasma to create a timing model for a modern chipset. The resulting model is fast, flexible, and useful for both design and verification. David Raymond Lutz, B. Kahne |
IPCCC | 1 |
| 1997 | The Half-Adder Form and Early Branch Condition ResolutionabstractThe authors present efficient methods to determine the four usual branch conditions for a sum or difference, before the result of the addition or subtraction is available. The methods lead to the design of an early branch resolver which integrates well with a regular adder/subtracter, adding only a small amount of circuitry and almost no delay. The methods exploit the properties of half-adder form. Sums in half-adder form can be computed vary quickly (with the delay of a half adder), yet they have enough structure so that many of the properties of the final sum can be easily detected. The reduced latency for evaluating branch conditions means that an addition or subtraction and a dependent conditional instruction can execute in the same cycle with a consequent increase in instruction-level parallelism, and improved performance for both single-issue and superscalar processors. David Raymond Lutz, Doddaballapur Narasimha-Murthy Jayasimha |
IEEE Symposium on Computer Arithmetic | 1 |
| 1996 | Early Zero DetectionabstractWe present an integrated adder/subtracter/zero-detector in which the zero detection completes well before the sum or difference is known. Previous zero detectors either required the sum to be available before they could complete, or were not well integrated with the ALU. We avoid these problems by exploiting the properties of half-adder form. Sums in half-adder form can be computed very quickly (with the delay of a half adder), yet they have enough structure so that many of the properties of the final sum can be easily detected. Our zero detector is faster than any previously described, requires only a small amount of additional circuitry in the ALU, and adds little or nothing to the overall delay of the ALU. We also examine some of the architectural implications of early zero detection: faster branching, more instruction-level parallelism, more powerful instructions, and reduced hardware needs for supporting speculative execution. David Raymond Lutz, Doddaballapur Narasimha-Murthy Jayasimha |
ICCD | 1 |