VLDB 2026 Research / reviewers in the wild / expert
Florent de Dinechin
dblp:74/1204 · also Florent Dupont de Dinechin
· DBLP profile ↗
57ranked-venue papers
25as first author
11since 2021 · last 2025
0000-0003-4927-3301ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 40 · 17 first-author · 6 since 2021Theory of computation · 17 · 8 first-author · 5 since 2021Software engineering, systems software and programming languages · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hardware Fixed-Point 2D and 3D normsabstractThis article studies the hardware implementation of the Euclidean norm in 2 and 3 dimensions with fixed-point inputs and outputs. It compares the CORDIC shift-and-add algorithm to a “naive” architecture combining squarers, adders and square root, with a common specification: faithful accuracy. This specification is used in both cases to determine bounds on architectural parameters such as the number of CORDIC iterations and the bit-width of internal data-paths. Several architectural variants of the “naive” architecture are investigated. Their relevance domains are discussed based on synthesis results on FPGAs. For two dimensions, CORDIC has lower area but longer latency than a well-researched naive version. 3D variants of CORDIC, however, are worse than the naive architecture both in area and delay. Romain Bouarah, Florent de Dinechin |
ARITH | 2 |
| 2025 | Double-Word Decomposition in a Combined FP16, BF16 and FP32 Dot Product Add OperatorabstractThis work presents a floating-point Fused Dot Prod-uct Add operator designed for two use cases: mixed-precision matrix multiply add for deep learning and single-precision arithmetic for numerical computing. We build on classic double- word techniques to emulate FP32 arithmetic as the unevalu-ated sum of lower-precision floating-point numbers. Specifically, we introduce E9S12, a 9-bit exponent, 12-bit significant, non- normalized floating-point format, to represent either the FP16 and BF16 multiplicands or the results of the FP32 multiplicand decomposition inside a Fused Dot Product Add operator. The re-sulting operator correctly rounds the FP16 products accumulated in FP32 and also matches the standard FP32 FMA arithmetic when the size-4 E9S12 dot product operates on decomposed FP32 multiplicands. We evaluated the implementation of this combined FP16, BF16 and FP32 Dot Product Add operator and compare it to correctly rounded Fused Dot Product Add operators by synthesis for a$4\text{nm}$technolgy node. Orégane Desrentes, Benoît Dupont de Dinechin, Florent de Dinechin |
ARITH | 3 |
| 2025 | RETROSPECTIVE: Table-Based Polynomials for Fast Hardware Function EvaluationabstractThe Higher Order Table-Based Method (HOTBM) was published [1] during Jérémie Detrey's PhD. It was at the time one of the most advanced generic methods for fixed-point hardware function evaluation. This article reviews its legacy, some of the subsequent progress in this field, and the challenges that remain. Florent de Dinechin |
ASAP | 1 |
| 2025 | Towards optimal reconfigurable constant multipliersabstractThis paper introduces a novel algorithm for generating run-time reconfigurable single constant multipliers (RSCMs) which are optimal within their model in terms of hardware cost. Optimality is ensured by an exhaustive exploration of the design space mixing constraint programming, depth-first search, and branch-and-prune techniques. The cost model of previous works is also refined. Compared to the state of the art, this approach enables much larger constant sets, and also significantly improves the performance of the resulting architectures. Applications to neural network inference and small floating-point multiplication units are evaluated. Bastien Barbe, Anastasia Volkova 0001, Florent de Dinechin |
DSD | 4 |
| 2023 | Exact Fused Dot Product Add OperatorsabstractThis article explores architectures of exact (correctly rounded) fused dot product and add operators suitable for the FP32 and FP64 binary floating-point representations with subnormal support, and other representations with a wide dynamic range such as bfloat16. The exact summation of terms before rounding requires a full-size accumulator, and this work discusses techniques to compress the identical bits of this accumulator. This requires the computation of the relative shift amounts of the terms, which is formulated as a parallel prefix algorithm, allowing for a low-latency implementation. Architectural options for the exact fused dot product and add operators with up to 16 products for FP32, FP64 and mixed-precision BF16 to FP32 are evaluated using the TSMC 16FFC technology node. Orégane Desrentes, Benoît Dupont de Dinechin, Florent de Dinechin |
ARITH | 3 |
| 2023 | Audio DSP to FPGA CompilationabstractThe implementation of real-time audio Digital Signal Processing (DSP) applications on FPGA has been extensively studied in the past. Up to now, Audio IPs11Throughout this paper, IP stands for Intellectual Property, i.e., a circuit component. were designed either “by hand” in VHDL or using predefined IPs in block synthesis environments. The advent of High Level Synthesis (HLS) allows for a real compilation flow from high-level audio DSP specifications down to FPGA bit-streams. This paper presents the principles and the implementation of the first “audio DSP compiler” targeting FPGAs. Our fully open-source system compiles audio DSP programs down to FPGA hardware and up to actual sound production. This compilation flow presents two important technological breakthroughs for audio programmers: achieving ultra-low latency real-time audio DSP (few micro-seconds) and the possibility of easily deploying systems with a large number of audio channels. Maxime Popoff, Romain Michon, Tanguy Risset, Pierre Cochard, Stéphane Letz, Yann Orlarey, Florent de Dinechin |
ASAP | 7 |
| 2022 | Low-precision logarithmic arithmetic for neural network acceleratorsabstractResource requirements for hardware acceleration of neural networks inference is notoriously high, both in terms of computation and storage. One way to mitigate this issue is to quantize parameters and activations. This is usually done by scaling and centering the distributions of weights and activations, on a kernel per kernel basis, so that a low-precision binary integer representation can be used. This work studies low-precision logarithmic number system (LNS) as an efficient alternative. Firstly, LNS has more dynamic than fixed-point for the same number of bits. Thus, when quantizing MNIST and CIFAR reference networks without retraining, the smallest format size achieving top-1 accuracy comparable to floating-point is 1 to 3 bits smaller with LNS than with fixed-point. In addition, it is shown that the zero bit of classical LNS is not needed in this context, and that the sign bit can be saved for activations. The proposed LNS neuron is detailed and its implementation on FPGA is shown to be smaller and faster than a fixed-point one for comparable accuracy. Secondly, low-precision LNS enables efficient inference architectures where 1 / multiplications reduce to additions; 2/ the weighted inputs are converted to classical linear domain, but the tables needed for this conversion remain very small thanks to the low precision; and 3/ the conversion of the output activation back to LNS can be merged with an arbitrary activation function. Maxime Christ, Florent de Dinechin, Frédéric Pétrot |
ASAP | 2 |
| 2022 | Resource Optimal Squarers for FPGAsabstractSquaring is an essential operation in computer arithmetic that can be considered as a special case of multiplication where several simplifications can be applied to reduce the complexity of the resulting circuit. However, the design of a squarer is not straightforward for modern FPGAs that provide embedded DSP blocks and look-up-tables (LUTs). This work proposes a flexible method to design resource optimal squarers, i.e., a squarer that uses a minimum number of LUTs for a user-defined number of DSP blocks. The method uses an integer linear programming (ILP) formulation based on a generalization of multiplier tiling. It is shown that the proposed squarer design method significantly improves the LUT utilization for a given number of DSPs over previous methods, while maintaining a similar critical path delay and latency. Andreas Böttcher, Martin Kumm, Florent de Dinechin |
FPL | 3 |
| 2022 | Using integer linear programming for correctly rounded multipartite architecturesabstractThis article introduces several improvements to the multipartite method, a generic technique for the hardware implementation of numerical functions. A multipartite architecture replaces a table of value with several tables and an adder tree. Here, the optimization of multipartite tables is formalized using Integer Linear Programming so that generic ILP solvers can be used. This improves the quality of faithfully rounded architectures compared to the state of the art. The proposed approach also enables correctly rounded multipartite architectures, providing errorless table compression. This improves the area by a factor 5 without any performance penalty compared with the state of the art in errorless compression. Another improvement of the proposed work is a cost function that attempts to predict the total cost of an architecture in FPGA architectural LUTs, where most of the previous works only count the size of the tables, thus ignoring the cost of the adder tree. Orégane Desrentes, Florent de Dinechin |
FPT | 2 |
| 2021 | Resource Optimal Truncated Multipliers for FPGAsabstractThis proposal presents the resource optimal design of truncated multipliers targeting field programmable gate arrays (FPGAs). In contrast to application specific integrated circuits (ASICs), the design for FPGAs has some distinct design challenges due to many possibilities of computing the partial products using logic-based or DSP-based sub-multipliers. To tackle this, we extend a previously proposed tiling methodology which translates the multiplier design into a geometrical problem: the target multiplier is represented by a board that has to be covered by tiles representing the sub-multipliers. The tiling with the least resources can be found with integer linear programming (ILP). Our extension considers the error of possibly unoccupied positions of the board and determines the tiling with the least resources that respects the maximal allowed error bound. This error bound is chosen such that a faithfully rounded truncated multiplier is obtained. Compared to previous designs that use a fixed number of guard bits or optimize at the level of the dot diagrams, this allows a much better use of sub-multipliers resulting in significant area savings without sacrificing the timing. Andreas Böttcher, Martin Kumm, Florent de Dinechin |
ARITH | 3 |
| 2021 | Towards Arithmetic-Centered Filter DesignabstractA hardware implementation can be defined to be faithful to the frequency specification of a linear time-invariant digital filter. Filter design and implementation then become a single global optimisation problem. To solve this problem, existing tools are reviewed, and the missing ones are framed. Florent de Dinechin, Silviu-Ioan Filip, Martin Kumm, Anastasia Volkova 0001 |
ARITH | 1 |
| 2020 | Next Generation Arithmetic for Edge ComputingabstractArithmetic is a key component and is ubiquitous in today’s digital world, ranging from embedded to high-performance computing systems. With machine learning at the fore in a wide range of application domains from wearables to automotive to avionics to weather prediction, sufficiently accurate yet low-cost arithmetic is the need for the day. Recently, there have been several advances in the domain of computer arithmetic, which includes high-precision anchored numbers from ARM, posit arithmetic, bfloat16, etc. as an alternative to IEEE 754-2008 compliant arithmetic. Optimizations on fixed-point and integer arithmetic are also being pursued actively for low-power computing architectures. Furthermore, approximate computing and transprecision/mixed-precision computing have been exciting areas of research forever. While academic research in the domain of computer arithmetic has a long history, industrial adoption of some of these new data types and techniques is in its early stages and expected to increase in the future. bfloat16 is an excellent example for this. In this paper, we bring academia and industry together to discuss the latest results and future directions for research in the domain of next-generation computer arithmetic, especially for edge computing. Andre Guntoro, Cecilia De la Parra, Farhad Merchant, Florent de Dinechin, John L. Gustafson, Martin Langhammer, Rainer Leupers, Sangeeth Nambiar |
DATE | 4 |
| 2020 | Application-Specific Arithmetic in High-Level Synthesis ToolsabstractThis work studies hardware-specific optimization opportunities currently unexploited by high-level synthesis compilers. Some of these optimizations are specializations of floating-point operations that respect the usual semantics of the input program without changing the numerical result. Some other optimizations, locally triggered by the programmer thanks to a pragma, assume a different semantics, where floating-point code is interpreted as the specification of computation with real numbers. The compiler is then in charge to ensure an application-level accuracy constraint expressed in the pragma and has the freedom to use non-standard arithmetic hardware when more efficient. These two classes of optimizations are prototyped in the GeCoS source-to-source compiler and evaluated on the Polybench and EEMBC benchmark suites. Latency is reduced by up to 93%, and resource usage is reduced by up to 58%. Yohann Uguen, Florent de Dinechin, Victor Lezaud, Steven Derrien |
ACM Trans. Archit. Code Optim. | 2 |
| 2019 | Dynamic Precision Numerics Using a Variable-Precision UNUM Type I HW CoprocessorabstractA very large internal accumulation register has been proposed to increase the accuracy of scientific code. However, there is a general class of iterative kernels where a vector of high-precision data must be saved from one iteration to the next. Saving the large internal accumulator to memory is impractical in such cases. This work proposes a Variable Precision (VP) Floating Point (FP) arithmetic co-processor architecture based on RISC-V, which 1/ supports legacy IEEE formats for input and output variables, 2/ uses variable length internal registers (up to 512 bits of mantissa) for inner loop multiply-add and 3/ supports loads and stores of intermediate results to cache memory with a dynamically adjustable precision (up to 256 bits of mantissa). It exploits the UNUM type I floating point format, proposing solutions to address some of its pitfalls such as the variable latency of the internal operation, and the variable memory footprint of the intermediate variables. This work is integrated on FPGA and demonstrated on a representative example. Andrea Bocco, Yves Durand, Florent de Dinechin |
ARITH | 3 |
| 2019 | Reflections on 10 Years of FloPoCoabstractThe FloPoCo open-source arithmetic core generator project started modestly in 2008 with a few parametric floating point cores. It has since then evolved to become a framework for research on hardware arithmetic cores at large. This article reviews some of the choices made in this past decade, and some of the challenges ahead for the next. Florent de Dinechin |
ARITH | 1 |
| 2019 | Table-Based versus Shift-And-Add Constant Multipliers for FPGAsabstractThe multiplication by a constant is a frequently used operation. To implement it on Field Programmable Gate Arrays (FPGAs), the state of the art offers two completely different methods: one relying on bit shifts and additions/subtractions, and another one using look-up tables and additions. So far, it was unclear which method performs best for a given constant and input/output data types. The main contribution of this work is a thorough comparison of both methods in the main application contexts of constant multiplication: filters, signal-processing transforms, and elementary functions. Most of the previous state of the art addresses multiplication by an integer constant. This work shows that, in most of these application contexts, a formulation of the problem as the multiplication by a real constant allows for more efficient architectures. Another contribution is a novel extension of the shift-and-add method to real constants. For that, an integer linear programming (ILP) formulation is proposed, which truncates each component in the shift-and-add network to a minimum necessary word size that is aligned with the approximation error of the coefficient. All methods are implemented within the open-source FloPoCo framework. Florent de Dinechin, Silviu-Ioan Filip, Martin Kumm, Luc Forget |
ARITH | 1 |
| 2019 | Evaluating the Hardware Cost of the Posit Number SystemabstractThe posit number system is proposed as a replacement of IEEE floating-point numbers. It is a floating-point system that trades exponent bits for significand bits, depending on the magnitude of the numbers. Thus, it provides more precision for numbers around 1, at the expense of lower precision for very large or very small numbers. Several works have demonstrated that this trade-off can improve the accuracy of applications. However, the variable-length exponent and significand encoding impacts the hardware cost of posit arithmetic. The objective of the present work is to enable application-level evaluations of the posit system that include performance and resource consumption. To this purpose, this article introduces an open-source hardware implementation of the posit number system, in the form of a C++ templatized library compatible with Vivado HLS. This library currently implements addition, subtraction and multiplication for custom-size posits. In addition, the posit standard also mandates the presence of the "quire", a large accumulator able to perform exact sums of products. The proposed library includes the first open-source parameterized hardware quire. This library is shown to improve the state-of-the-art of posit implementations in terms of latency and resource consumption. Still, standard 32 bits posit adders and multipliers are found to be much larger and slower than the corresponding floating-point operators. The cost of the posit 32 quire is shown to be comparable to that of a Kulisch accumulator for 32 bits floating-point. Yohann Uguen, Luc Forget, Florent de Dinechin |
FPL | 3 |
| 2019 | Byte-Aware Floating-point Operations through a UNUM Computing UnitabstractMost floating-point (FP) hardware support the IEEE 754 format, which defines fixed-size data types from 16 to 128 bits. However, a range of applications benefit from different formats, implementing different tradeoffs. This paper proposes a Variable Precision (VP) computing unit offering a finer granularity of high precision FP operations. The chosen memory format is derived from UNUM type I, where the size of a number is stored within the representation itself. The unit implements a fully pipelined architecture, and it supports up to 512 bits of precision for both interval and scalar computing. The user can conFigure the storage format up to 8-bit granularity, and the internal computing precision at 64-bit granularity. The system is integrated as a RISC-V coprocessor. Dedicated compiler support exposes the unit through a high level programming abstraction, covering all the operating features of UNUM type I. FPGA-based measurements show that the latency and the computation accuracy of this system scale linearly with the memory format length set by the user. Compared with a highly optimized software implementation, the proposed unit achieves speedups between 3.5 × and 18 ×, with comparable accuracy. Andrea Bocco, Tiago T. Jost, Albert Cohen 0001, Florent de Dinechin, Yves Durand, Christian Fabre |
VLSI-SoC | 4 |
| 2019 | Guest Editors Introduction: Special Section on Computer ArithmeticabstractThe papers in this special section examine the concept of computer arithmetic. Many services offered in the palm of our hand by today’s computing devices were undreamt of twenty years ago, and we probably don’t envision what services will be enabled twenty years from now. We even can’t be sure of the technology they will use, if good old silicon integration is no longer able to sustain Moore’s law. However, one can be confident that there will be computers, and that these computers will compute, and that at the core of these computations there will be adders, multipliers, elementary functions and other core arithmetic primitives. Computer arithmetic is the art of designing and using these core arithmetic primitives. It studies the representation of numbers in computers, and the transformation of these “machine numbers”. With the abacus and early mechanical calculators, it actually predates the computing era. Computer arithmetic has accompanied the evolutions of technology (from relays to vacuum tube to transistors and integrated circuits). It has also adapted to the evolution of applications: scientific computing, digital signal processing, cryptography, or machine learning use different kinds of numbers and operations. Formal proofs involving computer arithmetic components have become a major concern of many other applications. Javier D. Bruguera, Florent de Dinechin |
IEEE Trans. Computers | 2 |
| 2019 | Towards Hardware IIR Filters Computing Just Right: Direct Form I Case StudyabstractLinear Time Invariant (LTI) filters are often specified and simulated using high-precision software, before being implemented in low-precision fixed-point hardware. A problem is that the hardware does not behave exactly as the simulation due to quantization and rounding issues. This article advocates the construction of LTI architectures that behave as if the computation was performed with infinite accuracy, then converted to the low-precision output format with an error smaller than its least significant bit. This simple specification guarantees the numerical quality of the hardware, even for critical LTI systems. Besides, it is possible to derive the optimal values of all the internal data formats that ensure that the specification is met. This requires a detailed error analysis that captures not only the quantization and rounding errors, but also their infinite accumulation in recursive filters. This generic methodology is detailed for the case of low-precision LTI filters in the Direct Form I implemented in FPGA logic. It is demonstrated by a fully automated and open-source architecture generator tool, and validated on a range of Infinite Impulse Response filters. Anastasia Volkova 0001, Matei Istoan, Florent de Dinechin, Thibault Hilaire |
IEEE Trans. Computers | 3 |
| 2018 | Karatsuba with Rectangular Multipliers for FPGAsabstractThis work presents an extension of Karatsuba's method to efficiently use rectangular multipliers as a base for larger multipliers. The rectangular multipliers that motivate this work are the embedded 18 × 25-bit signed multipliers found in the DSP blocks of recent Xilinx FPGAs: The traditional Karatsuba approach must under-use them as square 18 × 18 ones. This work shows that rectangular multipliers can be efficiently exploited in a modified Karatsuba method if their input word sizes have a large greatest common divider. In the Xilinx FPG A case, this can be obtained by using the embedded multipliers as 16 × 24 unsigned and as 17 × 25 signed ones. The obtained architectures are implemented with due detail to architectural features such as the pre-adders and post-adders available in Xilinx DSP blocks. They are synthesized and compared with traditional Karatsuba, but also with (non-Karatsuba) state-of-the-art tiling techniques that make use of the full rectangular multipliers. The proposed technique improves resource consumption and performance for multipliers of numbers larger than 64 bits. Martin Kumm, Oscar Gustafsson, Florent de Dinechin, Johannes Kappauf, Peter Zipf |
ARITH | 3 |
| 2017 | Automating the pipeline of arithmetic datapathsabstractThis article presents the new framework for semi-automatic circuit pipelining that will be used in future releases of the FloPoCo generator. From a single description of an operator or datapath, optimized implementations are obtained automatically for a wide range of FPGA targets and a wide range of frequency/latency trade-offs. Compared to previous versions of FloPoCo, the level of abstraction has been raised, enabling easier development, shorter generator code, and better pipeline optimization. The proposed approach is also more flexible than fully automatic pipelining approaches based on retiming: In the proposed technique, the incremental construction of the pipeline along with the circuit graph enables architectural design decisions that depend on the pipeline. Matei Istoan, Florent de Dinechin |
DATE | 2 |
| 2017 | A High-Level Synthesis Approach Optimizing Accumulations in Floating-Point Programs Using Custom Formats and OperatorsabstractMany case studies have demonstrated the potential of Field-Programmable Gate Arrays (FPGAs) as accelerators for a wide range of applications. FPGAs offer massive parallelism and programmability at the bit level. This enables programmers to exploit a range of techniques that avoid many bottlenecks of classical von Neumann computing. However, development costs for FPGAs are orders of magnitude higher than classical programming. A solution would be the use of High-Level Synthesis (HLS) tools, which use C as a hardware description language. However, the C language was designed to be executed on general purpose processors, not to generate hardware. Its datatypes and operators are limited to a small number (more or less matching the hardware operators present in mainstream processors), and HLS tools inherit these limitations. To better exploit the freedom offered by hardware and FPGAs, HLS vendors have enriched the C language with integer and fixed-point types of arbitrary size. Still, the operations on these types remain limited to the basic arithmetic and logic ones. In floating point, the current situation is even worse. The operator set is limited, and the sizes are restricted to 32 and 64 bits. Besides, most recent compilers, including the HLS ones, attempt to follow established standards, in particular C11 and IEEE-754. This ensures bit-exact compatibility with software, but greatly reduces the freedom of optimization by the compiler. For instance, a floating point addition is not associative even though its real equivalent is. In the present work we attempt to give the compiler more freedom. For this, we sacrifice the strict respect of the IEEE-754 and C11 standards, but we replace it with the strict respect of a high-level accuracy specification expressed by the programmer through a pragma. The case study in this work is a program transformation that applies to floating-point additions on a loop's critical path. It decomposes them into elementary steps, resizes the corresponding subcomponents to guarantee some user-specified accuracy, and merges and reorders these components to improve performance. The result of this complex sequence of optimizations could not be obtained from an operator generator, since it involves global loop information. For this purpose, we used a compilation flow involving one or several source-to-source transformations operating on the code given to HLS tools (Figure 1).The proposed transformation already works very well on 3 of the 10 FPMarks where it improves both latency and accuracy by an order of magnitude for comparable area. For 2 more benchmarks, the latency is not improved (but not degraded either) due to current limitations of HLS tools. This defines short-term future work. The main result of this work is that HLS tools also have the potential to generate efficient designs for handling floating-point computations in a completely non-standard way. In the longer term, we believe that HLS flows can not only import application-specific operators from the FPGA literature, they can also improve them using high-level, program-level information. Yohann Uguen, Florent de Dinechin, Steven Derrien |
FCCM | 2 |
| 2017 | Bridging high-level synthesis and application-specific arithmetic: The case study of floating-point summationsabstractFPGAs are well known for their ability to perform non-standard computations not supported by classical microprocessors. Many libraries of highly customizable application-specific IPs have exploited this capability. However, using such IPs usually requires handcrafted HDL, hence significant design efforts. High Level Synthesis (HLS) lowers the design effort thanks to the use of C/C++ dialects for programming FPGAs. However, high-level C language becomes a hindrance when one wants to express non-standard computations: this languages was designed for programming microprocessors and carries with it many restrictions due to this paradigm. This is especially true when computing with floating-point, whose data-types and evaluation semantics are defined by the IEEE-754 and C11 standards. If the high-level specification was a computation on the reals, then HLS imposes a very restricted implementation space. This work attempts to bridge FPGA application-specific efficiency and HLS ease of use. It specifically targets the ubiquitous floating-point summation-reduction pattern. A source-to-source compiler transforms selected floating-point additions into sequences of simpler operators using non-standard arithmetic formats. This improves performance and accuracy for several benchmarks, while keeping the ease of use of a high-level C description. Yohann Uguen, Florent de Dinechin, Steven Derrien |
FPL | 2 |
| 2017 | Hardware Division by Small Integer ConstantsabstractThis article studies the design of custom circuits for division by a small positive constant. Such circuits can be useful for specific FPGA and ASIC applications. The first problem studied is the Euclidean division of an unsigned integer by a constant, computing a quotient and remainder. Several new solutions are proposed and compared against the state-of-the-art. As the proposed solutions use small look-up tables, they match well with the hardware resources of an FPGA. The article then studies whether the division by the product of two constants is better implemented as two successive dividers or as one atomic divider. It also considers the case when only a quotient or only a remainder is needed. Finally, it addresses the correct rounding of the division of a floating-point number by a small integer constant. All these solutions, and the previous state-of-the-art, are compared in terms of timing, area, and area-timing product. In general, the relevance domains of the various techniques are different on FPGA and on ASIC. H. Fatih Ugurdag, Florent de Dinechin, Serhan Gener, Sezer Gören 0001, Laurent-Stéphane Didier |
IEEE Trans. Computers | 2 |
| 2016 | Computing floating-point logarithms with fixed-point operationsabstractElementary functions from the mathematical library input and output floating-point numbers. However it is possible to implement them purely using integer/fixed-point arithmetic. This option was not attractive between 1985 and 2005, because mainstream processor hardware supported 64-bit floating-point, but only 32-bit integers. This has changed in recent years, in particular with the generalization of native 64-bit integer support. The purpose of this article is therefore to reevaluate the relevance of computing floating-point functions in fixed-point. For this, several variants of the double-precision logarithm function are implemented and evaluated. Formulating the problem as a fixed-point one is easy after the range has been (classically) reduced. Then, 64-bit integers provide slightly more accuracy than 53-bit mantissa, which helps speed up the evaluation. Finally, multi-word arithmetic, critical for accurate implementations, is much faster in fixed-point, and natively supported by recent compilers. Thanks to all this, a purely integer implementation of the correctly rounded double-precision logarithm outperforms the previous state of the art, with the worst-case execution time reduced by a factor 5. This work also introduces variants of the logarithm that input a floating-point number and output the result in fixed-point. These are shown to be both more accurate and more efficient than the traditional floating-point functions for some applications. Julien Le Maire, Nicolas Brunie, Florent de Dinechin, Jean-Michel Muller |
ARITH | 3 |
| 2015 | Code Generators for Mathematical FunctionsabstractA typical floating-point environment includes support for a small set of about 30 mathematical functions such as exponential, logarithm, trigonometric and hyperbolic functions. These functions are provided by mathematical software libraries (libm), typically in IEEE754 single, double and quad precision. This article suggests to replace this libm paradigm by a more general approach: the on-demand generation of numerical function code, on arbitrary domains and with arbitrary accuracies. First, such code generation opens up the libm function space available to programmers. It may capture a much wider set of functions, and may capture even standard functions on non-standard domains and accuracy/performance points. Second, writing libm code requires fine-tuned instruction selection and scheduling for performance, and sophisticated floating-point techniques for accuracy. Automating this task through code generation improves confidence in the code while enabling better design space exploration, and therefore better time to market, even for the libm functions. This article discusses the new challenges of this paradigm shift, and presents the current state of open-source function code generators available on http://www.metalibm.org/. Nicolas Brunie, Florent de Dinechin, Olga Kupriianova, Christoph Quirin Lauter |
ARITH | 2 |
| 2015 | Hardware Implementations of Fixed-Point Atan2abstractThe atan2 function computes the polar angle arctan(y/x) of a point given by its cartesian coordinates. It is widely used in digital signal processing to recover the phase of a signal. This article studies for this context the implementation of atan2 with fixed-point inputs and outputs. It compares the prevalent CORDIC shift-and-add algorithm to two multiplier-based techniques. The first one computes the bivariate atan2 function as the composition of two univariate functions: the reciprocal, and the arctangent, each evaluated using bipartite or polynomial approximation methods. The second technique directly uses piecewise bivariate polynomial approximations of degree 1 or 2. Each of these approaches requires a relevant argument reduction, which is also discussed. All the algorithms are last-bit accurate, and implemented with similar care in the open-source FloPoCo framework. Based on synthesis results on FPGAs, their relevance domains are discussed. Florent de Dinechin, Matei Istoan |
ARITH | 1 |
| 2014 | Sum-of-product architectures computing just rightabstractMany digital filters and signal-processing transforms can be expressed as a sum of products with constants (SPC). This paper addresses the automatic construction of low-precision, but high accuracy SPC architectures: these architectures are specified as last-bit accurate with respect to a mathematical definition. In other words, they behave as if the computation was performed with infinite accuracy, then rounded only once to the low-precision output format. This eases the task of porting double-precision code (e.g. Matlab) to low-precision hardware or FPGA. The paper further discusses the construction of the most efficient architectures obeying such a specification, introducing several architectural improvements to this purpose. This approach is demonstrated in a generic, open-source architecture generator tool built upon the FloPoCo framework. It is evaluated on Finite Impulse Response filters for the ZigBee protocol. Florent de Dinechin, Matei Istoan, Abdelbassat Massouri |
ASAP | 1 |
| 2013 | Arithmetic core generation using bit heapsabstractA bit heap is a data structure that holds the unevaluated sum of an arbitrary number of bits, each weighted by some power of two. Most advanced arithmetic cores can be viewed as involving one or several bit heaps. We claim here that this point of view leads to better global optimization at the algebraic level, at the circuit level, and in terms of software engineering. To demonstrate it, a generic software framework is introduced for the definition and optimization of bit heaps. This framework, targeting DSP-enabled FPGAs, is developed within the open-source FloPoCo arithmetic core generator. Its versatility is demonstrated on several examples: multipliers, complex multipliers, polynomials, and discrete cosine transform. Nicolas Brunie, Florent de Dinechin, Matei Istoan, Guillaume Sergent, Kinga Illyes, Bogdan Popa |
FPL | 2 |
| 2013 | On Ziv's rounding testabstractA very simple test, introduced by Ziv, allows one to determine if an approximation to the value f(x) of an elementary function at a given point x suffices to return the floating-point number nearest f(x) . The same test may be used when implementing floating-point operations with input and output operands of different formats, using arithmetic operators tailored for manipulating operands of the same format. That test depends on a “magic constant” e . We show how to choose that constant e to make the test reliable and efficient. Various cases are considered, depending on the availability of an fma instruction, and on the range of f(x) . Florent de Dinechin, Christoph Quirin Lauter, Jean-Michel Muller, Serge Torres |
ACM Trans. Math. Softw. | 1 |
| 2013 | Floating-Point Exponentiation Units for Reconfigurable ComputingabstractThe high performance and capacity of current FPGAs makes them suitable as acceleration co-processors. This article studies the implementation, for such accelerators, of the floating-point power function x y as defined by the C99 and IEEE 754-2008 standards, generalized here to arbitrary exponent and mantissa sizes. Last-bit accuracy at the smallest possible cost is obtained thanks to a careful study of the various subcomponents: a floating-point logarithm, a modified floating-point exponential, and a truncated floating-point multiplier. A parameterized architecture generator in the open-source FloPoCo project is presented in details and evaluated. Florent de Dinechin, Pedro Echeverría, Marisa López-Vallejo, Bogdan Pasca 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2011 | The Arithmetic Operators You Will Never See in a MicroprocessorabstractIt has been shown that FPGAs could outperform high-end microprocessors even on floating-point computations, thanks to massive parallelism. Too often, however, such studies re-implement in the FPGA the operators present in a processor. An FPGA can do much better: it can accomodate hardware operators that would make no economical sense in a general-purpose processor, and it can taylor them just right to the needs of the application. This talk tries to survey this idea systematically, discussing its potential, exhibiting some exotic (but useful) operators developed in the FloPoCo project, and listing some of the challenges ahead. Florent de Dinechin |
IEEE Symposium on Computer Arithmetic | 1 |
| 2011 | An FPGA architecture for solving the Table Maker's DilemmaabstractSolving the Table Maker's Dilemma, for a given function and a given target floating-point format, requires testing the value of the function, with high precision, at a very large number of consecutive values. We give an algorithm that allows for performing such computations on a very regular architecture, and present an FPGA implementation of that algorithm. Florent de Dinechin, Jean-Michel Muller, Bogdan Pasca 0001, Alexandru Plesco |
ASAP | 1 |
| 2011 | Certifying the Floating-Point Implementation of an Elementary Function Using GappaabstractHigh confidence in floating-point programs requires proving numerical properties of final and intermediate values. One may need to guarantee that a value stays within some range, or that the error relative to some ideal value is well bounded. This certification may require a time-consuming proof for each line of code, and it is usually broken by the smallest change to the code, e.g., for maintenance or optimization purpose. Certifying floating-point programs by hand is, therefore, very tedious and error-prone. The Gappa proof assistant is designed to make this task both easier and more secure, due to the following novel features: It automates the evaluation and propagation of rounding errors using interval arithmetic. Its input format is very close to the actual code to validate. It can be used incrementally to prove complex mathematical properties pertaining to the code. It generates a formal proof of the results, which can be checked independently by a lower level proof assistant like Coq. Yet it does not require any specific knowledge about automatic theorem proving, and thus, is accessible to a wide community. This paper demonstrates the practical use of this tool for a widely used class of floating-point programs: implementations of elementary functions in a mathematical library. Florent de Dinechin, Christoph Quirin Lauter, Guillaume Melquiond |
IEEE Trans. Computers | 1 |
| 2010 | Automatic generation of polynomial-based hardware architectures for function evaluationabstractPolynomial approximation is a very general technique for the evaluation of a wide class of numerical functions of one variable. This article details an architecture generator that inputs the specification of a function and outputs a synthe-sizable description of an architecture evaluating this function with guaranteed accuracy. It improves upon the literature in two aspects. Firstly, it uses better polynomials, thanks to recent advances related to constrained-coefficient polynomial approximation. Secondly, it refines the error analysis of polynomial evaluation to reduce the size of the multipliers used. An open-source implementation is provided in the FloPoCo project, including architecture exploration heuristics designed to use efficiently the embedded memories and multipliers of high-end FPGAs. High-performance pipelined architectures for precisions up to 64 bits can be obtained in seconds. Florent de Dinechin, Mioara Joldes, Bogdan Pasca 0001 |
ASAP | 1 |
| 2010 | Multiplicative Square Root Algorithms for FPGAsabstractMost current square root implementations for FPGAs use a digit recurrence algorithm which is well suited to their LUT structure. However, recent computing-oriented FPGAs include embedded multipliers and RAM blocks which can also be used to implement quadratic convergence algorithms, very high radix digit recurrences, or polynomial approximation algorithms. The cost of these solutions is evaluated and compared, and a complete implementation of a polynomial approach is presented within the open-source FloPoCo framework. This polynomial approach allows a shorter latency and higher frequency than the digit recurrence approach, and improves over previous multiplicative approaches. However, the cost of IEEE-compliant correct rounding is shown to be very high. Florent de Dinechin, Mioara Joldes, Bogdan Pasca 0001, Guillaume Revy |
FPL | 1 |
| 2010 | Pipelined FPGA AddersabstractInteger addition is a universal building block, and applications such as quad-precision floating-point or elliptic curve cryptography now demand precisions well beyond 64 bits. This study explores the trade-offs between size, latency and frequency for pipelined large-precision adders on FPGA. It compares three pipelined adder architectures: the classical pipelined ripple-carry adder, a variation that reduces register count, and an FPGA-specific implementation of the carry-select adder capable of providing lower latency additions at a comparable price. For each of these architectures, resource estimation models are defined, and used in an adder generator that selects the best architecture considering the target FPGA, the target operating frequency, and the addition bit width. Florent de Dinechin, Hong Diep Nguyen, Bogdan Pasca 0001 |
FPL | 1 |
| 2010 | Floating-point exponential functions for DSP-enabled FPGAsabstractThis article presents a generator of floating-point exponential operators targeting recent FPGAs with embedded memories and DSP blocks. A single-precision operator consumes just one DSP block, 18Kbits of dual-port memory, and 392 slices on Virtex-4. For larger precisions, a generic approach based on polynomial approximation is used and proves more resource-efficient than the literature. For instance a double-precision operator consumes 5 BlockRAM and 12 DSP48 blocks on Virtex-5, or 10 M9k and 22 18×18 multipliers on Stratix III. This approach is flexible and is demonstrated to scale up to quadruple-precision, while enabling frequencies close to the FPGA's nominal frequency. All the proposed architectures are last-bit accurate for all the floating-point range. They are available in the open-source FloPoCo framework. Florent de Dinechin, Bogdan Pasca 0001 |
FPT | 1 |
| 2010 | Efficient implementation of parallel BCD multiplication in LUT-6 FPGAsabstractDecimal multiplication is one of the most frequent operations used by many financial, business and user-oriented applications but current implementations in FPGAs are very inefficient in terms of both area and latency when compared to binary multipliers. In this paper we present a new method for implementing BCD multiplication more efficiently than previous proposals in current FPGA devices with 6-input LUTs. In particular, a combinational implementation maps quite well into the slice structure of the Xilinx Virtex-5/Virtex-6 families and it is highly pipelineable. The synthesis results for a Virtex-6 device indicate that our proposal outperforms the area and latency figures of previous implementations in FPGAs. Álvaro Vázquez, Florent de Dinechin |
FPT | 2 |
| 2009 | Generating high-performance custom floating-point pipelinesabstractCustom operators, working at custom precisions, are a key ingredient to fully exploit the FPGA flexibility advantage for high-performance computing. Unfortunately, such operators are costly to design, and application designers tend to rely on less efficient off-the-shelf operators. To address this issue, an open-source architecture generator framework is introduced. Its salient features are an easy learning curve from VHDL, the ability to embed arbitrary synthesizable VHDL code, portability to mainstream FPGA targets from Xilinx and Altera, automatic management of complex pipelines with support for frequency-directed pipeline, and automatic test-bench generation. This generator is presented around the simple example of a collision detector, which it significantly improves in accuracy, DSP count, logic usage, frequency and latency with respect to an implementation using standard floating-point operators. Florent de Dinechin, Cristian Klein, Bogdan Pasca 0001 |
FPL | 1 |
| 2009 | Large multipliers with fewer DSP blocksabstractRecent computing-oriented FPGAs feature DSP blocks including small embedded multipliers. A large integer multiplier, for instance for a double-precision floating-point multiplier, consumes many of these DSP blocks. This article studies three non-standard implementation techniques of large multipliers: the Karatsuba-Ofman algorithm, non-standard multiplier tiling, and specialized squarers. They allow for large multipliers working at the peak frequency of the DSP blocks while reducing the DSP block usage. Their overhead in term of logic resources, if any, is much lower than that of emulating embedded multipliers. Their latency overhead, if any, is very small. Complete algorithmic descriptions are provided, carefully mapped on recent Xilinx and Altera devices, and validated by synthesis results. Florent de Dinechin, Bogdan Pasca 0001 |
FPL | 1 |
| 2008 | Integer and floating-point constant multipliers for FPGAsabstractReconfigurable circuits now have a capacity that allows them to be used as floating-point accelerators. They offer massive parallelism, but also the opportunity to design optimised floating-point hardware operators not available in microprocessors. Multiplication by a constant is an important example of such an operator. This article presents an architecture generator for the correctly rounded multiplication of a floating-point number by a constant. This constant can be a floating-point value, but also an arbitrary irrational number. The multiplication of the significands is an instance of the well-studied problem of constant integer multiplication, for which improvement to existing algorithms are also proposed and evaluated. Nicolas Brisebarre, Florent de Dinechin, Jean-Michel Muller |
ASAP | 2 |
| 2008 | When FPGAs are better at floating-point than microprocessorsabstractIt has been shown that FPGAs could outperform high-end microprocessors on floating-point computations thanks to massive parallelism. However, most previous studies re-implement in the FPGA the operators present in a processor. This conservative approach is relatively straightforward, but it doesn't exploit the greater flexibility of the FPGA. We survey the many ways in which the FPGA implementation of a given floating-point computation can be not only faster, but also more accurate than its microprocessor counterpart. Techniques studied here include custom precision, mixing and matching fixed- and floating-point, specific accumulator design, dedicated architectures for coarser operators implemented as software in processors (such as elementary functions or Euclidean norms), operator specialization such as constant multiplication, and others. The FloPoCo project (http://www.ens-lyon.fr/LIP/Arenaire/Ware/FloPoCo/) aims at providing such non-standard operators. As a conclusion, current FPGA fabrics could be enhanced to improve floating-point performance. However, these enhancements should not take the form of hard FPU blocks as others have suggested. Instead, what is needed is smaller building blocks more generally useful to the implementation of floating-point operators, such as cascadable barrel shifters and leading zero counters Florent de Dinechin, Jérémie Detrey, Octavian Cret, Radu Tudoran |
FPGA | 1 |
| 2008 | An FPGA-specific approach to floating-point accumulation and sum-of-productsabstractThis article studies two common situations where the flexibility of FPGAs allows one to design application-specific floating-point operators which are more efficient and more accurate than those offered by processors and GPUs. First, for applications involving the addition of a large number of floating-point values, an ad-hoc accumulator is proposed. By tailoring its parameters to the numerical requirements of the application, it can be made arbitrarily accurate, at an area cost comparable to that of a standard floating-point adder, and at a higher frequency. The second example is the sum-of-product operation, which is the building block of matrix computations. A novel architecture is proposed that feeds the previous accumulator out of a floating-point multiplier whose rounding logic has been removed, again improving the area/accuracy tradeoff. These architectures are implemented within the FloPoCo generator, freely available under the LGPL. Florent de Dinechin, Bogdan Pasca 0001, Octavian Cret, Radu Tudoran |
FPT | 1 |
| 2007 | Return of the hardware floating-point elementary functionabstractThe study of specific hardware circuits for the evaluation of floating-point elementary functions was once an active research area, until it was realized that these functions were not frequent enough to justify dedicating silicon to them. Research then turned to software functions. This situation may be about to change again with the advent of reconfigurable co-processors based on field-programmable gate arrays. Such co-processors now have a capacity that allows them to accomodate double-precision floating-point computing. Hardware operators for elementary functions targeted to such platforms have the potential to vastly outperform software functions, and will not permanently waste silicon resources. This article studies the optimization, for this target technology, of operators for the exponential and logarithm functions up to double-precision. These operators are freely available from www.ens-lyon.fr/LIP/ Arenaire/. Jérémie Detrey, Florent de Dinechin, Xavier Pujol |
IEEE Symposium on Computer Arithmetic | 2 |
| 2007 | Floating-Point Trigonometric Functions for FPGAsabstractField-programmable circuits now have a capacity that allows them to accelerate floating-point computing, but are still missing core libraries for it. In particular, there is a need for an equivalent to the mathematical library (libm) available with every processor and providing implementations of standard elementary functions such as exponential, logarithm or sine. This is all the more important as FPGAs are able to outperform current processors for such elementary functions, for which no dedicated hardware exists in the processor. FPLibrary, freely available from www.ens-lyon.fr/LIP/Arenaire/, is a first attempt to address this need for a mathematical library for FPGAs. This article demonstrates the implementation, in this library, of high-quality operators for floating-point sine and cosine functions up to single-precision. Small size and high performance are obtained using a specific, hardware-oriented algorithm, and careful datapath optimisation and error analysis. Operators fully compatible with the standard software functions are first presented, followed by a study of several more cost-efficient variants. Jérémie Detrey, Florent de Dinechin |
FPL | 2 |
| 2006 | Floating Point or LNS: Choosing the Right Arithmetic on an Aapplication BasisabstractFor applications requiring a large dynamic range, real numbers may be represented either in floating-point (FP), or in the logarithm number system (LNS). Which system is best for a given application is difficult to know in advance, because the cost and performance of LNS operators depend on the target accuracy in a highly non linear way. In doubt, designers will choose floating-point. This article demonstrates a methodology for a better informed choice thanks to FPLibrary, a freely available, dual FP/LNS arithmetic operator library. FPLibrary may be used in the prototype phase of an application to obtain, with low design effort, accurate measures of performance, cost and accuracy of both LNS and FP approaches. Two case studies demonstrate the benefits of this methodology. Caroline Collange, Jérémie Detrey, Florent de Dinechin |
DSD | 3 |
| 2005 | Towards the Post-Ultimate libmabstractThis article presents advances on the subject of correctly rounded elementary functions since the publication of the libultim mathematical library developed by Ziv at IBM. This library showed that the average performance and memory overhead of correct rounding could be made negligible. However, the worst-case overhead was still a factor 1000 or more. It is shown that, with current processor technology, this worst-case overhead can be kept within a factor of 2 to 10 of current best libms. This low overhead has very positive consequences on the techniques for implementing and proving correctly rounded functions, which are also studied. These results lift the last technical obstacles to a generalisation of (at least some) correctly rounded double precision elementary functions. Florent de Dinechin, Alexey V. Ershov, Nicolas Gast |
IEEE Symposium on Computer Arithmetic | 1 |
| 2005 | Table-based polynomials for fast hardware function evaluationabstractMany general table-based methods for the evaluation in hardware of elementary functions have been published. The bipartite and multipartite methods implement a first-order approximation of the function using only table lookups and additions. Recently, a single multiplier second order method of similar inspiration has also been published. This paper extends such methods to approximations of arbitrary order, using adders, small multipliers, and very small ad hoc powering units. We obtain implementations that are both smaller and faster than previously published approaches. This paper also deals with the FPGA implementation of such methods. Previous work have consistently shown that increasing the approximation degree lead to not only smaller but also faster designs, as the reduction of the table size meant a reduction of its lookup time, which compensated for the addition and multiplication time. The experiments in this paper suggest that this still holds when going from order 2 to order 3, but no longer when using higher order approximations, where a tradeoff appears. Jérémie Detrey, Florent de Dinechin |
ASAP | 2 |
| 2005 | A Parameterized Floating-Point Exponential Function for FPGAs
Jérémie Detrey, Florent de Dinechin |
FPT | 2 |
| 2005 | Multipartite Table MethodsabstractA unified view of most previous table-lookup-and-addition methods (bipartite tables, SBTM, STAM, and multipartite methods) is presented. This unified view allows a more accurate computation of the error entailed by these methods, which enables a wider design space exploration, leading to tables smaller than the best previously published ones by up to 50 percent. The synthesis of these multipartite architectures on Virtex FPGAs is also discussed. Compared to other methods involving multipliers, the multipartite approach offers the best speed/area tradeoff for precisions up to 16 bits. A reference implementation is available at http://www.ens-lyon.fr/LIP/Arenaire/. Florent de Dinechin, Arnaud Tisserand |
IEEE Trans. Computers | 1 |
| 2004 | Second Order Function Approximation Using a Single Multiplication on FPGAs
Jérémie Detrey, Florent de Dinechin |
FPL | 2 |
| 2001 | Some Improvements on Multipartite Table Methods abstractThis paper presents an unified view of most previous table-lookup-and-addition methods: bipartite tables, SBTM, STAM and multipartite methods. This new definition allows a more accurate computation of the error entailed by these methods. Being more general, it also allows an exhaustive design space exploration which has been implemented, and leads to tables smaller than previously published ones by up to 50%. Some results have been synthesised for Virtex FPGAs, and are discussed. Florent de Dinechin, Arnaud Tisserand |
IEEE Symposium on Computer Arithmetic | 1 |
| 1999 | Towards Adaptable Hierarchical Placement for FPGAsabstractField Programmable Gate Arrays (FPGAs) are usually programmed using languages and methods inherited from the domain of VLSI synthesis. These methods, however, have not always been adapted to the new possibilities opened by FPGAs, nor to the new constraints they impose on a design. This paper addresses in particular the issue of laying out the various components of an architecture on an FPGA. The problem is to embed placement information in FPGA-oriented hardware description languages, in a way that is both expressive enough to be useful, and abstract enough to be portable from one FPGA architecture to the other. A generic placement framework is defined to address this problem, and two prototype implementations of this framework are presented, for Xilinx 6200 and Xilinx 4000 devices, on the example of a bit-serial complex multiplier. Florent de Dinechin, Wayne Luk, Steve McKeever |
FPGA | 1 |
| 1997 | Libraries of schedule-free operators in AlphaabstractThis paper presents a method, based on the formalism of affine recurrence equations, for the synthesis of digital circuits exploiting parallelism at the bit-level. In the initial specification of a numerical algorithm, the arithmetic operators are replaced with their yet unscheduled (schedule-free) binary implementation as recurrence equations. This allows a bit-level dependency analysis yielding a bit-parallel array. The method is demonstrated on the example of the matrix-vector product, and discussed. Florent de Dinechin |
ASAP | 1 |
| 1996 | Hierarchical Static Analysis Of Structured Systems Of Affine Recurrence EquationsabstractThe ALPHA language, based on systems of affine recurrence equations over polyhedral domains, allows the expression of complex algorithms as hierarchical, parameterized structures of such systems. This paper discusses the static analysis of ALPHA programs, an extended type-checking process based on the single assignment rule. We present techniques ensuring, on one hand, that a system is valid (with respect to this rule) within a certain range of its parameters, and on the other, that no system is used with invalid values of its parameters. These techniques have been implemented in a versatile tool for ALPHA which is demonstrated on the example of the Gaussian elimination. Florent de Dinechin |
ASAP | 1 |